Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2608.05800v1 Announce Type: new Abstract: Artificial intelligence (AI) systems increasingly mediate decisions affecting individuals and societies. Existing data protection frameworks address certain privacy-related harms, particularly those arising from data leakage, re-identification, and profiling. However, they inadequately capture a more fundamental risk: unreliable or unjustified inference produced by AI systems even when data collection and processing are legitimate. This article argues that modern AI raises distinct concerns of construct validity, confounding, representativeness, distribution shift, and fairness trade-offs that require specialised regulatory attention. In the context of AI, transparency and explainability acquire distinct and significantly more challenging meanings than in conventional software. A substantial body of work in critical data studies and the measurement-theoretic literature has diagnosed these epistemological limitations. This article's contri
arXiv:2608.05656v1 Announce Type: new Abstract: Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks favor technical methods, such as model benchmarks and LLM simulations, often sidelining empirical research with human subjects. To examine this apparent gap in the acceptance of human research, we conduct an expert survey (n=93) and expert interviews (n=17) with AI Safety & Ethics (AISE) researchers from Technical, Sociotechnical, Governance, and Normative backgrounds. Our findings suggest that although there is a consensus that human research is valuable for generating evidence for AISE, its adoption and acceptance are constrained by perceived validity issues, tangible resource barriers, epistemic and personal preferences in methods, and infrastructural constraints from the broader research community. In particular, Technical researchers tend to value human research less and collaborate acros
arXiv:2608.05583v1 Announce Type: new Abstract: As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients' own actions contribute to illness. We investigate how LLMs reason about responsibility and its consequences, tracing their judgments across successive levels, from the behavior, to the resulting illness, to the denial of care. We evaluate a wide range of LLMs, spanning different model families and capability levels, on various clinical vignettes adapted from prior studies. Our results identify a judgment-consequence gap: LLMs largely agree with humans that patients bear responsibility for health-harming behaviors, yet overwhelmingly refuse to let that judgment influence how they allocate scarce resources. Specifically, LLMs default to random allocation, whereas humans consistently favor the less-c
arXiv:2608.05545v1 Announce Type: new Abstract: Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by encouraging uncritical acceptance of AI-generated reasoning. This creates a need for mechanisms that preserve human agency by augmenting metacognition during AI-assisted intellectual work. To address this, we propose the Synthesis-Analysis Reciprocity Model, which views intellectual construction as a reciprocal interaction between Synthesis, which combines components into an artifact, and Analysis, which critically evaluates them against objective indicators and constrains subsequent synthesis. Grounded in this model, we present the Vibe Compiler, a research-logic compiler that helps researchers transform vague ideas (Vibes) into coherent research logic. The system compiles these ideas using a research paper ontology of sixteen academic parameters. Compilation failures indicate missing logical components;
arXiv:2608.05438v1 Announce Type: new Abstract: Generative AI has infiltrated every stage of the research lifecycle: how scholarship is conducted, written, published, and reviewed. Recent policy responses, such as ACM's authorship policy, address an immediate concern about responsible and transparent disclosure of AI use. We argue that a focus on authorship and disclosure, although necessary, risks obscuring and ballooning a set of entrenched problems and strains within publication systems. The central question is not about how papers and other research artifacts should incorporate AI, but how scientific communication itself should evolve when all relevant parties (authors, reviewers, readers) may rely on AI assistance. We draw on our experience within these and other roles to illustrate two contrasting but feasible visions of 2036 with four entwined questions, namely about the purpose of papers as artifacts, reviews, human reviewers, and the incentives that bind all of them. We argue
arXiv:2608.05427v1 Announce Type: new Abstract: School network reorganization is a strategic planning problem that requires balancing demographic trends, territorial accessibility, educational requirements, and institutional constraints while ensuring an efficient allocation of public resources. This paper proposes an optimization framework for school dimensioning decisions based on a novel Integer Linear Programming formulation integrating geographical, administrative, and educational criteria. A synthetic benchmark generator is introduced to evaluate the scalability and computational performance of the model on artificial instances, while a real-world case study involving the complete public school network of the Calabria region (Italy) is conducted using actual institutional, territorial, and demographic data. The proposed approach effectively identifies optimal aggregation plans under different policy scenarios while preserving the structural characteristics of the educational syst
arXiv:2608.05181v1 Announce Type: new Abstract: Building Information Modeling (BIM) has transformed workflows across the Architecture, Engineering, and Construction (AEC) industry, yet its relationship with employee job satisfaction remains insufficiently understood. This pilot study investigates whether BIM engagement or demographic characteristics better predict job satisfaction among AEC professionals. Survey responses from 104 participants were analyzed using Spearman rank correlations, logistic regression, and Classification and Regression Tree (CART) modeling. 27 items Job Satisfaction Index demonstrated excellent internal reliability. Across all analytical approaches, BIM engagement emerged as a stronger predictor of job satisfaction than demographic factors. Specifically, the proportion of project work completed using BIM was the only significant predictor of job satisfaction, whereas age, gender, education level, and professional experience showed no significant relationships.
arXiv:2608.05180v1 Announce Type: new Abstract: The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-appropriate preferences in high-stakes contexts. We introduce the Nuclear Decision-Making Benchmark (NDM Bench), a targeted evaluation framework of 151 scenarios authored by PhD-credentialed scholars in international relations spanning four domains: escalation (76), arms control (25), non-proliferation (25), and proliferation (25). Scenarios are actor-agnostic, enabling multiple country pairs to be exchanged, and we introduce experimental phrasing variants to probe sensitivity to narrative framing. We apply the benchmark to seven frontier AI systems: DeepSeek-V3.2, ERNIE 4.5-300B, Gemini 3 Pro, GLM-4.6, GPT-5.2, Llama 4 Maverick-17B Instruct, and Qwen3-235B. We find significant overall inter-model variation in all four domains, with 91.7% of pairwise inter-model
arXiv:2608.05179v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review. End-to-end AI scientist systems can now produce paper-like manuscripts, but their claims are often harder to verify than their code is to run. This survey studies that gap in computational AI/ML research, where code, benchmarks, experiments, and write-ups are most visible. We screen 125 candidate works and include 35, with full-text coding of 26 entries: 24 runnable systems and two study or position papers. We code seven audit dimensions: lifecycle stage, autonomy level, evaluation method, released artifacts, human-in-the-loop points, novelty verification, and result-selection disclosure. The main pattern is that code release is now common, but reproducibility-grade and claim-verification artifacts remain much less common. In the 24 runnab
arXiv:2608.05178v1 Announce Type: new Abstract: Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, datasets, or professional materials such as curriculum vitae (CV) must be shared selectively. We introduce a controlled simulation framework in which large language model (LLM)-based professors must grant access to only one requestor. Across prompts, requesters vary systematically by global region (Global North vs. Global South) and academic seniority (undergraduate student, PhD candidate, postdoctoral researcher, and tenured professor), while all other factors remain constant. Across varying evaluation scenarios, LLMs exhibit contrasting academic status biases, with some prioritizing PhD candidates, while others favor tenured professors. However, when global regions differ, a distinct divergence emerges based on model architecture: while many frontier LLMs systematically favor requesters from t
arXiv:2608.05177v1 Announce Type: new Abstract: Research indicates students require customized written composition feedback to enhance critical thinking and writing competence, yet teachers face barriers to delivering timely personalized guidance due to heavy workloads. To address this gap, this study developed the Writing Improvement and Smart Evaluation Agent (WISE Agent), an artificial intelligence (AI) feedback tool targeting textual logic and perspective biases in student essays. We conducted a three-month intervention with 260 Chinese sixth-grade students, each completing seven themed essays and receiving targeted WISE Agent feedback shortly after submission. Assessment used a critical thinking rubric adapted from the California Critical Thinking Disposition Inventory (CCTDI), covering seven core dimensions including cognitive maturity and open-mindedness. Results indicate structural optimizations in critical thinking dimensions rather than a uniform increase in total scores. Whi
arXiv:2608.05176v1 Announce Type: new Abstract: Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what they teach, how they teach it, and why. We now stand at what may be the most consequential of such turning points. Three deeply intertwined transformations have been converging simultaneously: 1. The very nature of music has changed: how it is made, distributed, consumed, and valued; 2. The public for music has changed: listening habits are now shaped by streaming algorithms and the boundary between consumer and creator has blurred; 3. Music-making itself has changed: digital audio workstations (DAWs) have for two decades been reshaping compositional practice. In addition, generative AI has now irrupted, capable of producing complete, stylistically coherent musical pieces from a short text prompt. These changes are not independent of one another, and they all bear directly on musical education - both
arXiv:2608.05175v1 Announce Type: new Abstract: Large language models can complete most of the assignments in an introductory artificial intelligence course. This paper is an experience report on redesigning one such course, CSS~382 at the University of Washington Bothell, in response. Rather than freeze the curriculum, the redesign retained the course's classical core (search, adversarial search, Markov decision processes, reinforcement learning) and added a strand in which students build a large language model from scratch, so that a tool they are required to use is also one they are required to understand. Assessment was rebuilt around tasks that resist unattributed automation: in-class exercises, reflective writing, and a defended team project, with examinations removed entirely. The policy on AI was inverted, from unmentioned in 2023 to required in 2026. The center of the paper is a participatory ethics sequence in which a cohort of students deliberated on and endorsed a "Student
arXiv:2608.05174v1 Announce Type: new Abstract: Artist-led open-source libraries such as Processing or openFrameworks have had a major impact on artists and designers who use code as a creative medium. In this work, we conduct the first large-scale empirical study of public code repositories that use these libraries. Our study dives into 1,613,571 code repositories collected from the Software Heritage archive. Combining quantitative and qualitative methods, we investigate the diversity of practices and practitioners in terms of code hosting, geographical distribution, characteristics of code-based creative works and the purposes of these works. Key findings include evidence of the worldwide presence of generative art and creative coding, as well as the adoption of these practices in both education and across creative industries. We illustrate these findings with concrete examples of repositories and contributor profiles around the globe, spanning the spectrum from university curricula
arXiv:2608.05173v1 Announce Type: new Abstract: As AI capabilities advance, AI systems will pose greater risks to national security and potentially humanity as a whole. Governments may eventually conclude that these risks warrant restraining AI development. This motivates the question: will governments still be able to restrain AI development in the future, should they want to do so? In this paper we analyze which world events and changes to the state of AI development would make future governance more difficult or even effectively impossible. Our analysis surfaces likely pathways that would lead to these difficulties, including hardware proliferation, continued algorithmic progress, and the release of catastrophically dangerous AI models. Due to the field's lack of understanding of AI development, it may be difficult or impossible to know when we will hit a "point of no return", and we therefore recommend a conservative approach. The window may be closing, but governments currently ha
arXiv:2608.05172v1 Announce Type: new Abstract: The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affects work: a new technology changes the cost or time each task requires and these task-level effects aggregate to occupation-level effects. We study how tasks should be weighted in this aggregation. Prior work has relied on idiosyncratic or ill-justified choices for task weights. While recent work suggests weighting tasks by time spent, existing time shares are either based on coarse ONET data not intended for this purpose or estimated via black-box language models. We address this gap by proposing a principled method for estimating time shares for nearly 18,000 tasks that constitute nearly all U.S. jobs. Our estimates factor a task's time into (i) the expected frequency of the task, derived from ONET, and (ii) the time to complete a single instance of it. To estimate the latter, we solve a constraint s
arXiv:2608.05171v1 Announce Type: new Abstract: Generative AI (GAI) creates new opportunities for collaborative problem-solving (CPS), yet its role in shaping student interaction remains unclear. To address this gap, we conducted a six-week quasi-experimental study with 201 fifth-grade students in two conditions: with and without GAI. Chi-square analysis showed significant differences in CPS behavior distributions between groups. Compared with the control group, the GAI-supported group demonstrated more social behaviors, particularly engagement and conflict management, but less frequent cognitive behaviors such as task planning and solution reasoning. Lag sequential analysis further revealed distinct interaction patterns: while the control group followed a more conventional transition from listening to planning, the GAI group showed a robust pathway from task planning to conflict management to solution reasoning. Thematic analysis of AI interaction logs suggested that students used GAI
“The premise that cybersecurity is a back-office or administrative expense and that something might not happen — that needs to be changed,” says Fadi Fadhil, field CIO and director of field strategy at Palo Alto Networks. “CISOs and CIOs can steer that change by engaging in simplified conversations with university leadership. It’s a strategic effort, helping them understand how the investment reduces institutional risk.” When it comes to budgeting for their cybersecurity programs, higher education CISOs must overcome some unique hurdles, ranging from the federated nature of university IT…
Raising your hand in class and patiently waiting until you’re called before speaking. Sharing with classmates in a group project. Understanding what you’re feeling and how best to express it safely. These are a few examples of what social-emotional skills look like in the classroom. Social-emotional learning (SEL) houses a variety of skills, all of which have always been embedded in the K–12 experience. As recent research points more directly to the value of weaving these learning moments into the K–12 curriculum, educational technology has risen to meet the demands. The Evidence for Social-…
Article URL: https://twitter.com/gergelyorosz/status/2062861559009820976 Comments URL: https://news.ycombinator.com/item?id=48411421 Points: 10 # Comments: 1
On any given Tuesday afternoon, a dean at Morgan State University can pull live enrollment trend data without submitting a ticket, waiting for a report or following up with the IT department. At most higher education institutions, that same request can take about three weeks. The difference isn’t the data platform, however. It’s how the historically Black college is prioritizing data literacy. Timothy Summers, vice president of IT and CIO at the Baltimore-based institution, is betting the university’s artificial intelligence strategy on employees’ ability to effectively interpret, question…
The landscape for specialized colleges and universities such as art schools is shifting as higher education continues to evolve to fit emerging job markets and student interest. Founded in 1882, Cleveland Institute of Art continuously challenges itself to stay modern and relevant. Years ago, the school’s leadership had the vision to partner with the city to revitalize an area due for reinvigoration. The result was the Interactive Media Lab, which brings together the university, the city and private industry into a satellite campus that gives students and the community a space to…
Innovative Leader Award - Lauren Harwood of Dighton-Rehoboth Regional School District shares how she focuses her efforts on AI, CTE program, and cybersecurity
The recent ransomware incident involving Canvas has renewed attention on one of the most difficult decisions schools and technology providers can face: how to respond when sensitive student, faculty, or institutional data is stolen and threatened with public release. The post The Canvas ransomware attack shows why schools must focus on containment, not just recovery appeared first on eCampus News .
The House VA Committee voted 19-0 to subpoena Oracle Health executives Larry Ellison and Mike Sicilia after learning the VA’s EHR contract ballooned from $10 billion to $27 billion — despite Oracle’s promise to Congress that it would absorb any costs beyond the original cap. The post Oracle Health Execs Subpoenaed After VA Contract Costs Nearly Triple appeared first on MedCity News .
Ionis Pharmaceuticals’ zilganersen, brand name Zanvastro, is the first disease-modifying therapy approved for the neurological disorder Alexander disease. While Ionis has experience developing medicines for rare neuroscience indications, Zanvastro will be the first neurology product the company brings to the market without a commercialization partner. The post FDA Approves Ionis Pharma Drug, the First for Ultra-Rare Alexander Disease appeared first on MedCity News .
Before the beginning of each school year, teachers spend time getting their classrooms ready for students. Everything has its place, from whiteboards to pens, pencils and pushpins. With the adoption of cloud computing by school districts, the IT departments of K–12 schools also must figure out what goes where, but on a much larger level. Many districts operate in hybrid cloud environments, where some data and applications reside on servers that are on-premises and others reside with public cloud providers. It’s not a decision that should be taken lightly. Click the banner below to see how a…
The reality of K–12 IT teams is that most are very small — and, as a result, stretched thin. More than half of K–12 districts say they’re understaffed for everyday classroom technical support needs. On a given day, technicians may handle dozens of support requests for password resets, basic device troubleshooting, network connectivity issues, broken devices and more. When a district’s technical needs outweigh its team’s capacity to support it, something has to give. The solution? Agentic artificial intelligence. Agentic AI isn’t here to replace IT staff. Instead, it gives small departments a…
Three things ski patrol can teach about building a business The post Bringing Ski Patrol Lessons to Healthtech Entrepreneurship appeared first on MedCity News .
With the rapid spread of Gen AI, we’re entering a time when human connection is becoming more valuable and more vulnerable. The post Can high-tech scale high-touch? appeared first on eCampus News .
arXiv:2608.19491v2 Announce Type: replace-cross Abstract: Most modern optimizers form their momentum as an exponential moving average (EMA) of past gradients, forgetting every direction at one fixed rate. However, the inputs a deep network sees during training can be highly anisotropic, with a few directions queried frequently while most are seen rarely. Preconditioning methods address this anisotropy by wrapping extra processing around this buffer and leave the momentum update itself unchanged. We propose Activation-Keyed Momentum (AK-Momentum), which builds direction-awareness into the momentum update rule. The gradient of a linear layer splits into an input activation that acts as a key and an output-side error that acts as a value. Keying on that activation, AK-Momentum updates the momentum buffer by the canonical delta rule, so each direction is forgotten at a rate set by how often it appears. We prove that it is a valid momentum, that it applies the input-side curvature correctio
arXiv:2606.21657v2 Announce Type: replace-cross Abstract: Do people perceive the same facial expression in the same way? Should we expect vision models to be flexible in how they perceive facial expressions? Facial expressions are nonverbal social signals used in human interaction, but facial expression recognition datasets often focus on a single deterministic annotation per sample. We introduce Chehre, an emoji-prompted video dataset with a wide range of dynamic facial expressions for exploring perceptual variation. In Chehre, 203 participants were prompted to express and record 40 facial emojis. Later, their facial motions were transferred onto synthetic faces to preserve privacy. A separate group annotated the videos, resulting in 2,111 videos annotated by 1,242 perceivers, with ~30 annotators per video. Chehre enables us to define a new task: "distributional expression recognition", which tests whether a model can reproduce the variation observed across annotator responses. We tes
arXiv:2606.19264v2 Announce Type: replace-cross Abstract: The knowledge encoded in large language models (LLMs) can serve as a substrate for structured reasoning over variables describing a complex world, but accessing this knowledge in a probabilistically coherent manner poses a difficult inference problem. We propose Large Language Gibbs, a scheme for structured probabilistic inference that uses conditional distributions of an LLM as transition operators. Rather than sampling structured objects through single-pass autoregressive generation, we iteratively resample individual variables conditioned on others using an LLM's next-token conditionals. This approach avoids order-dependent biases and produces a stationary distribution that reflects a compromise between all local conditionals. We apply this approach to sampling from synthetic distributions, consistent reasoning tasks, and Bayesian structure learning. The results suggest that the use of LLM conditionals in MCMC is a practical
arXiv:2606.18388v2 Announce Type: replace-cross Abstract: RL post-training strategies are dataset-dependent and reveal a recurring empirical pattern: capacity parameters accumulate monotonically across stages, while regularization parameters predominantly oscillate in response to shifting training dynamics. This distinction highlights a potential flaw in fixed training schedules: by forcing all parameters along rigid paths, they fail to capture the dynamic exploration-exploitation tradeoffs that regularization must track. We uncover this through LLMZero, an agentic system that optimizes training trajectories via tree search by diagnosing pathologies at each checkpoint and proposing coordinated multi-parameter transitions. Across four diverse GRPO tasks, LLMZero discovers strategies that improve over the base model by 9% to 140% and over grid search by 6% to 15% (relative), consistently outperforming random search and a skill-based agent under a matched compute budget. The capacity--reg
arXiv:2606.07451v2 Announce Type: replace-cross Abstract: Vision-language models such as CLIP are highly useful for diverse tasks due to their shared image-text embedding space. Despite this, the image and text embeddings are often poorly aligned, affecting downstream performance. Recent work has hypothesized that this can be attributed to an information imbalance: images contain more information than their captions describe. In this work, we propose TEVI, a framework that uses captions as a signal for what to retain from image embeddings. Specifically, we use sparse autoencoders to disentangle image embeddings and train a masking module to selectively reconstruct the embedding based on a given caption. In a controlled setup with synthetic captions, we show that TEVI is effective at preserving caption-described attributes while discarding others. We find that this extends to CLIP models trained on natural images, where TEVI learns to mask meaningfully and allows retrieval based on cond
arXiv:2606.02914v3 Announce Type: replace-cross Abstract: Background: Oral diseases affect nearly 3.5 billion people worldwide, yet the comparative clinical potential of large-scale AI models in dentistry remains poorly understood. Three distinct model categories have emerged: language-generative models, discriminative vision foundation models, and dental-specific foundation models, with no unified review examining their relationships and collective limitations. Methods: Following PRISMA-ScR guidelines, we systematically searched four databases (PubMed, Google Scholar, Scopus, arXiv), screened independently by two reviewers. After applying inclusion/exclusion criteria, 97 studies (2020-2026) were included. We propose a two-dimensional classification framework organizing models by architectural paradigm and dental specialization degree. Results: Language-generative models excel at text-based tasks (clinical reasoning, licensing exams, patient communication) but show inconsistent perform
arXiv:2606.02372v2 Announce Type: replace-cross Abstract: Equipping language agents with world models enables them to anticipate environment dynamics and evaluate candidate actions before execution. However, existing textual world models are typically fixed after training, preventing them from adapting to the on-policy state-action distributions induced by an evolving agent. Meanwhile, agent-improvement methods often rely on external rewards or verifiers, limiting their applicability in realistic interactive environments. In this paper, we propose COMAP, a novel framework that co-evolves textual world models and agent policies through closed-loop interaction. At each decision step, the world model predicts future state feedback for candidate actions, and the agent performs future-aware reflection by estimating the reliability of this feedback and refining its action accordingly. The resulting on-policy trajectories are then used to update the world model via self-distillation, allowing
arXiv:2605.30912v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) improves vision-language models (VLMs) by optimizing outcome rewards derived from final answers. However, such outcome-only rewards do not tell the model which image regions justify an answer. For questions that require visual grounding, these rewards cannot distinguish responses supported by relevant visual evidence from those produced by language-prior shortcuts or lucky guesses. We introduce EASE (Evidence-Anchored Spatial Attention), which augments multimodal RLVR with visual-evidence process supervision. EASE converts annotated evidence regions into a smoothed visual-token target and uses it to guide response-to-image attention during RL training, but only on high-reward trajectories. The annotations are used solely as privileged training labels, while inference requires only the original image and question. Across Qwen2.5-VL-7B, Qwen3-VL-4B, and Qwen3-VL-8B, EASE raises
arXiv:2602.07106v3 Announce Type: replace-cross Abstract: Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet extending them to jointly produce speech and 3D facial animation remains largely underexplored. A key challenge is the mismatch between the discrete semantic reasoning of LLMs and the dense temporal dynamics required for 3D facial motion. We propose Expressive Omni (Ex-Omni), a framework that augments OLLMs with speech-accompanied 3D facial animation. Ex-Omni decouples semantic reasoning from temporal generation through a speech-unit generator with blendshape co-supervision and a non-autoregressive blendshape decoder, where speech units provide temporal scaffolding and hidden speech representations carry facially relevant cues. We further introduce a token-as-query gated fusion (TQGF) interface for controlled semantic injection, as well as InstructS2SF-1200K, a 1.2M-sample weakly supervised dataset for speech-accompanied facial ani
arXiv:2602.06065v4 Announce Type: replace-cross Abstract: Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning. Studies of the internal representations of Large Language Models (LLMs) support their ability to parse text when predicting the next word, while representing semantic notions independently of surface form. Yet, which data statistics make these feats possible, and how much data is required, remain largely unknown. Probabilistic context-free grammars (PCFGs) provide a tractable testbed for studying these questions. However, prior work has focused either on the post-hoc characterization of the parsing-like algorithms used by trained networks; or on the learnability of PCFGs with fixed syntax, where parsing is unnecessary. Here, we (i) introduce a tunable class of PCFGs in which both the degree of ambiguity and the correlation structure across scales can be controlled; (ii) provide a l
arXiv:2503.03313v4 Announce Type: replace-cross Abstract: Text-Attributed Graphs (TAGs), where each node is associated with text descriptions, are ubiquitous in real-world scenarios. They typically exhibit distinctive structure and domain-specific knowledge, motivating the development of a Graph Foundation Model (GFM) that generalizes across diverse graphs and tasks. Despite large efforts to integrate Large Language Models (LLMs) and Graph Neural Networks (GNNs) for TAGs, existing approaches suffer from decoupled architectures with two-stage alignment, limiting their synergistic potential. Even worse, existing methods assign out-of-vocabulary (OOV) tokens to graph nodes, leading to graph-specific semantics, token explosion, and incompatibility with task-oriented prompt templates, which hinders cross-graph and cross-task transferability. To address these challenges, we propose PromptGFM, a versatile GFM for TAGs grounded in graph vocabulary learning. PromptGFM comprises two key componen
arXiv:2405.21047v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) struggle with reliably generating highly structured outputs, such as program code, mathematical formulas, or well-formed markup. Constrained decoding approaches mitigate this problem by greedily restricting what tokens an LLM can output at each step to guarantee that the output matches a given constraint. Specifically, in grammar-constrained decoding (GCD), the LLM's output must follow a given grammar. In this paper, we demonstrate that GCD techniques (and in general constrained decoding techniques) can distort the LLM's distribution, leading to outputs that are grammatical but appear with likelihoods that are not proportional to the ones given by the LLM, and so ultimately are low-quality. We call the problem of aligning sampling with a grammar constraint, grammar-aligned decoding (GAD), and propose adaptive sampling with approximate expected futures (ASAp), a decoding algorithm that guarantees the
arXiv:2606.20954v2 Announce Type: replace Abstract: Long-running language-model systems accumulate interaction history that outgrows the context window, so they must continually evict. When an eviction policy drops a task-critical detail, for example an access token issued at login or a path the next call needs, the action fails. We present LRE (Learned Relevance Eviction), a kilobyte-scale, CPU-only, language-model-free scorer that learns which units of history are task-critical and keeps them by verbatim extraction. Under a matched-budget comparison, in our experiment, no baseline dominates LRE on the accuracy-cost plane. On agents, LRE recovers 93% of the aggregate accuracy of keeping the entire history (41.1 vs. 44.0) and exceeds it by 27% on the simplest tasks, while requiring zero compressor calls and cutting the worst-case peak prompt by 52%. A controlled study trace shows LRE completes tasks where the others loop, finishing one such task in 37% fewer calls than keeping everythi
arXiv:2606.16583v2 Announce Type: replace Abstract: Safe deployment of clinical vision-language models (VLMs) requires reliable uncertainty estimation (UE): a signal indicating when predictions should be trusted or escalated to a clinician. We test whether current UE methods actually deliver this signal. Benchmarking 8 methods across 12 VLMs on clinical visual question-answering (VQA), we find that UE quality is not an intrinsic property of the UE method: it tracks model accuracy, degrading precisely where the model performance is weakest, and therefore where reliability is most needed. When we stress-test models by hiding the correct option among the multiple-choice answers (NOTA perturbations), accuracy collapses while uncertainty barely changes, leaving models systematically miscalibrated. Yet, we find that uncertainty on the unperturbed input reliably anticipates which predictions will collapse under NOTA, indicating that UE in current VLMs carries diagnostic information about mode
arXiv:2606.10581v2 Announce Type: replace Abstract: Speech carries more information than just words: a child's voice, a fearful tone, or a noisy background should all lead a sufficiently competent spoken-dialogue assistant to different replies. Current Speech Language Models (SLMs) can recognize such paralinguistic cues but often ignore them in open-ended dialogue. We observe that a simple paralinguistic instruction scaffold at the inference stage narrows this perception-behavior gap, suggesting that the relevant cues are already latent in the model. Such scaffolds, however, remain brittle under multi-turn context and competing instructions. Therefore, we propose \textbf{ParaBridge}, an on-policy self-distillation method that turns a brittle inference-time scaffold into stable model behavior. During training, the scaffold serves only as a temporary privileged view; the scaffold-free model rolls out its own response, while the scaffolded view supplies dense, full-vocabulary next-token t
arXiv:2606.09655v2 Announce Type: replace Abstract: Despite remarkable progress in machine translation (MT), non-AI communities have raised growing concerns about MT systems, suggesting a noticeable gap between technical advancement and the needs of real-world users. For instance, while NLP researchers focus on benchmark performance, end users care about ethical concerns, trust, reliability, costs, and more. We argue that listening to various user communities is essential so that research efforts would be directed towards the problems that the communities care about. To this end, we present a large-scale analysis, for the first time, that investigates what four stakeholder communities (AI developers, professional translators, language learners, and language service providers) post about MT technology on social media. To do so, we construct a dataset of 79,286 posts and comments from Reddit, Facebook, Bluesky, and Mastodon from 2019 to 2025, and analyse where these communities disagree,
arXiv:2606.07313v2 Announce Type: replace Abstract: Detecting AI-generated text is especially difficult under distribution shift, such as transfer across domains, source models, and editing attacks. We propose an AI-generated text detector based on steering vectors extracted from the hidden representations of a frozen language model. At each layer, we construct a direction that separates human-written from AI-generated text, and represent each input by its layer-wise alignment with these directions. A lightweight classifier trained on these projection features yields the final detection score. Our method achieves strong performance both in-distribution and under distribution shift, including across domains, source models, and machine-editing transformations such as polishing and rewriting. Interpretation analyses show that the learned directions align with recognizable stylistic cues while capturing substantial additional signal beyond surface features. These results position AI-genera
arXiv:2606.06350v2 Announce Type: replace Abstract: Reliable rubric grading requires more than accurate score prediction. Each judgement must be grounded in the mark scheme and evidence from the student answer. Existing credit-assignment and intervention methods, primarily designed for self-contained reasoning tasks such as mathematics reasoning, struggle in this setting because they do not identify where grading reasoning goes wrong or how the model's belief about the final mark changes during reasoning. We propose Evidence-Diagnosed Intervention Training (EDIT), a two-phase framework for training more rubric-faithful LLM graders. First, EDIT-SFT locates problematic reasoning steps using internal model signals: posterior belief over the final mark and input-grounding scores. It then revises only these local steps with help from a rubric checklist. Second, EDIT-RL calibrates the grader with belief-guided reward shaping, penalising large harmful belief drifts while still allowing helpfu
arXiv:2606.05553v2 Announce Type: replace Abstract: Role-playing language agents (RPLAs) simulate specific characters and personas across applications such as entertainment, companionship, interactive storytelling, and education. Faithful role-play requires more than producing plausible, in-character responses: as a character's values and behavior change over a narrative, an RPLA should reflect the character's state at the relevant stage. However, existing benchmarks largely treat characters as fixed personas or test only what they know at a given point in the narrative. We introduce ArcANE (Arc-Aware Narrative Evaluation), a benchmark for evaluating whether an RPLA follows a character's development across a narrative. ArcANE first builds an Arc that maps how a character's values, motivations, or relationships change over the story. The benchmark then scores how well an RPLA's responses fit the corresponding stages of the Arc, covering three distinct scenario types: scenes from the nov
arXiv:2606.03780v2 Announce Type: replace Abstract: Activation patching can identify a mixture-of-experts (MoE) block whose clean output restores a corrupted factual prediction. However, because the block output combines contributions from multiple routed experts, block-level rescue does not establish whether the recovery localizes to an individual expert or depends on the routed expert set. We study this question on single-token COUNTERFACT contrasts by corrupting subject-token embeddings, restoring clean block outputs, and then restoring clean-minus-noised expert updates under fixed routing. In Qwen3-30B-A3B-Base, a discovery sweep selects layer 44, and held-out analysis identifies L44E069 as a recurrent routed contributor with positive specificity over same-layer active experts. Its effect is fact-matched and improves true-token probability and rank, which explains part of the layer rescue. In Mixtral-8x7B-v0.1, the selected recurrent singleton is not specific; matched-size controls