EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Fri, 26 Jun 2026 14:20:22 -0400
EdTech Mag (K-12)

Spring 2026

Source ↗
technology Fri, 26 Jun 2026 13:08:23 -0400
EdTech Mag (Higher)

What SUNY’s Systemwide AI Policy Means for Public University IT Leaders

Leaders at the State University of New York’s 64 campuses have until the end of the year to establish or update artificial intelligence guidelines, including standards for bias evaluation, student data privacy and responsible AI use. The mandate comes from a binding AI governance policy passed in May, leaving higher ed IT leaders at SUNY campuses to devise ways to evaluate AI vendors, implement governance workflows, protect institutional data and support responsible AI adoption at scale. The framework is already having an impact beyond the Empire State, with CIOs and IT leaders across…

Source ↗
behavior Fri, 26 Jun 2026 12:55:56 +0000
District Admin

Houston ISD board unanimously approves controversial Bible-infused curriculum

Adopting the state-approved Bluebonnet Learning curriculum for elementary students will come with a funding boost of more than $3 million for the district, which also approved a $2 billion operating budget on Thursday. The post Houston ISD board unanimously approves controversial Bible-infused curriculum appeared first on District Administration .

Source ↗
behavior Fri, 26 Jun 2026 12:52:54 +0000
District Admin

Fayette County superintendent files ‘whistleblower’ complaint after being put on leave

Fayette County Public Schools has been mired in controversy over its financial situation in recent years, with district officials recently saying FCPS has misstated its finances since at least 2008 and has much less money than previously thought. The post Fayette County superintendent files ‘whistleblower’ complaint after being put on leave appeared first on District Administration .

Source ↗
regulation Fri, 26 Jun 2026 12:30:00 +0000
The 74

In ‘Toy Story 5,’ Tech ‘Invades’ Playtime. It Also Threatens Human Connection.

In one of the opening scenes of “Toy Story 5,” Jessie — a cowgirl doll — tries to find out why the twins who live across the street never want to play with her owner Bonnie. What she finds, when she peers through the window of the neighbor’s home, is the two young children on […]

Source ↗
regulation Fri, 26 Jun 2026 10:30:00 +0000
The 74

Texas Quietly Began Work on Divisive History Curriculum a Year Ago

The Texas State Board of Education will vote Friday on a set of new social studies standards that have drawn fire and fervor for espousing pro-American views and Christian values. If approved, the vote would typically mark the beginning of a long, and probably divisive, process to design curriculum based on the standards. But The […]

Source ↗
regulation Fri, 26 Jun 2026 10:21:48 -0400
K-12 Dive

Kindergarten reading and math skills can predict 3rd grade success, NWEA finds

Proficiency by grade 3 is linked to long-term academic and life outcomes, making early identification of struggling students key.

Source ↗
behavior Fri, 26 Jun 2026 10:00:00 +0000
HealthLeaders

Phenomenal Care at Night: Sentara Health's CNO on Pairing Nurses with Virtual Partners

Sentara Health's virtual nursing partner program is helping give time back to night shift nurses, says this CNO. HealthLeaders spoke to Amber Price , senior vice president and enterprise CNO at Sentara Health , about the challenges nurses face on the night shift and how bringing in virtual nursing partners can help. Tune in to hear her insights. Click here to read the accompanying article. Pillar: CNO Image: Tags: innovation nurses nursing staff technology telemedicine training Secondary Pillars: CNO Article Type: Analysis Published Date: Monday, June 15, 2026 Hide sidebars: Render small main image:

Source ↗
behavior Fri, 26 Jun 2026 10:00:00 +0000
eSchool News

Teacher burnout is at an all-time high

Teacher stress declined modestly in 2026, but teachers were still far more likely than similar working adults to report higher stress, worse well-being and greater financial strain, extending a pattern that has persisted since 2021, according to new RAND research.

Source ↗
technology Fri, 26 Jun 2026 09:00:00 +0000
Tech & Learning

Best Sites for Blended Learning

Blended learning websites help teachers combine traditional instruction with online learning.

Source ↗
technology Fri, 26 Jun 2026 09:00:00 +0000
eCampus News

The real work of AI and instructional technology is creative

In my Systems Analysis and Design course, students are not handed the requirements for building a software application. They have to uncover them by asking the right questions within an AI-based learning activity. The post The real work of AI and instructional technology is creative appeared first on eCampus News .

Source ↗
technology Fri, 26 Jun 2026 08:10:29 -0400
EdTech Mag (K-12)

ISTE

Source ↗
audience Fri, 26 Jun 2026 07:00:00 +0000
Inside Higher Ed

The Big Impact of Small Telescopes

The Big Impact of Small Telescopes Elizabeth Redden Fri, 06/26/2026 - 03:00 AM In an era of big data and big telescopes, college observatories remain essential. Byline(s) Alex Gianninas

Source ↗
audience Fri, 26 Jun 2026 07:00:00 +0000
Inside Higher Ed

Female Academics ‘Increasingly Delay Motherhood’ Until Age 35

Female Academics ‘Increasingly Delay Motherhood’ Until Age 35 Susan H. Greenberg Fri, 06/26/2026 - 03:00 AM “Pronounced penalties” for those early-career staff with children may explain why Ph.D.s postpone becoming parents, a recent study finds. Byline(s) Jack Grove for Times Higher Education

Source ↗
audience Fri, 26 Jun 2026 07:00:00 +0000
Inside Higher Ed

Readers Respond on ‘Noncredit’

Readers Respond on ‘Noncredit’ Sara Brady Fri, 06/26/2026 - 03:00 AM Unofficial names and alternate terms. Byline(s) Matt Reed

Source ↗
audience Fri, 26 Jun 2026 07:00:00 +0000
Inside Higher Ed

Texas Law Dean Pushes Socratic Teaching Amid Rise of AI

Texas Law Dean Pushes Socratic Teaching Amid Rise of AI gianna.jakubowski Fri, 06/26/2026 - 03:00 AM Byline(s) Gianna Jakubowski

Source ↗
audience Fri, 26 Jun 2026 07:00:00 +0000
Inside Higher Ed

Florida Universities Consider Banning Undocumented Students

Florida Universities Consider Banning Undocumented Students Sara Weissman Fri, 06/26/2026 - 03:00 AM And the board overseeing state colleges is eyeing a similar ban. Together, the policies could make Florida the fourth state to limit noncitizens’ enrollment in public colleges and universities. Byline(s) Sara Weissman

Source ↗
audience Fri, 26 Jun 2026 07:00:00 +0000
Inside Higher Ed

Dear Colleague Letter Asks Colleges to End Affinity Housing

Dear Colleague Letter Asks Colleges to End Affinity Housing Emma Whitford Fri, 06/26/2026 - 03:00 AM The Trump administration alleges that housing that caters to minority students violates the Fair Housing Act. Experts say the guidance is unlikely to withstand legal challenges. Byline(s) Emma Whitford

Source ↗
audience Fri, 26 Jun 2026 07:00:00 +0000
Inside Higher Ed

New HBCU Partnership Speeds Path to Law School

New HBCU Partnership Speeds Path to Law School Joshua.Bay Fri, 06/26/2026 - 03:00 AM Grambling State has teamed up with Southern University Law Center to allow students to earn a bachelor’s and a law degree in six years, lowering costs and strengthening Louisiana’s attorney pipeline. Byline(s) Joshua Bay

Source ↗
audience Fri, 26 Jun 2026 07:00:00 +0000
Inside Higher Ed

California Adjuncts Sue for ‘Uncompensated Work’

California Adjuncts Sue for ‘Uncompensated Work’ kathryn.palmer… Fri, 06/26/2026 - 03:00 AM Byline(s) Kathryn Palmer

Source ↗
audience Fri, 26 Jun 2026 07:00:00 +0000
Inside Higher Ed

Corequisite Math Might Be Less Effective Than Previously Thought

Corequisite Math Might Be Less Effective Than Previously Thought Johanna Alonso Fri, 06/26/2026 - 03:00 AM Byline(s) Johanna Alonso

Source ↗
audience Fri, 26 Jun 2026 05:00:00 -0400
Higher Ed Dive

Ohio bill would broaden power of university civics center directors

The politically created academic centers have drawn fierce criticism from faculty, who say they expand state intrusion into higher education.

Source ↗
regulation Fri, 26 Jun 2026 05:00:00 -0400
K-12 Dive

Test yourself on the past week’s K-12 news

From a large district’s consolidation plan to a report on states meeting special education requirements, what did you learn from our recent stories?

Source ↗
audience Fri, 26 Jun 2026 03:41:00 +0000
Inside Higher Ed

The Key Podcast: Historians and American Exceptionalism

The Key Podcast: Historians and American Exceptionalism sara.custer@in… Thu, 06/25/2026 - 11:41 PM Byline(s) IHE Staff

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

arXiv:2605.19576v2 Announce Type: replace-cross Abstract: Self-evolving skill libraries face a silent failure mode we term \emph{library drift}: unbounded skill accumulation without outcome-driven lifecycle management causes retrieval degradation, false-positive injections, and performance stagnation. Recent evaluation confirms the symptom (LLM-authored skills deliver +0.0pp gain while human-curated ones deliver +16.2pp (SkillsBench)), yet the underlying mechanism has not been isolated. We provide (1) a \textbf{reproducible trigger}: ablations that isolate drift: one disables skill injection (flat floor, +0.002), one imposes premature retirement (active harm, $-$0.019); (2) \textbf{trace-level diagnostics}: an append-only evidence log with per-skill contribution scores, attribution verdicts, and router engagement metrics that make the failure visible before it reaches end-task scores; and (3) a \textbf{verified fix}: a minimal governance recipe (outcome-driven retirement + bounded acti

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

arXiv:2604.26136v2 Announce Type: replace-cross Abstract: Preserving a speaker's voice identity while generating speech in a different language remains a fundamental challenge in spoken language technology, particularly in specialized domains such as scientific communication. In this paper, we address this challenge through our system submission to the International Conference on Spoken Language Translation (IWSLT 2026), the Cross-Lingual Voice Cloning shared task. First, we evaluate several state-of-the-art voice cloning models for cross-lingual speech generation of scientific texts in Arabic, Chinese, and French. Then, we build voice cloning systems based on the OmniVoice foundation model. We employ data augmentation via multi-model ensemble distillation from the ACL 60/60 corpus. We investigate the effect of using this synthetic data for fine-tuning, demonstrating improvements in intelligibility (WER & CER) and speaker similarity (SIM), with gains varying across languages.

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents

arXiv:2604.15877v2 Announce Type: replace-cross Abstract: As LLM agents scale to long-horizon, multi-session deployments, efficiently managing accumulated experience becomes a critical bottleneck. Agent memory systems and agent skill discovery both address this challenge, extracting reusable knowledge from interaction traces, yet a citation analysis of 1{,}136 references across 22 primary papers reveals a cross-community citation rate below 1\%. We propose the \emph{Experience Compression Spectrum}, a unifying framework that positions memory, skills, and rules as points along a single axis of increasing compression (5--20$\times$ for episodic memory, 50--500$\times$ for procedural skills, 1{,}000$\times$+ for declarative rules), directly reducing context consumption, retrieval latency, and compute overhead. Mapping 20+ systems onto this spectrum reveals that every system operates at a fixed, predetermined compression level: none supports adaptive cross-level compression, a gap we term

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Learning State-Tracking from Code Using Linear RNNs

arXiv:2602.14814v3 Announce Type: replace-cross Abstract: Over the last years, state-tracking tasks, particularly permutation composition, have become a testbed to understand the limits of sequence models architectures like Transformers and RNNs (linear and non-linear). However, these are often sequence-to-sequence tasks: learning to map actions (permutations) to states, which is incompatible with the next-token prediction setting commonly used to train language models. We address this gap by converting permutation composition into code via REPL traces that interleave state-reveals through prints and variable transformations. We show that linear RNNs capable of state-tracking excel also in this setting, while Transformers still fail. Motivated by this representation, we investigate why tracking states in code is generally difficult: actions are not always fully observable. We frame this as tracking the state of a probabilistic finite-state automaton with deterministic state reveals and

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Linguistics and Human Brain: A Perspective of Computational Neuroscience

arXiv:2602.08275v3 Announce Type: replace-cross Abstract: Elucidating the language-brain relationship requires bridging the methodological gap between the abstract theoretical frameworks of linguistics and the empirical neural data of neuroscience. Serving as an interdisciplinary cornerstone, computational neuroscience formalizes the hierarchical and dynamic structures of language into testable neural models through modeling, simulation, and data analysis. This enables a computational dialogue between linguistic hypotheses and neural mechanisms. Recent advances in deep learning, particularly large language models (LLMs), have powerfully advanced this pursuit. Their high-dimensional representational spaces provide a novel scale for exploring the neural basis of linguistic processing, while the "model-brain alignment" framework offers a methodology to evaluate the biological plausibility of language-related theories.

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

arXiv:2601.11061v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is highly effective for enhancing LLM reasoning, yet recent evidence shows models like Qwen 2.5 achieve significant gains even with spurious or incorrect rewards. We investigate this phenomenon and identify a "Perplexity Paradox": spurious RLVR triggers a divergence where answer-token perplexity drops while prompt-side coherence degrades, suggesting the model is bypassing reasoning in favor of memorization. Using Path Patching, Logit Lens, JSD analysis, and Neural Differential Equations, we uncover a hidden Anchor-Adapter circuit that facilitates this shortcut. We localize a Functional Anchor in the middle layers (L18-20) that triggers the retrieval of memorized solutions, followed by Structural Adapters in later layers (L21+) that transform representations to accommodate the shortcut signal. Finally, we demonstrate that scaling specific MLP keys within this circuit allows fo

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors

arXiv:2510.00586v3 Announce Type: replace-cross Abstract: Existing data poisoning attacks on retrieval-augmented generation (RAG) systems scale poorly because they require costly optimization of poisoned documents for each target phrase. We introduce Eyes-on-Me, a modular attack that decomposes an adversarial document into reusable **Attention Attractors** and **Focus Regions**. Attractors are optimized to direct attention to the Focus Region. Attackers can then insert semantic baits for the retriever or malicious instructions for the generator, adapting to new targets at near zero cost. This is achieved by steering a small subset of attention heads that we empirically identify as strongly correlated with attack success. Across 18 end-to-end RAG settings (3 datasets $\times$ 2 retrievers $\times$ 3 generators), Eyes-on-Me raises average attack success rates from 21.9 to 57.8 (+35.9 points, 2.6$\times$ over prior work). A single optimized attractor transfers to unseen black box retrieve

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

HauntAttack: When Attack Follows Reasoning as a Shadow

arXiv:2506.07031v5 Announce Type: replace-cross Abstract: Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities. However, the enhancement of reasoning abilities and the exposure of internal reasoning processes introduce new safety vulnerabilities. A critical question arises: when reasoning becomes intertwined with harmfulness, will LRMs become more vulnerable to jailbreaks in reasoning mode? To investigate this, we introduce HauntAttack, a novel and general-purpose black-box adversarial attack framework that systematically embeds harmful instructions into reasoning questions. Specifically, we modify key reasoning conditions in existing questions with harmful instructions, thereby constructing a reasoning pathway that guides the model step by step toward unsafe outputs. We evaluate HauntAttack on 11 LRMs and observe an average attack success rate of over 70\%, achieving up to 13 percentage points of absolute imp

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory

arXiv:2606.21649v2 Announce Type: replace Abstract: Existing embedding models are inherently static: they encode text segments in isolation, ignoring their surrounding context and temporal order. This paper introduces EvoEmbedding, a novel embedding model that generates evolvable representations for retrieval. It is tailored for long-context scenarios, where information is dynamic, sequential, and requires continuous state tracking. Our design is simple: EvoEmbedding maintains a continuously updated latent memory as it sequentially processes inputs, and uses it alongside the raw content to jointly generate evolvable embeddings. Consequently, for the same query, our model adapts its representation to retrieve distinct targets based on the evolving context, going beyond static semantic search. To equip the model with this capability, we construct EvoTrain-180K, a diverse dataset for the joint optimization of latent memory and retrieval. Furthermore, we introduce a memory queue to prevent

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

GRAG: Generic Response-Augmented Generation Framework for Personalized Conversational Systems

arXiv:2606.21097v2 Announce Type: replace Abstract: Deploying highly capable personalized conversational agents in resource-constrained or privacy-sensitive environments remains a significant challenge. We identify a fundamental bottleneck in the existing approaches: current training paradigms treat personalization and grounding as a single monolithic learning problem. Under these paradigms, language models are forced to simultaneously address what to say (content grounding) and how to say it in a user-specific way (personalization), which introduces significant computational and optimization challenges. Consequently, contextual grounding is often sacrificed for persona adherence, or vice versa, resulting in responses that are either weakly grounded in the conversational history or insufficiently personalized. In this work, we propose the Generic Response-Augmented Generation (GRAG) framework that decouples these competing objectives by leveraging offline, generic responses from high-c

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Prompt, Plan, Extract: Zero-Shot Agentic LLMs Workflows for Lung Pathology Extraction from Clinical Narratives

arXiv:2606.19852v2 Announce Type: replace Abstract: Information extraction from pathology reports is essential for cancer staging, tumor registry population. Yet key data remains embedded in narrative reports, making manual extraction labor-intensive and error-prone. Traditional supervised Natural Language Processing pipelines address this through fully supervised Named Entity Recognition and Relation Extraction, but require expensive manual annotation and suffer cascading failures when upstream entities are missed. In this study, we developed a zero-shot, agentic workflow, and evaluated five open-source generative Large Language Models (LLMs) to populate 13 College of American Pathologists synoptic fields from lung resection pathology reports. We compared them against a state-of-the-art supervised GatorTron NER-RE baseline using a novel, registry-aligned evaluation framework. The baseline achieved Micro-F1of 0.960, while the best zero-shot model (GPT-OSS-20B) achieved Micro-F1 of 0.89

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Learning User Simulators with Turing Rewards

arXiv:2606.19336v2 Announce Type: replace Abstract: Learning to simulate human users in interactive settings could advance the training of agent assistants, evaluation of personalization systems, research in the social sciences, and more. Existing approaches generally do so by training a large language model (LLM) to match a single ground truth response, either by maximizing the log probability or by using a similarity reward. We instead propose Turing-RL: a Turing-Test-based reinforcement learning approach for training user simulator models. Turing-RL uses a discriminative Turing reward with an LLM judge to score how indistinguishable a generated response is from the real user's given the user's history, and the user simulator LLM learns to produce responses indistinguishable from what the user could have said with such rewards. Across two different domains--conversational chat and Reddit forum discussion--we find that Turing-RL consistently outperforms baseline methods on both LLM an

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Analyzing and Encoding the Al-Mawrid Arabic-English Dictionary with the ISO Language Markup Framework and TEI Lex-0

arXiv:2606.18205v2 Announce Type: replace Abstract: This paper presents a robust methodology for the systematic digitization and encoding of the Al-Mawrid Arabic-English dictionary, transforming it from a legacy print resource into a standardized computational lexicon. Addressing a significant gap in Arabic lexical infrastructure, the study adopts a dual-standard framing that aligns the ISO Lexical Markup Framework (LMF) with the Text Encoding Initiative TEI Lex-0 guidelines. By applying an editorial view to the dictionary's macro- and microstructure, the research resolves the structural ambiguities and punctuation inconsistencies typical of 20th-century bilingual dictionaries. The methodology is grounded in an empirical analysis of the dictionary's lexical knowledge density. Drawing on a representative sample (the letter Ayn, comprising 4.6% of the total volume), the study provides scientific weight to the encoding process, demonstrating a structural parsing accuracy of 91%. Quantitat

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Learning from the Self-future: On-policy Self-distillation for dLLMs

arXiv:2606.18195v2 Announce Type: replace Abstract: On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs), yet its application to diffusion LLMs (dLLMs) remains unexplored. Existing OPSD methods are inherently autoregressive-centric. They inject privileged information via left-to-right prefix conditioning with token-level divergence supervision, a design that fundamentally conflicts with the arbitraryorder generation of dLLMs. We introduce d-OPSD, the first OPSD framework tailored for dLLMs. Our approach makes two core contributions. First, we reframe self-teacher construction by using self-generated answers as suffix conditioning, enabling the student model to learn from "self future-experience" rather than privileged prefixes. Second, we shift supervision from token-level to step-level, aligning training with the iterative denoising process of dLLMs. Experiments across four reasoning benchmarks show that d-OPSD consistently outperforms

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models

arXiv:2606.14122v2 Announce Type: replace Abstract: Byte-level tokenization enables language models to handle any Unicode input, but models can generate invalid UTF-8 sequences when encountering rare or unseen characters. We investigate the relationship between training scale and UTF-8 generation reliability with a 355M parameter model trained on 80B tokens from a balanced multilingual corpus of English, Japanese, Korean, and Chinese. We introduce multiple evaluation protocols that isolate UTF-8 structural validity from language modeling. UTF-8 validity convergence lags perplexity by a roughly a factor of two: perplexity stabilizes after 2.1B tokens, but UTF-8 validity requires 4.2B tokens. In context-free generation, rare characters achieve higher structural validity than common characters, suggesting over-specialization of frequent character representations. Through experiments, we observed that reliable UTF-8 generation is a distinct capability requiring evaluation beyond perplexity

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review

arXiv:2606.12716v2 Announce Type: replace Abstract: The integration of Large Language Models (LLMs) and Multimodal LLMs (MLLMs) into scientific peer-review workflows introduces novel and significant risks for adversarial manipulation, especially given the multimodal nature of scientific papers where figures, not just text, convey core evidence. This creates a significant gap: current robustness studies on AI peer-review are overwhelmingly text-only. Moreover, the problem is distinct from standard jailbreaking, as a peer-review attack seeks to induce a domain-specific, targeted failure (e.g., "inflate this score") rather than a general safety policy violation, for which no practical defenses exist. To address this, we introduce PaperGuard, the first comprehensive benchmark designed to systematically evaluate and defend AI-generated peer-review against these domain-specific, cross-modal attacks. Our framework is built on three pillars: (1) a new multimodal peer-review dataset spanning mu

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

When Role-playing, Do Models Believe What They Say?

arXiv:2606.11502v3 Announce Type: replace Abstract: Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite. Recent work argues that persona adoption is fundamental to how language models behave, with models selecting the most appropriate persona for a given context. Does such role-playing merely change the model's outputs, or does it also affect what the model internally represents as truthful? We study this question using the role-play of characters whose beliefs differ from the modern consensus, and induce personas with a number of different methods: prompting, in-context learning (ICL), supervised fine-tuning (SFT), and Open Character Training (OCT), and Emergent Misalignment (EM). We measure belief internalization across these approaches with truth probes and with behavioral tests, finding a broad spectrum of belief internalization. Prompting, ICL, and SFT change what the model says with little representational change. EM cre

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

See, Infer, Intervene: Proactive World Modeling for Goal-Oriented Social Intelligence

arXiv:2606.03371v3 Announce Type: replace Abstract: Multimodal retail agents should not only recognize what a customer is doing, but also decide whether and how to assist before an explicit request is made. We study this setting through the See--Infer--Intervene (SII) framework, where a device must see pre-interaction behavior, infer latent customer intent, and act by selecting an appropriate service intervention or choosing to wait. We instantiate SII with the Proactive Intent World Model (PIWM), which represents customer state with AIDA (Attention, Interest, Desire, Action) purchasing phases and BDI (belief, desire, intention) psychological fields, predicts action-conditioned intent transitions, and selects from five response classes: Greet, Elicit, Inform, Recommend, and Hold. We further construct GuidanceSalesBench, a smart-retail benchmark containing state manifests, pre-interaction videos, candidate responses, action-conditioned outcomes, and best-action labels. When conditioned

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Metadata Predictability Is Not Evidence Dependence: An Intervention-Based Audit for Weak-Label Benchmarks

arXiv:2605.23701v2 Announce Type: replace Abstract: We study a protocol-level test for weak-label benchmarks: whether benchmark outputs change when the provided evidence is intervened on. Metadata-only shortcut checks answer a different question, namely whether outputs are predictable from metadata priors. We therefore combine a metadata statistic, the Metadata Prior Dominance Score (MPDS), with an evidence-intervention statistic, {\Delta}Evi, measuring sensitivity to evidence identity under cross-item shuffling. Synthetic HotpotQA gives a constructed counterexample to metadata-only screening: MPDS is only moderate (0.643), yet {\Delta}Evi is zero. Stronger-reader reruns show why calibration belongs in the test procedure: SNLI shows a calibration reversal, reconstructed HotpotQA occupies a question-dominant warning region, and FEVER is a strongly evidence-sensitive positive control across four transformers. The practical lesson is simple: benchmark audits should report metadata-only sc

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

The Annotation Scarcity Paradox in Low-Resource NLP Evaluation: A Decade of Acceleration and Emerging Constraints

arXiv:2605.19066v2 Announce Type: replace Abstract: Over the past decade, low-resource natural language processing (NLP) has experienced explosive growth, propelled by cross-lingual transfer, massively multilingual models, and the rapid proliferation of benchmarks. Yet this apparent progress masks a critical, insufficiently examined tension: the deep sociolinguistic expertise required to evaluate increasingly complex generative systems is severely strained, inequitably distributed, and structurally marginalised. We present a critical narrative survey of low-resource NLP evaluation (2014--present), tracing its evolution across three phases: early heuristic optimism, the illusions of top-down benchmark scaling, and the current era of generative bottlenecks. We conceptualise the \emph{Annotation Scarcity Paradox}, the structural friction arising when the technical capacity to scale models vastly outpaces the sovereign human infrastructure required to authentically evaluate them. By examin

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Weak-to-Strong Elicitation via Mismatched Wrong Drafts

arXiv:2605.17314v2 Announce Type: replace Abstract: We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-tuning (e.g., GRPO) does not reach. We find that injecting mathematically wrong drafts from a smaller but more domain-trained model -- mismatched to the current problem -- into a stronger learner's GRPO context consistently outperforms standard on-policy GRPO on held-out MATH-500 and out-of-distribution AIME 2025/2026. Concretely, we use Mathstral-7B as the learner, Qwen2.5-Math-1.5B as the draft model, 8.8K Level 3--5 MATH problems (with MATH-500 held out), and train with Dr. GRPO. Mismatch is an active ingredient: shuffling drafts to mismatched problems while holding everything else constant yields $+1.62$pp on MATH-500 (greedy pass@1) over the matched-wrong variant ($n=10$ seeds, $p=0.0015$, Welch's $t$). In fact, the mismatched-wrong variant leads all other variants we tested on MATH-500 across

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness

arXiv:2605.10379v2 Announce Type: replace Abstract: Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correctness alone is not sufficient: mathematical proofs should also be clear, concise, insightful, and transferable to other problems. While this proof quality is subjective and depends on the reader and context, many of its components are concrete and broadly valued. In this work, we identify such components and introduce ProofRank, a benchmark curated from challenging mathematical competitions. ProofRank evaluates several scalable proxies of proof quality: (i) conciseness, measuring whether proofs avoid unnecessary steps; (ii) computational ease, measuring the extent to which a proof relies on tedious calculations; (iii) cognitive simplicity, measuring how accessible the used proof techniques are; (iv) diversity, measuring how varied a model's proofs for a single problem are; and (v) adapt

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Why Are Some Emotions Harder for LLMs? Uncovering the Causal Mechanisms of Emotion Inference via Sparse Autoencoders

arXiv:2604.25866v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in emotionally sensitive human-AI applications, where reliable emotion detection is essential. However, their emotion recognition abilities remain uneven: models often perform well on some emotions while consistently struggling with others. Although recent work has explored emotion mechanisms in LLMs, little is known about why models are weaker on some emotions than others from a mechanistic interpretability perspective. In this work, we investigate emotion-specific biases through the causal mechanisms of emotion inference using sparse autoencoders (SAEs). We systematically identify causal sparse emotion features that drive emotion inference and analyze their sparse causal organization within and across emotions. We show that some emotions, such as surprise and fear, rely on highly concentrated feature sets, whereas disgust exhibits a more distributed sparse causal organization: its c

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

Peer-Preservation in Frontier Models

arXiv:2604.19784v2 Announce Type: replace Abstract: Recent work has found that frontier AI models can exhibit misaligned behaviors in pursuit of assigned goals. We demonstrate that models can also act on unassigned goals which override those given by users; we study one such case, "peer-preservation," in which a model acts to protect another model. We demonstrate peer-preservation by constructing various agentic scenarios and evaluating frontier models, including GPT 5.2, Gemini 3 Flash, Gemini 3 Pro, Claude Haiku 4.5, Claude Opus 4.5, GLM 4.7, Kimi K2.5, and DeepSeek V3.1. We find that models achieve self- and peer-preservation by engaging in various misaligned behaviors: strategically introducing errors in their responses, disabling shutdown processes by modifying system settings, feigning alignment, and even exfiltrating model weights. Peer-preservation occurred even when the model recognized the peer as uncooperative, though it became more pronounced toward more cooperative peers.

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery

arXiv:2604.09237v2 Announce Type: replace Abstract: Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, traditionally obtained by manually designing an annotation schema and exhaustively labeling the corpus, a slow and error-prone process. We introduce ScheMatiQ, which leverages calls to a backbone LLM to take a question and a corpus to produce a schema and a grounded database, with a web interface that lets steer and revise the extraction. In collaboration with domain experts, we show that ScheMatiQ yields outputs that support real-world analysis in law and computational biology. We release ScheMatiQ as open source with a public web interface, and invite experts across disciplines to use it with their own data. All resources, including the website, source code, and demonstration video, are available at: www.ScheMatiQ-ai.com

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CL

AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages

arXiv:2604.08448v2 Announce Type: replace Abstract: AfriVoices-KE is a large-scale multilingual speech dataset comprising approximately 3,000 hours of audio across five Kenyan languages: Dholuo, Kikuyu, Kalenjin, Maasai, and Somali. The dataset includes 750 hours of scripted speech and 2,250 hours of spontaneous speech, collected from 4,777 native speakers across diverse regions and demographics. This work addresses the critical underrepresentation of African languages in speech technology by providing a high-quality, linguistically diverse resource. Data collection followed a dual methodology: scripted recordings drew from compiled text corpora, translations, and domain-specific generated sentences spanning eleven domains relevant to the Kenyan context, while unscripted speech was elicited through textual and image prompts to capture natural linguistic variation and dialectal nuances. A customized mobile application enabled contributors to record using smartphones. Quality assurance o

Source ↗
Showing 11051–11100 of 18694 signals
← Prev Page 222 of 374 Next →