EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CY

The Governance Inversion Hypothesis: Why More AI Regulation May Produce Less Organisational Control

arXiv:2606.26117v1 Announce Type: new Abstract: This paper introduces the Governance Inversion Hypothesis (GIH) to explain a growing paradox in artificial intelligence (AI) governance: under conditions of increasing regulatory expansion and technological complexity, organisations may become more formally governed while simultaneously experiencing a decline in operational control over AI systems. Existing AI governance frameworks generally assume that stronger regulation improves accountability, oversight, and organisational control. This paper challenges that assumption by arguing that governance formalisation itself may contribute to the erosion of control in AI-intensive environments. Drawing on institutional theory, organisational governance research, accountability scholarship, and emerging AI governance literature, the paper develops a conceptual framework explaining how regulatory expansion may weaken operational authority through four interconnected mechanisms: authority fragmen

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CY

Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation

arXiv:2606.26116v1 Announce Type: new Abstract: A brand whose customers use both ChatGPT and Claude for product recommendations faces a strategic choice: a single optimization playbook, or one per provider? Across 215 commercially-framed prompts in four measurement batches, the two providers disagree on which brands they recommend roughly two-thirds of the time (cross-provider recommendation Jaccard 0.35, below the 0.50-0.61 same-prompt rerun baseline). The picks diverge. But when neither provider recommends a brand, we classify the failure into one of three modes -- discoverability (the brand never reaches the model), compellingness (it reaches the model but isn't mentioned), or positioning (it's mentioned but not recommended) -- and on 7,763 such joint failures, both providers diagnose the same failure mode 95.1% of the time (clustered 95% CI [94.3%, 95.7%]). Agreement rises monotonically with falling brand prominence, from 81% [78.2%, 84.0%] on category leaders to 99.6% [99.3%, 99.9

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CY

A Multi-Layer AI Framework for Information Landscape Analysis

arXiv:2606.26115v1 Announce Type: new Abstract: This paper proposes a multi-layer AI framework for information landscape analysis in the context of information disorder. Rather than treating misinformation detection as a binary fact-checking task, the framework analyzes political and media content across multiple dimensions, including source reliability, factual structure, framing, bias, emotional activation, manipulation patterns, and propagation dynamics. The goal is to move beyond isolated claim verification toward a structured representation of the informational environment surrounding an event, entity, or narrative. We argue that AI systems for media analysis should support epistemic mapping: a transparent, multi-dimensional account of how facts, interpretations, actors, and narratives interact over time. The paper presents the conceptual architecture, analytical layers, and methodological rationale of the framework, with the aim of supporting more nuanced, explainable, and critic

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CY

Dream machine -- the next creative economy

arXiv:2606.26114v1 Announce Type: new Abstract: We examine the structural transformation of creative industries under generative artificial intelligence, drawing on 374 primary sources spanning policy documents, industry data, creator surveys, and platform analytics. Beginning with the December 2024 release of OpenAI's Sora video model as a watershed event, we trace the historical pattern of creative resistance to technological disruption, then develop an analytical framework -- the Human-AI Agency Continuum for mapping the spectrum of human and machine collaboration in creative work. We present evidence for the "slop ceiling," an audience-imposed quality threshold that constrains AI-generated content to approximately 1--3% of platform streams despite comprising 44% of uploads. Analysis of the UK Government's 2025 consultation on AI and copyright (over 11,500 responses, 88% opposing expanded AI training rights) reveals deep structural tensions between technology firms and creative work

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CY

Generative AI and Copyright Infringement: A Legal-Technical Analysis of AI Music Generation Systems Under 17 U.S.C. Title 17

arXiv:2606.26111v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) has enabled users to synthesize music with text prompts, combining copyrighted lyrics, AI-composed melodies, and synthetic vocals that imitate real artists. This paper examines the legal and technical dimensions of AI-based music creation (e.g., Google Gemini's music tools) under U.S. copyright law. We analyze whether a user who inputs one artist's protected lyrics into a GenAI system, directs it to use another artist's voice or style, publishes the resulting song, and monetizes it violates 17 U.S.C. Section 106's exclusive rights [3]. The analysis integrates Title 17 doctrine (rights of reproduction, derivative works, distribution), 17 U.S.C. Section 114's narrow sound recording protection [4], and the new voice-cloning laws emerging at the state level [20]. We argue that unauthorized lyric copying poses a high risk of infringement of the musical composition, whereas mere AI-generated voice imit

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CY

Simulating Eating Disorder Patients with LLMs: Evaluating Psychological Persona Stability in Multi-Turn Conversations

arXiv:2606.26109v1 Announce Type: new Abstract: Large language model (LLM)-based simulations of clinical patients are increasingly used for research and training, yet their validity requires persona stability: coherent maintenance of an assigned psychological profile across and within conversations. We evaluate this prerequisite using eating disorder personas grounded in five published case vignettes, a dual-assessment framework (self-report + independent observer ratings), and validated psychometric instruments (EDE-Q) with known ground-truth scores. Across six LLMs and two experiments (between-conversation stability (Exp. I) and within-conversation stability (Exp. II)), we find that LLMs are paradoxically too stable and too inaccurate: variability is negligible, yet all models systematically overshoot ground-truth severity by 12-30% of the scale range (0.7-1.8 points on a 0-6 scale). The mechanism is selective stereotyping: models differentiate cases on behavioural items (dietary res

Source ↗
technology Fri, 26 Jun 2026 00:00:00 -0400
arXiv cs.CY

Benchmarking Open-Weight Foundation Models for Global AI Technical Governance

arXiv:2606.26099v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in artificial intelligence (AI) governance analysis across national and international organisations. There is, however, growing evidence that such models produce significantly less accurate responses for countries that are underrepresented in their training data-a pattern described in existing literature as geographic bias. Existing studies examining this phenomenon are subject to three methodological limitations that together undermine their findings: (1) reliance on proprietary systems whose weights are not publicly released, which prevents independent replication; (2) evaluation of model knowledge about years that fall after data collection for model training had concluded, leading to geographic ignorance in addition to the natural limits of each model's knowledge; and (3) use of coarse binary response classification that cannot distinguish models' confident fabrication (HF) from t

Source ↗
technology Fri, 24 Jul 2026 17:14:00 +0000
MedCity News

Priority Health Launches 2 Cancer Support Solutions with Color Health, Grail

Priority Health partnered with Color Health and Grail to offer self-funded employers virtual cancer navigation, support and multi-cancer early detection beginning in 2027. The post Priority Health Launches 2 Cancer Support Solutions with Color Health, Grail appeared first on MedCity News .

Source ↗
technology Fri, 24 Jul 2026 16:18:15 +0000
HN: education

The Direct and Indirect Effects of Genetics and Education

Article URL: https://arxiv.org/abs/2607.19562 Comments URL: https://news.ycombinator.com/item?id=49037832 Points: 4 # Comments: 0

Source ↗
technology Fri, 24 Jul 2026 13:45:00 +0000
MedCity News

AI Will Reshape the Patient Experience

We’re about to enter a world beyond just talking to a chat tool about your health. AI tools are on the precipice of being able to take action on their advice. The post AI Will Reshape the Patient Experience appeared first on MedCity News .

Source ↗
technology Fri, 24 Jul 2026 13:20:25 +0000
MedCity News

Healthcare’s Next Source of Alpha Is Moving Upstream

The question is no longer whether healthcare innovation is accelerating. The question is where investors gain access to generate the strongest returns. The post Healthcare’s Next Source of Alpha Is Moving Upstream appeared first on MedCity News .

Source ↗
technology Fri, 24 Jul 2026 11:02:59 +0000
HN: education

Gen Z protests in India demanding resignation of Education Minister

Article URL: https://www.bbc.com/news/videos/c3307vgneveo Comments URL: https://news.ycombinator.com/item?id=49033802 Points: 6 # Comments: 4

Source ↗
technology Fri, 24 Jul 2026 09:00:00 +0000
Tech & Learning

Helping Students Explore Space And The World Beyond Their Classrooms

Innovative Leader Award - NASA Solar System Ambassador Tim Needles shares how exploring the world and space can provide learning experiences to infinity and beyond

Source ↗
technology Fri, 24 Jul 2026 09:00:00 +0000
eCampus News

Intentional governance: Responsibility, judgment, and leadership in the age of intelligent systems

Artificial intelligence is changing the conditions under which institutional purpose is interpreted, pursued, and realized through the exercise of institutional judgment. As institutions increasingly adopt AI, discussions related to governance have expanded to include fairness, bias, transparency, explainability, privacy, accountability, and regulatory compliance. The post Intentional governance: Responsibility, judgment, and leadership in the age of intelligent systems appeared first on eCampus News .

Source ↗
technology Fri, 24 Jul 2026 06:30:54 +0000
MedCity News

CRISPR Biotech Scribe Therapeutics Writes a New Chapter With $129M IPO

Scribe Therapeutics’ IPO will support a pipeline led by a genetic medicine that represses expression of its target gene without making permanent changes. Common chronic cardiometabolic conditions are Scribe’s initial focus as the biotech works to bring genetic medicines beyond rare diseases. The post CRISPR Biotech Scribe Therapeutics Writes a New Chapter With $129M IPO appeared first on MedCity News .

Source ↗
technology Fri, 24 Jul 2026 04:46:02 +0000
MedCity News

Healthcare Investors Are Rewriting the ROI Playbook from Cost-Cutting to Cash Flow

Unlike past healthcare tech waves sold on cost savings, AI is proving it can generate new revenue — with returns as high as 100x, according to a panel of investors at MedCity News’ Bullseye event. The post Healthcare Investors Are Rewriting the ROI Playbook from Cost-Cutting to Cash Flow appeared first on MedCity News .

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents

arXiv:2606.21262v2 Announce Type: replace-cross Abstract: Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad. Rubric-based rewards improve interpretability through natural-language criteria, but existing methods share two limitations: they score at the trajectory level, offering no guidance for individual steps; and their scorer is closed-source and static, so it cannot adapt as the agent evolves during training. We propose ARCO (Adaptive Rubric CO-evolution), which generates a per-step rubric and predicts a rubric-conditioned step-level reward for each action, and continually updates this rubric model on on-policy rollouts so that its criteria and scores co-evolve with the agent's improving behavior. Across HotpotQA, 2WikiMultiHopQA, and MuSiQue with two open-source backbones, ARCO achieves the highest EM in all settings over outcome-, rubric-, and process-reward baselines, and analys

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Understanding and Accelerating the Training of Masked Diffusion Language Models

arXiv:2605.13026v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models (ARMs) for language modeling. However, MDMs are known to learn substantially more slowly than ARMs, which may become problematic when scaling MDMs to larger models. Therefore, we ask the following question: how can we accelerate standard MDM training while maintaining its final performance? To this end, we first provide a detailed analysis of why MDM training is slow. We find that the main factor is the locality bias of language: the predictive information for a token is concentrated in nearby positions. We further investigate how this bias slows learning and suggest a simple yet effective remedy: bell-shaped time sampling as a training strategy. Notably, MDMs trained with our training recipe reach the same validation negative log-likelihood (NLL) up to $\sim4\times$ faster than standard training on One Billion Word Benchmark (LM1B).

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training

arXiv:2603.06642v2 Announce Type: replace-cross Abstract: Test-Time Training (TTT) language models replace the KV-cache with fast weights updated during inference, achieving O(1) memory but suffering catastrophic failure on exact-recall tasks. Version 1 of this work proposed SR-TTT, which routes high-surprisal tokens to a sparse exact-attention Residual Cache, and reported large Needle-in-a-Haystack gains. We show those gains were evaluation artifacts: the loss and metric read logits at the answer positions rather than one position earlier, training both models to copy an answer already visible in their input (a model trained on retrieval-impossible data reaches 100% accuracy under the flawed metric); additionally, the cache attended non-causally over future tokens, including the answer itself. We release a corrected implementation with startup causality self-tests, then ask whether the SR-TTT hypothesis survives correction. It does not, and the failure decomposes into two independent,

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance

arXiv:2511.12997v2 Announce Type: replace-cross Abstract: Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex browsing tasks across diverse domains. However, current agents struggle with repetitive errors and lack the ability to learn from past experiences across sessions, limiting their long-term robustness and sample efficiency. We introduce WebCoach, a model-agnostic self-evolving framework that equips web browsing agents with persistent cross-session memory, enabling improved long-term planning, reflection, and continual learning without retraining. WebCoach consists of three key components: (1) a WebCondenser, which standardizes raw navigation logs into concise summaries; (2) an External Memory Store, which organizes complete trajectories as episodic experiences; and (3) a Coach, which retrieves relevant experiences based on similarity and recency, and decides whether to inject task-specific advice

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Content Anonymization for Privacy in Long-form Audio

arXiv:2510.12780v3 Announce Type: replace-cross Abstract: Voice anonymization techniques have been found to successfully obscure a speaker's acoustic identity in short, isolated utterances in benchmarks such as the VoicePrivacy Challenge. In practice, however, utterances seldom occur in isolation: long-form audio is commonplace in domains such as interviews, phone calls, and meetings. In these cases, many utterances from the same speaker are available, which pose a significantly greater privacy risk: given multiple utterances from the same speaker, an attacker could exploit an individual's vocabulary, syntax, and turns of phrase to re-identify them, even when their voice is completely disguised. To address this risk, we propose a new approach that performs a contextual rewriting of the transcripts in an ASR-TTS pipeline to eliminate speaker-specific style while preserving meaning. We present results in a long-form telephone conversation setting demonstrating the effectiveness of a cont

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Simple Policy Gradients for Reasoning with Diffusion Language Models

arXiv:2510.04019v3 Announce Type: replace-cross Abstract: Diffusion large language models (dLLMs) represent a promising alternative to autoregressive LLMs; however, the lack of effective post-training techniques, including reinforcement learning (RL), remains a key challenge for dLLMs, especially for downstream applications. Existing approaches often rely on a sequence-level view that requires biased likelihood approximations. In this work, we propose Amortized Group Relative Policy Optimization (AGRPO), a policy gradient algorithm that leverages the Markovian nature of dLLMs, optimizing individual denoising steps rather than full sequences. Our approach improves alignment between the trained policy and the inference process and also admits efficient, unbiased gradient updates via a novel timestep estimation scheme. We demonstrate AGRPO's effectiveness on different math and reasoning tasks, achieving absolute accuracy gains of +59.4\% and +69.7\% on Countdown and Sudoku over the base L

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

arXiv:2508.05502v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings. We argue that this failure is not merely a linguistic limitation: culture-specific visual knowledge depends on native visual-textual alignments that translation-centric pipelines rarely provide. We present MELLA, a multimodal dataset across eight low-resource languages, designed to support linguistic fluency and cultural groundedness. MELLA uses a dual-source strategy that combines native web image-alt-text pairs for culture-grounded supervision with generated-and-translated image descriptions for linguistically rich supervision, explicitly separating two learning signals often conflated in multilingual multimodal data. Through controlled diagnostic fine-tuning on multiple MLLM backbones, we show that MELLA mitigates cultural hallucination by helping models re

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

How Robust Is Homogeneity Bias in LLMs? Evidence Across Models, Decoding Settings, and Identity Signals

arXiv:2501.02211v3 Announce Type: replace-cross Abstract: Large language models (LLMs) reproduce homogeneity bias -- the tendency to portray marginalized groups as more internally similar than dominant groups -- but whether this bias generalizes across models, is stable under different inference settings, or depends on how group identity is signaled remains unstudied. We map homogeneity bias across seven open-weight instruction-tuned LLMs (7-20B parameters), a 5x5 temperature x top-p decoding grid, and two paradigms for signaling group identity (explicit labels vs. racially distinctive names). In six of seven models, Hispanic and Asian Americans are portrayed as significantly more homogeneous than White Americans at the default configuration, and the effect remains positive on average at every temperature and top-p tested; African American and gender bias instead vary in direction across models. A conservative cell-level re-analysis confirms Hispanic and Asian homogeneity as robust whi

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning

arXiv:2607.09328v2 Announce Type: replace Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed safety check that jointly explain a disaster may appear dozens of sections apart; in a novel, a character's true motive may surface only through scenes far removed from the moment it becomes relevant. This source-internal evidence integration is central to real-world long-document analysis, yet existing benchmarks largely sidestep it. Needle probes, planted facts, and reverse-engineered multi-hop chains embed evidence that may differ from the host text in distribution, placement, or register, making it unclear whether strong performance reflects genuine source reasoning or distributional artifacts. We introduce WILDTRACE, a benchmark of 481 tasks over 214 naturally occurring long-form sources such as technical

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs

arXiv:2605.28346v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly evaluated for whether they identify the right visual content, but little is known about whether they express such content in a discourse-appropriate form. We address this research gap using information structure (IS), testing whether VLMs distinguish discourse-old Topics from discourse-new Foci in visually grounded question answering. We exploit Hungarian, a language in which Topic and Focus map onto dedicated syntactic positions, making IS choices observable in text. Comparing six VLMs with human participants, we find that models produce IS-relevant constructions, but over-regularise this sensitivity. Under the interacting pressures of discourse status, grammatical role (preference for subject Topics) and definiteness (preference for indefinite Foci), humans choose variable strategies for IS realisation. VLMs, by contrast, collapse onto narrow response templates, resembling mode collapse

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation

arXiv:2605.27390v3 Announce Type: replace Abstract: Speculative decoding accelerates Large Language Model inference through draft-then-verify generation, yet lightweight draft models face coupled efficiency and quality limitations: large-vocabulary output projection is costly, while limited draft capacity and static parameters reduce acceptance under specialized or shifting inputs. Vocabulary pruning lowers projection cost, but static variants miss locally important long-tail tokens, while dynamic variants remain sensitive to preset selection policies and budgets. Moreover, limited draft capacity can leave the draft distribution misaligned even when the target token is covered. Online alignment improves draft quality, but full-parameter updates introduce substantial memory and latency overhead. We introduce EvoSpec, which jointly adapts the active vocabulary and lightweight draft parameters from verification feedback. EvoSpec asynchronously retrieves semantic and statistical token neig

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation

arXiv:2605.25572v2 Announce Type: replace Abstract: The growing complexity of quantum programming frameworks has exposed a critical limitation in existing large language model (LLM)-based code assistants: general-purpose models hallucinate PennyLane-specific gate names, misplace device configurations, and produce structurally invalid circuits when faced with specialized quantum coding challenges. We present PennySynth, a retrieval-augmented generation framework that addresses this gap by conditioning LLM inference on a curated knowledge base of 13,389 PennyLane instruction-code pairs, built via a three-stage extraction, verification, and deduplication pipeline over official PennyLane repositories, community GitHub sources, and QHack competition archives. PennySynth introduces a code-aware embedding strategy using st-codesearch-distilroberta-base, trained for natural-language-to-code retrieval, increasing average retrieval cosine similarity from 0.45 to 0.726 compared to a general-purpo

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora

arXiv:2605.22660v2 Announce Type: replace Abstract: Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slang, and cultural references introduce hard-to-avoid translation artifacts. Yet automated moral classification depends on language-specific annotated corpora that exist almost exclusively in English. We investigate whether LLM-based translation can bridge this gap, taking Polish as a test case. Using $\sim$50k morally annotated social media posts from a diverse range of topics, we apply a principled four-method validation pipeline: LaBSE cross-lingual embedding similarity, Centred Kernel Alignment (CKA), LLM-as-judge evaluation, and deep learning classifier parity tests. We show that despite shortcomings in handling slang, vulgarity, and culturally loaded expressions, direct translation preserves subtle moral cues well enough to be harvested by cross-lingual machine learning --- with a mean cosine si

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routing

arXiv:2605.22057v2 Announce Type: replace Abstract: Enterprise routers assign queries to expert agents, yet deployed profiles stay static while agents evolve (prompts, tools, models), and developers rarely keep descriptions or exemplars current. We present FlyRoute, a self-evolving profiling framework that grows capability evidence from real traffic: dispatch candidates, quality-gate successful pairs into each agent's success store, periodically distill evidence into learned capability descriptions, and inject those descriptions together with BM25-retrieved successes into an LLM router. To make this flywheel data-efficient, FlyRoute introduces a targeted exploration policy that combines profile uncertainty, BM25 relevance, and lexical novelty, prioritizing under-profiled agents only for plausible queries and avoiding redundant evidence collection. In experiments on our proprietary enterprise developer-support dataset of real routed queries, FlyRoute improves a same-backbone zero-shot L

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

arXiv:2604.11666v2 Announce Type: replace Abstract: As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dialogue partners (i.e., form and use a theory-of-mind, or ToM) becomes increasingly critical for safe interaction with potentially adversarial partners. We propose a novel privacy-themed ToM challenge, ToM for Steering Beliefs (ToM-SB), in which a defender must act as a Double Agent to steer the beliefs of an attacker with partial prior knowledge within a shared universe. To succeed on ToM-SB, the defender must engage with and form a ToM of the attacker, with a goal of fooling the attacker into believing they have succeeded in extracting sensitive information. We find that strong frontier models like Gemini3-Pro and GPT-5.4 struggle on ToM-SB, often failing to fool attackers in hard scenarios with partial attacker prior knowledge, even when prompted to reason about the attacker's beliefs (T

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues

arXiv:2604.01925v2 Announce Type: replace Abstract: Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly. Existing benchmarks use name based proxies to detect implicit biases, which carry weak associations with many social demographics and cannot extend to dimensions like age or socioeconomic status. We introduce ImplicitBBQ, a QA benchmark that evaluates implicit bias through characteristic based cues, demographically associated attributes that signal implicitly, across age, gender, region, religion, caste, and socioeconomic status. Evaluating 11 models, we find that implicit bias in ambiguous contexts is over six times higher than explicit bias in open weight models. Notably, this bias is distributed unevenly across demographics: caste emerges as the most severe while gender is the least affected. Safety prompting and chain-of-thought reasoning fail to subs

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

LinearARD: Linear-Memory Attention Distillation for RoPE Restoration

arXiv:2604.00004v2 Announce Type: replace Abstract: The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight Continual Pre-Training (CPT). While effective for processing long sequences, this paradigm often disrupts original model capabilities, leading to performance degradation on standard short-text benchmarks. We propose LinearARD, a self-distillation method that restores Rotary Position Embeddings (RoPE)-scaled students through attention-structure consistency with a frozen native-RoPE teacher. Rather than matching opaque hidden states, LinearARD aligns the row-wise distributions of dense $Q/Q$, $K/K$, and $V/V$ self-relation matrices to directly supervise attention dynamics. To overcome the quadratic memory bottleneck of $n \times n$ relation maps, we introduce a linear-memory kernel. This kernel leverages per-token log-sum-exp statistics and fuses logit recomputation into the backward pass to compute

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Gumbel Distillation for Parallel Text Generation

arXiv:2603.22216v2 Announce Type: replace Abstract: The slow, sequential nature of autoregressive (AR) language models has driven the adoption of parallel decoding methods. However, these non-AR models often sacrifice generation quality as they struggle to model the complex joint distribution of token sequences. To narrow this performance gap, we introduce Gumbel Distillation, a novel distillation technique that enables parallel decoders to learn this distribution effectively. Our method leverages the Gumbel-Max trick to create a deterministic mapping from a latent Gumbel noise space to the output tokens of a high-performing AR teacher. As a model-agnostic technique, Gumbel Distillation seamlessly integrates with diverse parallel decoding architectures, including MDLM and BD3-LM. Experiments on LM1B and OpenWebText show that Gumbel Distillation substantially improves the generation quality of parallel language models, achieving a 30.0% improvement in MAUVE score and 10.5% in generative

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining

arXiv:2603.11838v2 Announce Type: replace Abstract: Large language models pretrained on internet-scale data risk lookahead bias in forecasting tasks, as they may have already seen the true outcome during training. To address this, we present DatedGPT, a family of twelve 1.3B-parameter language models trained from scratch on approximately 100 billion tokens each with strict annual data cutoffs spanning 2013 to 2024, together with DatedInstruct, an instruction dataset grounded in each year's documents to prevent leakage during post-training. The models are competitive with open models of similar scale, and perplexity-based probing confirms that each model's knowledge is bounded by its cutoff year. On stock return prediction over 61,000 firm-day news headlines, DatedGPT-instruct achieves an annualised Sharpe ratio of $3.20$ under the lookahead-bias-free setup. Lookahead-biased models, whose training data covers the outcome period, add a lookahead premium of $26.4$ b.p. per standard deviat

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Non-Zipfian Distribution of Stopwords or Function Words and Subset Selection Models

arXiv:2603.04691v2 Announce Type: replace Abstract: Stopwords and function words are relatively less informative for the content of a language and more often play a structural role in a sentence. Stopwords are ubiquitous words and may contain verbs, adjectives and adverbs. On the other hand, function words are strictly prepositions, conjunctions, pronouns, determiners, qualifiers, articles, interrogatives, and a limited number of auxiliary verbs. In contrast to the well known Zipf's law for rank-frequency plot for all words, the rank-frequency plots for stopwords or function words are best fitted by the Beta Rank Function (BRF). On the other hand, the rank-frequency plots of non-stopwords or non-function-words also deviate from the Zipf's law, but are better described by a quadratic function of log-token-count over log-rank than by BRF. Based on the observed rank of stopwords or function words in the full word list, we propose a stopword/function word/subset selection model that the pr

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents

arXiv:2601.01885v3 Announce Type: replace Abstract: Large language model (LLM) agents face fundamental limitations in long-horizon reasoning due to finite context windows, making effective memory management critical. Existing methods typically handle long-term memory (LTM) and short-term memory (STM) as separate components, relying on heuristics or auxiliary controllers, which limits adaptability and end-to-end optimization. In this paper, we propose Agentic Memory (AgeMem), a unified framework that integrates LTM and STM management directly into the agent's policy. AgeMem exposes memory operations as tool-based actions, enabling the LLM agent to autonomously decide what and when to store, retrieve, update, summarize, or discard information. To train such unified behaviors, we propose a three-stage progressive reinforcement learning strategy and design a step-wise GRPO to address sparse and discontinuous rewards induced by memory operations. Experiments on five long-horizon benchmarks

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation

arXiv:2512.07540v4 Announce Type: replace Abstract: Error Span Detection (ESD) extends automatic machine translation (MT) evaluation by localizing translation errors and labeling their severity. Current generative ESD methods typically use Maximum a Posteriori (MAP) decoding, assuming that the model-estimated probabilities are perfectly correlated with similarity to the human annotation, but we often observe higher likelihood assigned to an incorrect annotation than to the human one. We instead apply Minimum Bayes Risk (MBR) decoding to generative ESD. We use a sentence- or span-level similarity function for MBR decoding, which selects candidate hypotheses based on their approximate similarity to the human annotation. Experimental results on the WMT24 Metrics Shared Task show that MBR decoding significantly improves span-level performance and generally matches or outperforms MAP at the system and sentence levels. To reduce the computational cost of MBR decoding, we further distill its

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances

arXiv:2511.03354v2 Announce Type: replace Abstract: Generative artificial intelligence (GenAI) is transforming bioinformatics by advancing genomics, proteomics, transcriptomics, structural biology, and drug discovery. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses framework, this review addresses six research questions to evaluate influential GenAI strategies in terms of methodological innovation, predictive performance, specialization, limitations, and data use. RQ1 shows that GenAI supports sequence analysis, molecular design, and integrative data modelling, often outperforming traditional methods through improved pattern recognition and generation. RQ2 finds that specialized architectures generally outperform general-purpose models because of domain-specific pretraining and context-aware design. RQ3 identifies benefits in molecular analysis and biological data integration, including improved accuracy and reduced analytical error. RQ4 reports advance

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs

arXiv:2509.14257v3 Announce Type: replace Abstract: Large Language Model agents achieve strong performance on multi-step reasoning and tool-use tasks, but their impressive capabilities typically rely on extremely large backbones. Existing distillation approaches train smaller students to imitate full teacher trajectories, yet reasoning and knowledge gaps between the teacher and student can cause compounding errors. We propose SCoRe, a student-centered framework in which the student generates training trajectories and the teacher corrects only the earliest error, producing training data matched to the student's abilities and exposing specific weaknesses. The student is first fine-tuned on corrected trajectories. Subsequently, short-horizon reinforcement learning starts from the verified prefix preceding the earliest error, with target rewards assigned at that step. This design enables the student to solve problems through unconstrained RL exploration rather than teacher imitation, while

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Token-Level Entropy Reveals Demographic Disparities in Large Language Models

arXiv:2501.19337v5 Announce Type: replace Abstract: A name alone measurably reshapes a language model's next-token distribution before a single token is sampled. We measure full-vocabulary Shannon entropy of the next-token distribution across six open-weight model families on 5,760 sentence-completion prompts in which race and gender are signaled only by a first name. Black-associated names co-occur with higher first-token entropy and more diverse continuations than White-associated names -- directionally consistent in all six instruction-tuned models under shared raw-text input, all six base checkpoints, and, for output diversity, five of six models under native chat formatting -- opposite to the homogeneity bias documented under explicit group labels (Lee et al., 2024). The gap persists under tokenization and frequency controls and on a frequency-matched name subset; per-prompt effects are small (d = 0.06-0.16) but uniformly signed (template-level paired d = 0.66-1.08). Gender points

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

OpenForgeRL: Train Harness-native Agents in Any Environment

arXiv:2607.21557v1 Announce Type: cross Abstract: Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. OpenForgeRL achieves this with a lightweight proxy that serves the harness's model calls while recording them as training data for a standard RL codebase (e.g., veRL), and a Kubernetes orchestrator that runs each rollout in its own remote container, together enabling training on any harness in any environment at scale. By decoupling training and inference, OpenForgeRL allows researchers to easily train, study, and improve agents directly in the r

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

The Boundaries of Automation: A Theory of Persistent Human Participation

arXiv:2607.21547v1 Announce Type: cross Abstract: The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with algorithms wherever possible. Implicit in this pursuit is the assumption that humans remain in the loop only because current AI systems are not yet sufficiently capable. This paper challenges that assumption. Rather than asking how far automation can extend, we ask where its conceptual limits lie and argue that human participation may persist even with highly capable AI systems for three distinct reasons. Technical or complementarity grounds arise when humans contribute capabilities or perspectives unavailable to AI. Normative or developmental grounds arise when participation itself is valuable for human agency or learning. Most importantly, emergence grounds arise from target emergence: in some activities, the target is not fully specified in advance but instead emerges through the interaction itself. In these cases, hum

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

arXiv:2607.21535v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in Multi-Token-Prediction (MTP/NEXTN) draft head under the assumption that the draft is negligibly cheap. At million-token context this breaks: an MTP draft head typically runs full attention over the entire KV cache at every draft step, so its read grows linearly with context and comes to dominate the draft cost -- precisely where speculation is most valuable. The effect compounds with draft length (a deep native draft can turn net-negative, slower than no speculation) and sharpens under hybrid/linear-attention targets, where cheaper verification leaves the draft's full-attention read exposed. We apply a StreamingLLM-style sliding window plus attention sink to the draft's attention only (Windowed-MTP), leaving full-attention verification intact. It is training-fr

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

GS-Agent: Creating 4D Physical Worlds With Generative Simulation

arXiv:2607.21522v1 Announce Type: cross Abstract: Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effort to fine-tune materials, motions, and visual fidelity. Recent advances in generative foundation models have sparked interest in learning to generate such 4D worlds from large-scale data; however, existing methods still struggle to ensure physical plausibility and controllability. In this work, we take a different path by leveraging foundation models to construct an agentic system that emulates how humans traditionally create 4D worlds, yet automates the entire process. We present GS-Agent, an end-to-end multi-agent framework that integrates physics engines in the loop to generate realistic, dynamic, and controllable 4D physical worlds from natural language. Inspired by how humans build 4D worlds, GS-Agent decomposes the t

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

arXiv:2607.21482v1 Announce Type: cross Abstract: Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requirements that typically prohibit data transmission to external services. Locally deployable open-weight models offer an alternative since sensitive data never leave the local environment. We introduce an open-source framework for evaluating the efficacy of AI agents powered by open-weight LLMs on one of the most persistent bottlenecks in research on longitudinal population studies: data preparation. The framework comprises: a curated ground-truth dataset (cleaning scripts preparing six sweeps of data from a British cohort study), task definitions encompassing tasks such as category harmonization and multi-wave merging, and automated routines for evaluating the LLM-produced R code and outputted data. We benchmark L

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Error Certificates for KV-Cache Eviction via Randomized Design

arXiv:2607.21475v1 Announce Type: cross Abstract: Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest. We prove that this design cannot know what it destroyed: evicted values can be altered so that everything the serving system retains is unchanged while the true attention-output error grows arbitrarily, so no serving-time estimator of that error is consistent. Randomized eviction restores identifiability. With a Poisson-sampled tail at known inclusion probabilities, one logit offset performs the H\'ajek correction inside the softmax, and a survey-sampling variance estimator over the retained set becomes a per-step error certificate with 0.97 empirical coverage at no accuracy cost. On real workloads we pre-registered seven claims and lost three: question-aware eviction at 25--50\% budgets is nearly free; output log-probability predicts failure better than the certificate; certificate-gated budget escalation adds nothing. What survives

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog

arXiv:2607.21412v1 Announce Type: cross Abstract: Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning, especially in safety-critical or compliance-sensitive domains. Recent neuro-symbolic approaches address this gap by coupling neural models with external symbolic engines, yet most integrations are bespoke and lack a standardized interface for tool-augmented agents. This paper presents Euclid-MCP, an open-source MCP server that provides deterministic logical reasoning via SWI-Prolog. Euclid-MCP introduces Euclid-IR, an engine-agnostic intermediate representation for Horn-clause logic that is human-readable, easy for LLMs to generate, and straightforward to compile into Prolog or alternative backends. The server exposes a compact tool interface that supports a translate-run-inspect-repair loop, enabling LLM clients to delegate inference while retaining full access to proof traces and derivation logs.

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

Training Large Language Models for Self-Explanation Faithfulness

arXiv:2607.21090v1 Announce Type: cross Abstract: We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated reasoning accurately reflects its internal decision-making process. While existing work focuses on evaluating faithfulness or using inference-time prompting frameworks to improve an LLM's self-explanation's tractability, these approaches do not provide a mechanism to directly optimize a model's parameters to generate faithful self-explanations. We bridge this gap by modifying existing faithfulness metrics into an RL training objective. We investigate (1) if models can be trained to accurately detect factors that affect their decisions, and (2) whether RL can directly optimize for the disclosure of these factors thereby improving LLM self-explanations' faithfulness. We experiment with two intervention types: random-word insertions and user-bias insertions, using a per-sample reward derived f

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CL

VibeVoice-ASR-BitNet Technical Report

arXiv:2607.21075v1 Announce Type: cross Abstract: We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization (I8_S) with kernel fusion and SIMD optimization, while the autoregressive language model adopts BitNet-style ternary weights (I2_S). To preserve accuracy under aggressive compression, we employ a progressive quantization-aware training strategy. For inference, we implement custom SIMD kernels and fused operators within the ggml framework targeting both ARM and x86 platforms, achieving real-time recognition with RTF < 1 using as few as 3 CPU threads. VibeVoice-ASR-BitNet is 1.6-2.3x faster than Whisper.cpp at comparable model sizes (~1.6 GB), with only modest accuracy degradation compared to the FP16 baseline.

Source ↗
Showing 9551–9600 of 11035 signals
← Prev Page 192 of 221 Next →