EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18624 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

The Benefit of Collective Intelligence in Community-Based Content Moderation is Limited by Overt Political Signalling

arXiv:2601.22201v2 Announce Type: replace-cross Abstract: Social media platforms face increasing scrutiny over the rapid spread of misinformation. In response, many have adopted community-based content moderation systems, including Community Notes (formerly Birdwatch) on X (formerly Twitter), Community Notes on Meta, and Footnotes on TikTok. However, research shows that the current design of these systems can allow political biases to influence both the development of notes and the rating processes, reducing their overall effectiveness. We hypothesise that enabling users to collaborate on writing notes, rather than relying solely on individually authored notes, can enhance the overall quality of their notes. To test this idea, we conducted an online experiment in which participants jointly authored notes on politically misleading posts. We find that collaboration improves the helpfulness of notes, although the average effect depends on the interactional context. In particular, the bene

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification

arXiv:2511.03217v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information. At the same time, knowledge-graph-based fact-checkers deliver precise and interpretable evidence, yet suffer from limited coverage or latency. By integrating LLMs with knowledge graphs and real-time search agents, we introduce a hybrid fact-checking approach that leverages the individual strengths of each component. Our system comprises three autonomous steps: 1) a Knowledge Graph (KG) Retrieval for rapid one-hop lookups in DBpedia, 2) an LM-based classification guided by a task-specific labeling prompt, producing outputs with internal rule-based logic, and 3) a Web Search Agent invoked only when KG coverage is insufficient. Our pipeline achieves an F1 score of 0.93 on the FEVER benchmark on the Supported/Refuted split without task-specific fine-tuning. To address Not enough information cases, we conduct a

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Stories and Systems: Educational Interactive Storytelling to Teach Media Literacy and Systemic Thinking

arXiv:2508.11059v2 Announce Type: replace-cross Abstract: This paper explores how Interactive Digital Narratives (IDNs) can support learners in developing the critical literacies needed to address complex societal challenges, so-called wicked problems, such as climate change, pandemics, and social inequality. While digital technologies offer broad access to narratives and data, they also contribute to misinformation and the oversimplification of interconnected issues. IDNs enable learners to navigate nonlinear, interactive stories, fostering deeper understanding and engagement. We introduce Systemic Learning IDNs: interactive narrative experiences explicitly designed to help learners explore and reflect on complex systems and interdependencies. To guide their creation and use, we propose the CLASS framework, a structured model that integrates systems thinking, design thinking, and storytelling. This transdisciplinary approach supports learners in developing curiosity, critical thinking

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Is This AI? Longitudinal Analysis of Strategies Used for AI Detection on Two Subreddits

arXiv:2606.22689v2 Announce Type: replace Abstract: As AI-generated content (e.g., "slop") becomes more prevalent online, people are developing strategies to attempt to identify it (or, conversely, to gain confidence that something is not AI-generated). What strategies are people using, and how are they changing over time as generative AI models themselves change? In this work, we catalog and analyze 2 years and 8 months of the AI detection strategies discussed by users of two popular Reddit communities (r/isthisAI and r/RealOrAI) that use the wisdom of crowds to identify AI-generated media. Through a mixed-method analysis of 13,098 posts and 222,060 comments within these communities, we catalog and analyze the prevalence of 12 AI-detection strategies, including examining fine-grained physical details, recognizing trends in AI-created content, and the assumptions people make about what models are capable of producing. Furthermore, we find that these strategies and mental models shift o

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

arXiv:2604.24155v3 Announce Type: replace Abstract: The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Studies of agent-type value forks challenge this assumption by showing that people do not always judge humans and AI systems identically.This paper extends that challenge by examining two further possibilities: first, that evaluations of AI behavior change when its human origins are made visible; and second, that people judge the humans who program AI systems differently from either the machines or the human actors they are compared against. An experiment with 1,002 U.S. adults measured moral judgments in a runaway mine train scenario, varying the subject of evaluation across four conditions: a repairman, a repair robot, a repair robot programmed by company engineers, and compan

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Psychometric Comparability of LLM-Based Digital Twins

arXiv:2601.14264v2 Announce Type: replace Abstract: Large language models (LLMs) act as digital twins for human respondents, yet their psychometric comparability remains uncertain. We propose a construct validity framework spanning construct representation and the nomothetic span, benchmarking models against human gold standards. Across studies, digital twins achieved high aggregate-level accuracy and profile correlations, but showed attenuated item-level correlations. In word association tests, LLM networks exhibited humanlike small-world structure and theory-consistent communities, yet diverged lexically and in local structure. In decision-making and contextualized tasks, they under-reproduced heuristic biases, demonstrating normative rationality, compressed variance, and limited temporal sensitivity. Feature-rich and trait relevant conditioning improved Big Five personality prediction and nomothetic-span alignment, but network invariance remained limited, with partial configural sol

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Towards Automating Scientific Review with Google's Paper Assistant Tool

arXiv:2606.28277v1 Announce Type: cross Abstract: Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving. However, this rapid acceleration is creating a systemic challenge: traditional human peer review cannot scale to match the influx of AI-assisted science. Ultimately, to resolve this tension, we must also deploy AI to accelerate the verification and review process itself. To frame the discussion around this transition, we propose a taxonomy consisting of four progressive levels of AI-human collaboration in scientific evaluation, and discuss various trade-offs involved with each. As a step toward this future, we introduce the Paper Assistant Tool (PAT), an agentic AI framework built for deep scientific review and verification. PAT ingests full scientific manuscripts and produces a comprehensive evaluation, checking theoretical results, validating experiments, suggesting improvements,

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction

arXiv:2606.28186v1 Announce Type: cross Abstract: Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on costly human calibration or item-level textual representations, providing limited evidence about the cognitive processes that make items difficult. We argue that difficulty should be viewed not only as a property of item text, but also as an observable consequence of the problem-solving burden an item induces. Large Reasoning Models (LRMs) offer scalable process evidence through reasoning traces, but such evidence must be structured to support interpretable modeling. To this end, we introduce Epi2Diff (Episode to Difficulty), a framework that maps LRM reasoning traces into cognitively grounded episode sequences. These episodes group trace segments into functional problem-solving states, enabling difficulty to be modeled through reasoning scale, effort allocatio

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

AI Persuasive Framing in Collective Dilemmas

arXiv:2606.27951v1 Announce Type: new Abstract: AI agents are promising tools that can act as flexible behavioral nudges to enhance human cooperation in addressing large-scale societal problems. However, evidence on whether AI agents can effectively boost cooperation remains mixed. We recruited 1,283 participants to play iterated Collective Risk Games in small groups, testing whether AI assistants could nudge participants toward cooperation. By using persuasive framing personalized to each player's Social Value Orientation profile, the AI interventions significantly increased contributions and group success rates. These cooperative effects were short-lived, however, fading after the first few rounds. Strikingly, when the AI treatments were reconfigured to promote selfish behavior through exculpatory framing, the negative effects on contributions and group success were larger and substantially more persistent, particularly for personalized interventions. This asymmetry between prosocial

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

When AI Deceives: A Natural Experiment on the Causal Effects of Perceived Deception on Player Ratings in RPGs

arXiv:2606.27689v1 Announce Type: new Abstract: AI-driven deception mechanisms are increasingly prevalent in digital games, yet the direction and magnitude of their effects on player experience remain contested. Existing research has not sufficiently disentangled designer-intended deception intensity from players' actual perception of deception, and most prior work relies on low-ecological-validity experiments or cross-sectional surveys. The present study aims to independently examine the causal effects of design deception intensity (DDI) and player deception awareness (PDA) on player ratings within a naturalistic gaming environment, and to investigate the moderating role of player experience. Leveraging the 54 version updates of Baldur's Gate 3 between 2019 and 2025 as a quasi-natural experiment, it collected all English-language Steam reviews posted within 1 to 28 days following each update, and constructed a player-version two-way fixed effects panel dataset. DDI was coded by human

Source ↗
technology Mon, 27 Jul 2026 23:25:04 +0000
HN: education

Why communication might be the missing link in engineering education

Article URL: https://digilent.com/blog/why-communication-might-be-the-missing-link-in-engineering-education/ Comments URL: https://news.ycombinator.com/item?id=49076962 Points: 3 # Comments: 0

Source ↗
technology Mon, 27 Jul 2026 23:19:44 +0000
MedCity News

Why Endeavor Health Refuses to Silo Its Value-Based Care Strategy

Lakshmi Halasyamani, chief clinical officer at Endeavor Health, thinks value-based care isn’t really about payment models. To her, this work centers on understanding patients’ clinical, social and financial circumstances well enough to make care plans that actually work for their lives. The post Why Endeavor Health Refuses to Silo Its Value-Based Care Strategy appeared first on MedCity News .

Source ↗
technology Mon, 27 Jul 2026 21:16:49 +0000
MedCity News

Argenx’s $2B Forte Bio Acquisition Brings Autoimmune Playbook to Prevalent Disorders

Argenx’s first acquisition is the buyout of Forte Biosciences, a company with a lead drug that has early clinical validation in celiac disease and vitiligo. Argenx executives say this antibody complements its blockbuster product Vyvgart, offering a different mechanism of action but similar pipeline-in-a-product potential. The post Argenx’s $2B Forte Bio Acquisition Brings Autoimmune Playbook to Prevalent Disorders appeared first on MedCity News .

Source ↗
technology Mon, 27 Jul 2026 18:03:34 +0000
HN: education

To prevent LLMs from destroying education, the work must happen in class

Article URL: https://blainehansen.me/post/learning-is-for-students-not-llms/ Comments URL: https://news.ycombinator.com/item?id=49073349 Points: 7 # Comments: 1

Source ↗
technology Mon, 27 Jul 2026 16:22:47 -0400
EdTech Mag (Higher)

Review: Canva for Campus Provides a Shared Platform for Visual Communication

Universities produce an enormous amount of visual content. Faculty members create lecture presentations and course materials. Marketing teams produce event promotions and recruitment campaigns. Students generate slides, videos and graphics for coursework. But keeping all of those efforts coordinated while maintaining quality and brand consistency can be difficult. Canva for Campus is designed to simplify that process by providing a shared platform where students, faculty and staff can design, collaborate and publish visual materials without needing specialized design software. Canva for…

Source ↗
technology Mon, 27 Jul 2026 13:58:00 +0000
MedCity News

Prior Authorization Is Draining Revenue: Why Automation Has Become a Strategic Imperative

When staff are overextended and patients face delays or abandon care, the issue is no longer a narrow administrative burden. It becomes a business, financial, and access challenge that demands a different operating model. The post Prior Authorization Is Draining Revenue: Why Automation Has Become a Strategic Imperative appeared first on MedCity News .

Source ↗
technology Mon, 27 Jul 2026 13:51:00 +0000
MedCity News

How AI Inside Clinical Workflows Is Unlocking Patient Throughput

The question is not whether the technology is impressive. The question is whether it shows up at the right time, with the right information, inside the workflow where the decision is being made. Anything else is just another report. The post How AI Inside Clinical Workflows Is Unlocking Patient Throughput appeared first on MedCity News .

Source ↗
technology Mon, 27 Jul 2026 11:30:00 +0000
MedCity News

AI Readiness Starts with Solving Healthcare’s Data Fragmentation Problem

[Sponsored] On a recent webinar sponsored by Verato, panelists from SCAN Health Group and the Alliance of Community Health Plans discussed how their organizations are meeting the moment. The post AI Readiness Starts with Solving Healthcare’s Data Fragmentation Problem appeared first on MedCity News .

Source ↗
technology Mon, 27 Jul 2026 11:29:25 +0000
MedCity News

What Happens When Patients Trust AI Over Their Doctor?

As patients increasingly bring AI chatbot advice into exam rooms, where does liability land when that advice goes wrong? Health law attorney Meghan O’Connor said the answer still isn’t clear, but silence is the riskiest response for providers. The post What Happens When Patients Trust AI Over Their Doctor? appeared first on MedCity News .

Source ↗
technology Mon, 27 Jul 2026 09:00:00 +0000
Tech & Learning

What is LIFT and How Can I Use It To Teach?

LIFT, or Literacy Intervention for Teens, is built to help older struggling readers succeed.

Source ↗
technology Mon, 27 Jul 2026 09:00:00 +0000
eCampus News

Strengthening campus pathways to student mental health care

Among college students, 37 percent report moderate to severe depressive symptoms and 33 percent report moderate to severe anxiety. More than three-quarters (76 percent) of college students report moderate or high stress levels, and almost one-quarter (20.6 percent) experience significant psychological distress while on campus. The post Strengthening campus pathways to student mental health care appeared first on eCampus News .

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Reliability Scales Inversely: Bigger Language Models Compound Mistakes Faster

arXiv:2607.18292v2 Announce Type: replace-cross Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account -- more data, retrieval, or scale -- misses an auto-regressive risk residual that increases with scale: the model commits to a low-probability token, conditions on it as established, and snowballs. We track this through per-position disagreement $\delta = \log p_M - \log p_O$ against a stronger same-family oracle, whose second moment splits exactly into bias$^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ and risk $\mathrm{Var}[\delta]$. Across three model families, we present four findings: (i) under scaling, the knowledge gap falls up to $7\times$ while knowledge degradation grows up to $39\times$; (ii) at a fabrication, felt uncertainty $H(p_M)$ relaxes quickly while oracle-referenced risk persists up to $23\times$ longer, leaving a confident-but-precarious risk regime that bridges consecutive fabric

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Token-Operations-Oriented Inference Optimization Techniques for Large Models

arXiv:2606.20295v2 Announce Type: replace-cross Abstract: Large model inference optimization serves as a key foundation for supporting the scalable, low-cost, and highly stable operation of large model services. Centered on token-oriented inference optimization technology, this paper proposes for the first time a four-layer technical architecture consisting of Multi-model Fusion, Model Optimization, Compute-Model Fusion, and Compute-Network-Model Fusion. It systematically reviews the key technologies and current industry status across these four levels and analyzes the application value of related technologies in real-world business scenarios. This paper provides a practical technical path for reducing token production costs, improving token service efficiency, ensuring the stability of token supply, and driving the transition of large model services from being merely callable to being operable.

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models

arXiv:2605.17770v4 Announce Type: replace-cross Abstract: The advancement of Large Reasoning Models (LRMs) has catalyzed a paradigm shift from reactive ``fast thinking'' text generation to systematic, step-by-step ``slow thinking'' reasoning, unlocking state-of-the-art performance in complex mathematical and logical tasks. However, the field faces \textit{the fundamental gap between token-level behavioral analysis and internal reasoning mechanisms, and the instability of reinforcement learning (RL) for reasoning optimization relying on costly external verifiers}. We identify and formally define \textbf{Entropy-Gradient Inversion}, a robust negative correlation between token entropy and logit gradients that acts as a definitive geometric fingerprint for LRM reasoning capability. Building on this, we propose \textbf{Correlation-Regularized Group Policy Optimization (CorR-PO)}, which embeds this inversion signature into RL reward regularization. Extensive experiments on various reasoning

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning

arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~(LLMs). However, introducing additional process reward models incurs substantial computational overhead, and there is no unified theoretical framework for process-level advantage estimation. To bridge this gap, we propose \textbf{S}elf-Guided \textbf{P}rocess \textbf{R}eward \textbf{O}ptimization~(\textbf{SPRO}), a novel framework that enables process-aware RL through two key innovations: (1) we show that process rewards can be derived intrinsically from the policy model itself, and (2) we redefine step-wise advantage by introducing well-defined Cumulative Process Rewards~(\textbf{CPR}) and \textbf{M}asked \textbf{S}tep \textbf{A}dvantage~(\textbf{MSA}), which facilitates rigorous step-wise action advantage estimation within shared-prompt sampling groups. Our experimental results show that

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Hint-Guided Diversified Policy Optimization for LLM Reasoning

arXiv:2606.03021v2 Announce Type: replace Abstract: Recent developments in Large Language Models (LLMs) have showcased impressive reasoning capabilities, with Reinforcement Learning with Verifiable Rewards (RLVR) being a promising enhancement strategy. However, existing reward mechanisms are constrained to the outcome-level correctness and lack explicit signals to guide the model to consider diverse solutions. In contrast, human problem solving typically involves evaluating multiple potential approaches and selecting the most reliable solution, a cognitive process that current RLVR frameworks do not explicitly incentivize. Inspired by this, we propose Hint-Guided Diversified Policy Optimization (HDPO), allowing the model to first list all potential candidate solution outlines as hints and then select the most reliable one for further reasoning. HDPO comprises two stages of Cold Start for Structured Reasoning and Hint-Guided Diversified Reinforcement Learning to incentivize the model to

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation

arXiv:2605.03534v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer. We study evidence sufficiency verification for selective RAG answering, in which a verifier receives a question, a candidate answer, and retrieved evidence and decides whether the evidence supports, refutes, or is insufficient for the answer, answering only when support is established. We present SURE-RAG, an aggregation protocol that treats evidence sufficiency as a set-level property: missing hops and unresolved conflicts cannot be detected by scoring passages independently. A shared claim-evidence verifier produces a local relation distribution for each (claim, passage) pair, which SURE-RAG aggregates into four interpretable answer-level feature blocks (coverage, relation strength, uncertainty, and retrieval), producing a three-way decision and an auditable

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision

arXiv:2603.07025v2 Announce Type: replace Abstract: Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficult to train with supervised fine-tuning, requiring large, task-specific speech corpora. While recent distillation-based approaches train performant English-only Speech LLMs using only annotated ASR data by aligning text and speech using only a lightweight projector, these models under-perform when scaled to multilingual settings due to language interference in the shared projector. We address this by introducing language-aware distillation using a query bank and a gating network that selects or mixes query tokens using a Q-Former projector. Our approach shows gains of 14% over matched multilingual distillation baselines on instruction following. We further synthesize Audio-MLQA, a multilingual spoken QA benchmark built on MLQA with high-quality TTS questions. Our best model improves over e

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

SwiftMem: Fast Agentic Memory via Query-aware Indexing

arXiv:2601.08160v2 Announce Type: replace Abstract: Agentic memory systems have become critical for enabling LLM agents to maintain long-term context and retrieve relevant information efficiently. However, existing memory frameworks often perform query-agnostic retrieval over the full memory embedding space even when their storage layer is backed by efficient vector indexes such as HNSW. This full-scope retrieval path creates latency bottlenecks as memory grows, hindering real-time agent interactions. We propose SwiftMem, a query-aware agentic memory system that narrows retrieval to query-relevant memory subsets through specialized indexing over temporal and semantic dimensions. Our temporal index enables logarithmic-time range queries for time-sensitive retrieval, while the semantic DAG-Tag index maps queries to relevant topics through hierarchical tag structures. To address memory fragmentation during growth, we introduce an embedding-tag co-consolidation mechanism that reorganizes s

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

What Matters When Building Universal Multilingual Named Entity Recognition Models?

arXiv:2601.06347v2 Announce Type: replace Abstract: Recent progress in universal multilingual named entity recognition (NER) has been driven by multilingual transformer models, task-specific architectures, custom loss functions, and large-scale training datasets. However, despite substantial prior work, we find that many critical design decisions for such models are made without systematic justification, with individual components evaluated only in combination rather than in isolation. We argue that this lack of rigor impedes progress in the field by making it difficult to identify which choices improve multilingual generalization. In this work, we conduct extensive experiments on transformer backbones, architectures, training objectives, data composition, and threshold selection. Building on these findings, we present Otter, a universal multilingual NER model supporting over 100 languages. Otter achieves consistent improvements over strong multilingual NER baselines, outperforming sim

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning

arXiv:2601.03190v4 Announce Type: replace Abstract: Machine unlearning aims to forget sensitive knowledge from Large Language Models (LLMs) while maintaining general utility. However, existing approaches typically treat all tokens in a response indiscriminately and enforce uncertainty over the entire vocabulary. This global treatment results in unnecessary utility degradation and extends optimization to content-agnostic regions. To address these limitations, we propose PALU (Prefix-Aware Localized Unlearning), a framework driven by a local entropy maximization objective across both temporal and vocabulary dimensions. PALU reveals that (i) suppressing the sensitive prefix alone is sufficient to sever the causal generation link, and (ii) flattening only the top-$k$ logits is adequate to maximize uncertainty in the critical subspace. These findings allow PALU to alleviate redundant optimization across the full vocabulary and parameter space while minimizing collateral damage to general mo

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

InteractComp: Evaluating Search Agents With Ambiguous Queries

arXiv:2510.24668v2 Announce Type: replace Abstract: Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that user queries are complete and unambiguous. This assumption leaves under-tested a practical failure mode: agents may face ambiguous requests where the intended target cannot be identified without clarification. Yet most agents lack interactive mechanisms during the search process, and existing benchmarks cannot assess this capability. To address this gap, we introduce InteractComp, a benchmark designed to evaluate whether search agents can recognize query ambiguity and actively interact to resolve it during search. Following the principle of easy to verify, interact to disambiguate, we construct 210 expert-curated questions across 9 domains through a target-distractor methodology that creates controlled ambiguity resolvable only through interaction. Evaluation of 17 models reveals striking fa

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph

arXiv:2508.12393v3 Announce Type: replace Abstract: The rapid expansion of medical literature challenges the scalable structuring of domain knowledge. Knowledge Graphs (KGs) offer a solution, yet current construction methods lack generalizability and ignore the temporal dynamics of evolving knowledge. To address this, we introduce MedKGent, a Large Language Model (LLM) agent framework for building temporally evolving medical KGs. Using over 10 million PubMed abstracts from 1975 to 2023, MedKGent incrementally constructs a KG daily via two specialized agents. The Extractor Agent identifies knowledge triples and assigns confidence scores, while the Constructor Agent integrates these triples into a temporal graph, reinforcing recurring knowledge and resolving conflicts. The resulting KG contains 156,275 entities and 2,971,384 triples, making it, to our knowledge, the largest LLM-derived medical KG to date. Automated and expert assessments showed triple-validity rates approaching 90%. In d

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Learning to Reason for Factuality

arXiv:2508.05618v2 Announce Type: replace Abstract: Reasoning Large Language Models (R-LLMs) have significantly advanced complex reasoning tasks but often struggle with factuality, generating substantially more hallucinations than their non-reasoning counterparts on long-form factuality benchmarks. However, extending online Reinforcement Learning (RL), a key component in recent R-LLM advancements, to the long-form factuality setting poses several unique challenges due to the lack of reliable verification methods. Previous work has utilized automatic factuality evaluation frameworks such as FActScore to curate preference data in the offline RL setting, yet we find that directly leveraging such methods as the reward in online RL leads to reward hacking in multiple ways, such as producing less detailed or relevant responses. We propose a novel reward function that simultaneously considers the factual precision, response detail level, and answer relevance, and applies online RL to learn hi

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Interpretable Depression Detection from Social Media Text Using LLM-Derived Embeddings

arXiv:2506.06616v2 Announce Type: replace Abstract: Accurate and interpretable detection of depressive language in social media can support early identification of mental health conditions and inform timely interventions. In this paper, we investigate the use of large language models (LLMs) and traditional machine learning classifiers for three social media-based mental health prediction tasks: binary depression classification, depression severity classification, and differential diagnosis among depression, PTSD, and anxiety. We compare zero-shot LLMs with supervised classifiers trained on conventional text embeddings, psycholinguistic features, and embeddings derived from LLM-generated mental health summaries. Across multiple publicly available social media text datasets and five-fold cross-validation experiments, we find that zero-shot LLMs exhibit strong performance and generalization in binary depression classification, but struggle with fine-grained severity prediction. In contras

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization

arXiv:2607.22334v1 Announce Type: cross Abstract: Open-weight language models from different families exhibit complementary capabilities, motivating their consolidation into a compact student through on-policy distillation (OPD). However, full-vocabulary OPD typically assumes a shared tokenizer, while existing cross-tokenizer methods may discard teacher probability mass or assign it to student tokens with unrelated content. We introduce Byte-Prefix Marginalization (BPM), which re-expresses the teacher's next-token distribution over the student vocabulary in a shared byte space. Specifically, BPM assigns each teacher token's probability to the longest student token whose byte representation is a prefix of the teacher token's bytes, aggregates mass mapped to the same student token, and places otherwise unmatched mass in an explicit residual category. This produces a vocabulary-complete, byte-aligned, and mass-preserving target for dense OPD. The target exactly recovers the teacher-induce

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents

arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write interaction with a running database); observation-space scale and complexity (causal diagnosis across thousands of time series, business logs, and concurrent activity); solution-space openness (multiple remediations with different operational trade-offs); and scenario complexity and coverage (faults cascading across internal mechanisms and operational domains). We present DBA-Bench, a benchmark addressing these gaps through production fidelity, outcome-first evaluation, and controlled scenario reproducibility. It uses instrumented PostgreSQL environments with active workloads, persistent state, and multi-source observations; defines success by measurable recovery or fault elimination under safety constraints; and re

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

arXiv:2607.22083v1 Announce Type: cross Abstract: We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diversity of executable environments, task assets, and agentic scaffolds through real-world deployment and large-scale synthesis. Our RL pipeline applies mixed-mode RLHF over Think and Non-Think responses to improve overall model quality and reduce failure cases, length-controlled reasoning RL to balance accuracy and reasoning efficiency, and agentic RL with outcome and process rewards to stabilize long-horizon training. Extensive evaluations show tha

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation

arXiv:2607.22034v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor lighting. In such settings, a reliable uncertainty signal matters more than raw accuracy, because it determines when a system should defer rather than answer. We evaluate two small open-weight VLMs -- Qwen2-VL-2B-Instruct and SmolVLM-Instruct -- across six realistic photographic degradations at three severity levels, comparing two confidence signals: the confidence the model states in natural language, and the model's own mean token probability over its generated answer. Across 3,800 predictions, we find a large and consistent gap. Verbalized confidence in Qwen2-VL is almost constant (mean 0.87-0.90 across all conditions) and detects its own errors at chance level (AUROC 0.39-0.75, typically ~0.50), while internal token probability from the same model separates correct from incorrect answers

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

arXiv:2607.22014v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction. We introduce MissionBench, a benchmark for mission-level evaluation of MLLMs in aerial 3D environments. It comprises 120 missions across five simulated 3D environments and four task families. Agents must autonomously plan, navigate, and report outcomes using only egocentric observations and its action history, without aerial-specific fine-tuning. Across 22 open- and closed-source MLLMs, the strongest model succeeds on fewer than 35% of missions compared to 84.4% human performance, highlighting the difficulty of multi-step embodied tasks. Despite large variations between model families, we observe gains from scaling, indicating that larger general-purpose models possess stronger zero-shot embodied capabilities

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning

arXiv:2607.21971v1 Announce Type: cross Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We hypothesize that the success of such evolution frameworks hinges on meta-skills, such as self-reflection with environment feedback, that enable effective multi-round refinement, yet are largely neglected by traditional post-training. To bridge this gap, we present MetaEvolve, a framework designed to develop these meta-skills via a data synthesis pipeline, evolution-aware reinforcement learning (RL), and inference-time evolutionary search. Concretely, we ground MetaEvolve in coding, where program execution provides natural, continuous reward signals beyond binary correctness. Building on these signals, we synthesize evolution trajectories as training data, each containing a current program, its fitness score (combining correctness and efficiency), and a history of prior attempts, and train t

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Diffusion Models in Medical Image Inpainting: Challenges, Solution Taxonomy, and Future Directions

arXiv:2607.21904v1 Announce Type: cross Abstract: Image inpainting aims to reconstruct missing or corrupted regions of an image while preserving as much as possible, visual and semantic consistency. In medical imaging, this task is particularly important because artifacts, missing information, and pathological alterations can compromise diagnostic reliability and downstream clinical applications. Recently, diffusion models have emerged as state-of-the-art generative approaches for medical image inpainting due to their ability to generate anatomically consistent reconstructions. This survey presents a systematic review of diffusion-based methods for medical image inpainting, covering the main architectures, applications, datasets, and evaluation strategies reported across 60 studies. In addition, we propose a taxonomy for diffusion-based approaches. The analysis reveals a rapid growth of research interest in diffusion-based medical image inpainting, with denoising diffusion probabilisti

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

LeAct: Learning to Reason from Expert Actions

arXiv:2607.21856v1 Announce Type: cross Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines, classical planners, theorem provers), which routinely produce near-optimal actions across diverse domains. But these experts are silent: they commit to an action without writing down the chain of thought (CoT) behind it. Recovering that CoT as natural-language reasoning would distill expert knowledge into a student that generalizes beyond the demonstrated actions. We treat it as a latent variable and study how to recover it from the action alone. Our approach, LeAct (Learning to reason from Actions), optimizes this latent variable: the student samples candidate CoTs for each expert action, and we retain those that measurably improve its own probability of recovering the action. Across imperfect-information games at mu

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Adversarial Prompts for Acceptance Collapse in Speculative Decoding

arXiv:2607.21804v1 Announce Type: cross Abstract: Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model. However, this guarantee of semantic equivalence masks a severe operational vulnerability: draft-target alignment can be systematically attacked. In this paper, we introduce ADSD, which, to the best of our knowledge, is the first prompt-suffix attack that collapses verifier acceptance by pushing draft probability mass toward tokens the target is unlikely to accept. ADSD uses Soft-Collapse, a verifier-aligned surrogate derived from the asymmetric speculative acceptance rule, together with a target-preservation objective that discourages obvious task corruption. ADSD successfully generates highly effective adversarial suffixes. On the GSM8K dataset, our attack increases the mean sample time by 62.3% while preserving the task quality. We further show that this vul

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

From Seasonality to Semantics: Benchmarking a Hybrid Probabilistic Forecasting System for Roadblocks in Bolivia

arXiv:2607.21785v1 Announce Type: cross Abstract: Roadblocks in Bolivia are a social conflict phenomenon with devastating economic impacts, estimated at losses equivalent to 4% of the national Gross Domestic Product. Despite their recurrence and impact, there is a lack of local predictive systems to anticipate these events for logistical decision-making. This paper presents a hybrid probabilistic forecasting system that integrates time series decomposition (Prophet) with natural language processing (NLP) techniques applied to a six-year corpus of Bolivian news coverage. The methodology employs vector semantic embeddings and zero-shot classification models to capture signals of discursive escalation prior to the materialization of the roadblocks. Using an expanding walk-forward validation scheme applied over 1,762 days and seven forecasting horizons (H+1 to H+7), seven internal configurations and four external benchmarks were compared, including SARIMA and LightGBM. The results demonstr

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets

arXiv:2607.21692v1 Announce Type: cross Abstract: Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input. These selectors are often trained by distilling the attention patterns of a dense teacher, assuming that attention reveals which context the teacher actually uses. We test that assumption on retrieval tasks where the evidence for each answer is known exactly. By masking parts of the context and measuring whether the answer changes, we find that attention and causal dependence often disagree, and distilled selectors inherit the mismatch. Teachers attend to outdated facts they have learned to ignore, and their attention can vary across training runs even when they rely on the same evidence. In a two-step reference task, attention at the answer skips the intermediate step because it was resolved earlier in the forward pass: a selector trained on attention achieves 41% accuracy, while the same selector trained on causal eviden

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

arXiv:2607.21655v1 Announce Type: cross Abstract: Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing earlier progress. For this reason, recent studies have increasingly explored progress rewards that provide feedback during task execution. However, the current literature lacks a shared framework. Existing methods use different observations, goal specifications, output signals, supervision sources, and evaluation protocols. This makes it difficult to compare them and understand what their results actually validate. In this survey, we provide a unified view of progress reward modeling for robotic learning. We organize the field in three connected steps. We first study the interface of a progress model. This defines the problem from the outside by asking what information the model re

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

arXiv:2607.21653v1 Announce Type: cross Abstract: Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstream frameworks each change threads through layers of trainer, distributed backend, and rollout glue: the cost lands on the researcher at every iteration. Molt is a PyTorch-native training framework built to keep that cost small: a codebase compact and clean enough for a researcher to hold in their head, and for an AI coding assistant to read and reason about in its entirety, so the algorithm flow can be traced and changed end to end. The agent is an ordinary program, and one asynchronous loop trains multimodal and mixture-of-experts policies while never training on a token it did not generate, consistent in tokens, policy versions, and model semantics. Leanness does not cost performance: under a matched, fully asynchronous protocol, Molt is statistically comparable to a state-of-the-art Mega

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models

arXiv:2607.21617v1 Announce Type: cross Abstract: Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful transcribers: when text is imperfect, they often tend to rewrite it into a more plausible form - a behavior that clean-text OCR benchmarks cannot detect. We introduce FaithC4, a multilingual perturbation benchmark of 1,455 single-page documents (English, Chinese, Korean) with three perturbation families: scramble, random substitution, and visually similar substitution. We use the benchmark to evaluate 15 systems spanning general-purpose VLMs, OCR-specialized VLMs, and traditional OCR pipelines. These three categories differ in WER degradation under perturbation: general-purpose VLMs degrade by up to 4.5 points, OCR-specialized VLMs by 0.2-2 points, and traditional OCR by less than 0.6 points on English. Probing Qwen3-VL-4B layer-by-layer, we identify a consistent

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CL

The Hard Decision Layer: Evidence for Committed Inference in Transformers

arXiv:2607.21613v1 Announce Type: cross Abstract: We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Decision Layer_ (HDL), a natural architectural property where answer option rankings stabilize abruptly during inference. Empirical validation across four language models (Qwen, Llama, Granite, Mistral) and four benchmark datasets demonstrates consistent HDL emergence without learned routing policies. We also show that the HDL is invariant to fine-tuning. Our results reveal striking accuracy improvements at the HDL: up to +0.61 (Qwen on CommonsenseQA), after which performance stabilizes. Systematic ablations on label formats and problem complexity confirm the phenomenon is fundamental to model architecture. These findings offer mechanistic insights into transformer inference and suggest opportunities for efficient reasoning and model steering. All code and results required to reproduce this wo

Source ↗
Showing 7801–7850 of 11029 signals
← Prev Page 157 of 221 Next →