EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18624 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Mon, 20 Jul 2026 00:00:00 -0400
arXiv cs.CY

Complete Trip: A Linked Multimodal Human Mobility Dataset

arXiv:2607.15436v1 Announce Type: new Abstract: Human mobility data have become fundamental to research across transportation, public health, urban science, and disaster resilience. However, existing mobility datasets typically capture only isolated aspects of travel behavior and rarely provide linked multimodal journeys together with network-level route representations and population-level inference. Here we present Complete Trip, a mobility dataset that reconstructs linked multimodal travel behavior from passively collected smartphone location-based services (LBS) data. The first released implementation covers six counties in Utah throughout 2020 and represents journeys across car, bus, rail, and active transportation through a four-stage workflow consisting of trip identification, mode imputation, route reconstruction, and trip linking. Complete Trip preserves journey-level relationships by linking sequential travel segments where multiple segments belong to the same travel episode,

Source ↗
technology Mon, 20 Jul 2026 00:00:00 -0400
arXiv cs.CY

Clinical Audit Logs as Multi-Axial Traces of Care Delivery

arXiv:2607.15397v1 Announce Type: new Abstract: Electronic health record audit logs record timestamped actions through which clinical work is carried out. Generated as operational metadata, they now support research on clinician effort, patient outcomes, care-team coordination, and workflow structure. This Perspective explains that breadth by articulating audit logs as multi-axial event streams and drawing implications for representation learning, evaluation, and governance. Each logged action belongs simultaneously to multiple clinically meaningful relations: a clinician's work, a patient's trajectory, a team's activity, and a recurring workflow. This structure motivates foundation-model pretraining to learn reusable representations over the raw stream. Reading audit logs as multi-axial traces specifies what such representations must preserve, how their value should be tested, and how their use should be governed.

Source ↗
technology Mon, 20 Jul 2026 00:00:00 -0400
arXiv cs.CY

On the Effectiveness of Fact Checking Information from Politically Congruent and Incongruent Large Language Models

arXiv:2607.15364v1 Announce Type: new Abstract: Social media companies have shifted away from human fact-checkers and instead have embedded conversational Large Language Models (LLM) on their platforms. LLM chatbots differ from human fact-checkers in many ways that may shape user responses to corrections. Of particular interest in this study is that LLM chatbots can be ideologically configured via the content emphasized in their responses, the sources cited, and the configured persona. Using data from two within-subjects experiments (n=705), this paper investigates the effectiveness of fact checking information from ideologically configured LLM chatbots. We find that LLM fact-checkers significantly shift trust in true and false political news headlines, even when the chatbot is politically incongruent with the user. The perceived political congruency between the participant and the bot matters only when headlines are politically distant. That is, trust in correctly labeled true headlin

Source ↗
technology Mon, 18 May 2026 09:00:00 +0000
Tech & Learning

Elevating the Classroom: A Blueprint for Strategic AI Integration and Student Empowerment

Practical advice for district leaders implementing AI in their district.

Source ↗
technology Mon, 18 May 2026 09:00:00 +0000
Tech & Learning

What is Edcafe AI and How Can I Use It To Teach?

Edcafe AI is an eduction specific tool designed to help along the entire teaching cycle.

Source ↗
technology Mon, 17 Aug 2026 21:59:47 +0000
MedCity News

KFF: Insurers Denied 12%-18% of Prior Authorization Requests in 2025

KFF found that insurers denied 12%-18% of standard prior authorization requests in 2025, with significant variation among insurers and gaps in transparency. The post KFF: Insurers Denied 12%-18% of Prior Authorization Requests in 2025 appeared first on MedCity News .

Source ↗
technology Mon, 17 Aug 2026 20:13:47 +0000
MedCity News

Slate Medicines’ Merger and $245M Private Placement Fuel Mission in Migraine

Migraine drugs developer Slate Medicines is going public in a reverse merger with Fulcrum Therapeutics. The biotech’s lead antibody drug blocks two novel migraine targets, offering the potential for better efficacy compared to other next-generation migraine drugs in R&D. The post Slate Medicines’ Merger and $245M Private Placement Fuel Mission in Migraine appeared first on MedCity News .

Source ↗
technology Mon, 17 Aug 2026 17:30:00 +0000
MedCity News

The Rise of Consumer Diagnostics and Emerging Trends

Consumer diagnostics will be one of several discussions at the INVEST Digital Health conference focused on power shifting to the consumer in healthcare. The conference is scheduled for October 29 in Dallas at Pegasus Park, in partnership with Health Wildcatters. The post The Rise of Consumer Diagnostics and Emerging Trends appeared first on MedCity News .

Source ↗
technology Mon, 17 Aug 2026 16:02:54 -0400
EdTech Mag (Higher)

What the Workforce Pell Grant Launch Means for Community College IT Leaders

In 2025, the Workforce Pell Grant was announced as part of the Working Families Tax Cuts Act. On August 4, the first program was approved for the new Workforce Pell Grant, a milestone in higher education. To be eligible, institutions must prove a 70% completion rate and 70% job placement rate — in the industry or field of the program — within 180 days. Programs must also run between eight to 15 weeks in length. “If an institution is starting now, they’re six months too late — minimum,” says Wesley Bange, a chief information and technology officer at Bossier Parish Community College. Meeting…

Source ↗
technology Mon, 17 Aug 2026 14:21:35 -0400
EdTech Mag (K-12)

How Districts Can Rethink Cyber Risk in Terms of Financial Exposure

Quantifying cyber risk in terms of financial exposure helps K–12 IT teams justify their spending and make the case that security is an enterprise-critical endeavor. This approach pushes districts to audit existing security tools, prioritize threats based on potential financial impact and measure success in dollar-based risk exposure rather than technical metrics. Here’s how a quantified risk analysis can help K–12 security leaders identify which investments meaningfully reduce exposure and which add cost and complexity without real value. Click the banner below to learn how a risk…

Source ↗
technology Mon, 17 Aug 2026 13:45:00 +0000
MedCity News

Healthcare Is Deploying AI Tools — It’s Not Ready for AI Colleagues

Most healthcare organizations are experimenting with AI. Few are preparing to manage AI agents as participants in everyday healthcare workflows. The post Healthcare Is Deploying AI Tools — It’s Not Ready for AI Colleagues appeared first on MedCity News .

Source ↗
technology Mon, 17 Aug 2026 11:30:00 +0000
MedCity News

How Payers and Health Plans Think of Identity and the Risks of Fragmented Data

[Sponsored] A recent webinar, sponsored by Verato, offered insights from executives at SCAN Health Plan and the Alliance of Community Health Plans on a wide range of tech challenges in healthcare from fragmented data to technology architecture to patient identity. The post How Payers and Health Plans Think of Identity and the Risks of Fragmented Data appeared first on MedCity News .

Source ↗
technology Mon, 17 Aug 2026 09:00:00 +0000
Tech & Learning

What is Atlas and How Can I Use It To Teach?

Atlas is an AI-powered learning platform built to meet students where they are.

Source ↗
technology Mon, 17 Aug 2026 09:00:00 +0000
Tech & Learning

8 of the Best AI Teaching Tools

Use these best AI teaching tools to educate across subjects with intelligent digital assistance.

Source ↗
technology Mon, 17 Aug 2026 09:00:00 +0000
eCampus News

Schools are building AI rules before they know the destination

America’s schools are moving quickly to respond to the rise of artificial intelligence, with parents, teachers, administrators, and lawmakers working to wrap their arms around what this means for student education. The post Schools are building AI rules before they know the destination appeared first on eCampus News .

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

arXiv:2607.13124v2 Announce Type: replace-cross Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. Two observations trace this gap. First, greedy \textsc{pass}@$1$ nearly vanishes after compression, yet \textsc{pass}@$k$ recovers substantially under repeated sampling: useful generations are demoted, not erased. Second, the recoverable regime fails mainly through suffix repetition. Recovery should therefore train on the compressed model's own on-policy states with dense token-level supervision, which On-Policy Distillation (OPD) provides by reusing the pre-compression model as a frozen teacher. However, long on-policy rollouts spend early recovery budget on low-information repetitive suffixes, delaying loss descent. To mitigate this waste, we propose \textbf{\shortopd}, a short-to-long OP

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

arXiv:2605.29668v2 Announce Type: replace-cross Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-improvement methods accumulate natural-language guidance without checking that each new item preserves previously correct behavior, so a note that fixes one trajectory can silently regress another. We introduce GRASP (Gated Regression-Aware Skill Proposer), which treats agent improvement as a sequence of edits to a bounded skill library, admitting each candidate only if it produces a net improvement on a balanced held-out probe under a hard regression budget. We evaluate GRASP across five base models on two FHIR-based clinical benchmarks, which score procedural reliability against FHIR state rather than clinical correctness or patient outcomes. On MedAgentBench, GRASP lifts gpt-oss-120b from 40.6% to 88.8%, exceeds the strongest of five self-improvement b

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Understanding and Mitigating Over-refusal for Large Language Models via Representation Intervention

arXiv:2511.19009v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate powerful capabilities across various natural language processing tasks,yet their inherent safety vulnerabilities undermine the reliable application of LLMs in real-world scenarios. To enhance LLM safety, various jailbreak defense methods have been proposed to guard against harmful outputs. However, improvements in model safety often come at the cost of severe over-refusal, failing to strike a good balance between safety and usability. This phenomenon is a critical reliability degradation issue in LLM intelligent systems, failing to strike a good balance between safety defense effectiveness and system usability reliability. In this paper, we first analyze the causes of over-refusal from a representation perspective, revealing that LLMs are unable to effectively distinguish between over-refusal samples and malicious samples. Based on this, we propose to mitigate overrefusal by intervening i

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models

arXiv:2510.25577v2 Announce Type: replace-cross Abstract: Recent advances in Speech Foundation Models (SFMs) enable direct processing of raw audio, allowing models to respond to subtle paralinguistic variation. However, how these models interpret non-lexical cues remains largely unstudied. We introduce VQ-Bench, a controlled evaluation suite featuring a parallel dataset of synthesized modal, breathy, creaky, and end-creak phonation types. We evaluate SFM sensitivity through open-ended generation across four ecologically valid domains, alongside speech emotion recognition. Our results reveal performance gaps: while a leading commercial API failed basic biometric sanity checks, other models exhibited systematic shifts in agency, empathy, and leadership based on phonation. Our findings also highlight gender asymmetries in salary and leadership endorsements, demonstrating that SFMs may mirror or amplify human social biases. This work establishes a reproducible framework for ensuring respon

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

BAT: Learning to Reason about Spatial Sounds with Large Language Models

arXiv:2402.01591v4 Announce Type: replace-cross Abstract: Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and dista

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Lost in Historical Time? A Polish History Matura Benchmark for Large Language Models

arXiv:2608.12343v2 Announce Type: replace Abstract: Language models are widely used by students as knowledge sources, yet benchmarks rarely assess their interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exam (Matura) in history - three official papers from 2023-2025, comprising short-answer questions and extended essays - and compare model performance against the human examinee population. Although models score near the ceiling, aggregate scores mask distinct competency profiles: rankings are unstable across task types, source modalities, and geographical scopes, with a consistent penalty for Polish versus Global history content. Qualitative error analysis reveals two recurring failure modes - source decontextualization, when models reason from source content rather than treating it as an object of analysis, and temporal disorientation, when responses are historically misplaced. This study introduces the first LLM history benchmark groun

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration

arXiv:2606.04743v2 Announce Type: replace Abstract: Agents are widely deployed as assistants over documents, tools, and code. However, they typically act only on explicit user requests, which surface only the problems the user has noticed, while many other important problems coexist, hidden in plain sight, within the broader user context, with their total number unknown in advance. We frame this as the task of discovering multiple hidden problems from context, in which coexisting problems should be uncovered, grounded in supporting evidence, and paired with concrete actions. To this end, we introduce TIDE, a template-guided iterative framework with two complementary mechanisms. Specifically, motivated by the observation that single-pass prediction anchors on the most salient cases and yields generic claims, we propose iterative discovery, which surfaces a small batch of candidates per round while conditioning on what has already been found, so subsequent rounds extend coverage; and tho

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

PhoneWorld: Scaling Phone-Use Agent Environments

arXiv:2605.29486v2 Announce Type: replace Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Existing mobile-agent benchmarks have made important progress on evaluation, but they do not by themselves provide a scalable way to construct many new phone-use environments. We present PhoneWorld, a reusable pipeline that converts real GUI trajectories and screenshots into controllable phone-use environments, executable tasks, automatic verifiers, and training rollouts. Rather than hand-building one mobile benchmark at a time, PhoneWorld uses real trajectories to recover which screens matter, how screens connect, which interactions must change environment state, and which user goals admit automatic verification. From these signals, it builds runnable mock Android apps backed by read-only app content and mutable state, then derives executable tasks, rule-based verifiers, and training roll

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Fine-grained Claim-level RAG Benchmark for Law

arXiv:2605.21071v4 Announce Type: replace Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses. In high-stake domains such as law, retrieval-augmented generation (RAG) is commonly used to mitigate hallucinations in generated responses. Nonetheless, prior work shows that RAG systems, whether general-purpose or legal-specific, still hallucinate at varying rates, making fine-grained evaluation essential. Despite the need, existing evaluation frameworks for legal RAG systems lack the granularity required to provide detailed analysis of retrieval and generation performance separately. Moreover, current benchmarks are largely English-only and centered on legal expert queries, overlooking non-expert needs. We introduce ClaimRAG-LAW, a comprehensive dataset for legal RAG that supports French and English, targets both experts and non-experts, and includes diverse quest

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

When Gradient Importance Lies: Adaptive LoRA Rank Allocation Fails Under GRPO

arXiv:2605.07366v2 Announce Type: replace Abstract: Adaptive rank allocation for LoRA - allocating more parameters to important layers and fewer to unimportant ones - consistently improves efficiency under supervised fine-tuning (SFT). We test whether this success transfers to reinforcement learning, specifically Group Relative Policy Optimization (GRPO). Using gradient-magnitude profiling on Qwen 2.5 1.5B with GSM8K, we find that, in our setting, it does not: proportional rank allocation degrades accuracy by 4.5 points compared to uniform allocation (70.0% vs. 74.5%), despite using identical parameter budgets. We identify two mechanisms behind this failure. First, the gradient landscape under GRPO is fundamentally flatter than under SFT: the max-to-min layer importance ratio is only 2.17x, whereas the layer concentration reported by Shi et al. (2024) for SFT (top 30% of layers carrying >80% of the gradient signal) implies a max/min ratio well above 10x. All layers carry meaningful gra

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization

arXiv:2604.26460v2 Announce Type: replace Abstract: Stylistic personalization - making LLMs write in a specific individual's style, rather than merely adapting to task preferences - lacks evaluation grounded in authorship science. We show that grounding evaluation in authorship verification theory transforms what benchmarks can measure. Drawing on three measurement traditions - LUAR (a trained authorship verification model), an LLM-as-judge with decoupled trait matching, and classical function-word stylometrics - we evaluate four inference-time personalization methods across 50 authors and 1,000 generations. The theory-grounded metric (LUAR) provides what ad hoc alternatives cannot: calibrated baselines (human ceiling 0.756, cross-author floor 0.626) that give scores absolute meaning. All methods score below this floor (0.484-0.508), exposing an authorship gap invisible to uncalibrated metrics. The three metrics produce near-zero pairwise correlations (|r| < 0.07), confirming that with

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Learning to Interrupt in Language-based Multi-agent Communication

arXiv:2604.06452v2 Announce Type: replace Abstract: When a colleague starts explaining something you already understand, you interrupt them. This simple act, a listener taking control of the conversation, is natural in human communication but absent in current verbose LLM multi-agent systems. Current approaches address communication efficiency only from the speaker side, compressing messages before they are sent. We flip the perspective: rather than making speakers more concise, we let listeners decide when they have heard enough. We propose a new communication paradigm in which the listener can interrupt the speaker mid-generation. We find that LLMs, given this ability, are overconfident and interrupt too early before receiving sufficient information. This finding motivates HANDRAISER, a learning method that predicts the right moment to interrupt based on estimated future reward and communication cost. We evaluate our framework on three multi-agent tasks: 2-agent text pictionary, 3-ag

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Early Stopping for Large Reasoning Models via Confidence Dynamics

arXiv:2604.04930v2 Announce Type: replace Abstract: Large reasoning models rely on long chain-of-thought generation to solve complex problems, but extended reasoning often incurs substantial computational cost and can even degrade performance due to overthinking. A key challenge is determining when the model should stop reasoning and produce the final answer. In this work, we study the confidence of intermediate answers during reasoning and observe two characteristic behaviors: correct reasoning trajectories often reach high-confidence answers early, while incorrect rollouts tend to produce long, unproductive reasoning traces and exhibit less reliable confidence dynamics. Motivated by these observations, we propose CoDE-Stop (Confidence Dynamics Early Stop), an early stopping method that leverages the dynamics of intermediate answer confidence to decide when to terminate reasoning, requiring no additional training and easily integrating into existing models. We evaluate CoDE-Stop on di

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Adaptive Stopping for Multi-Turn LLM Reasoning

arXiv:2604.01413v3 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly rely on multi-turn reasoning and interaction, such as adaptive retrieval-augmented generation (RAG) and ReAct-style agents, to answer difficult questions. These methods improve accuracy by iteratively retrieving information, reasoning, or acting, but introduce a key challenge: \textbf{When should the model stop?} Existing approaches rely on heuristic stopping rules or fixed turn budgets and provide no formal guarantees that the final prediction still contains the correct answer. This limitation is particularly problematic in high-stakes domains such as finance and healthcare, where unnecessary turns increase cost and latency, while stopping too early risks incorrect decisions. Conformal prediction (CP) provides formal coverage guarantees, but existing LLM-CP methods only apply to a single model output and cannot handle multi-turn pipelines with adaptive stopping. To address this gap, we propos

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

TriageSim: A Conversational Emergency Triage Simulation Framework from Structured Electronic Health Records

arXiv:2603.10035v3 Announce Type: replace Abstract: Research in emergency triage is restricted to structured electronic health records (EHR) due to regulatory constraints on nurse-patient interactions. We introduce TriageSim, a simulation framework for generating persona-conditioned triage conversations from structured records. TriageSim enables multi-turn nurse-patient interactions with explicit control over disfluency and decision behaviour, producing a corpus of ~800 synthetic transcripts and corresponding audio. We use a combination of automated analysis for linguistic, behavioural and acoustic fidelity alongside manual evaluation for medical fidelity using a random subset of 50 conversations. The utility of the generated corpus is examined via conversational triage classification. We observe modest agreement for acuity levels across three modalities: generated synthetic text, ASR transcripts, and direct audio inputs. We provide the code for TriageSim at https://github.com/dipankar

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models

arXiv:2602.09992v2 Announce Type: replace Abstract: Several recent contributions have evaluated the Poverty of the Stimulus Hypothesis (PoSH) using Artificial Neural Networks (ANNs). The results suggest that ANN-based language models can acquire certain structure-dependent generalizations from limited input without structural inductive biases of the type traditionally hypothesized by linguists. However, existing studies have largely focused on individual phenomena and adopted different evaluation protocols, leaving it unclear whether previous findings generalize across phenomena and learning conditions. We introduce \poshbench, a unified benchmark covering four canonical PoS phenomena. Training Transformer, LSTM, and n-gram models, we find that ANN-based models can achieve above-chance generalization from surprisingly limited input (10M words), but they show less efficient learning than children as input scale grows. Moreover, cognitively motivated inductive biases substantially improv

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Seeing is Coding: On the Effectiveness of Vision Language Models in Code Understanding

arXiv:2602.01785v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable success in source code understanding, yet as software systems grow in scale, computational efficiency has become a critical bottleneck. Currently, these models rely on a text-based paradigm that treats source code as a linear sequence of tokens, which leads to a linear increase in context length and associated computational costs. The rapid advancement of Multimodal LLMs (MLLMs) introduces an opportunity to optimize efficiency by representing source code as rendered images. Unlike text, which is difficult to compress without losing semantic meaning, the image modality is inherently suitable for compression. By adjusting resolution, images can be scaled to a fraction of their original token cost while remaining recognizable to vision-capable models. To explore the feasibility of this approach, we conduct the first systematic study on the effectiveness of MLLMs for code understanding

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Decoding Student Minds: Leveraging Conversational Agents for Psychological and Learning Analysis

arXiv:2512.10441v2 Announce Type: replace Abstract: This paper presents a psychologically-aware conversational agent designed to enhance both learning performance and emotional well-being in educational settings. The system combines Large Language Models (LLMs), a knowledge graph-enhanced BERT (KG-BERT), and a bidirectional Long Short-Term Memory (LSTM) network with attention to classify students' cognitive and affective states in real time. Unlike prior chatbots limited to either tutoring or affective support, our approach leverages multimodal data-including textual semantics, prosodic speech features, and temporal behavioral trends-to infer engagement, stress, and conceptual understanding. A pilot study with 45 university students demonstrated improved motivation, reduced stress, and moderate academic gains compared to unimodal baselines. We explicitly discuss the exploratory nature of this small-sample pilot, report effect sizes and inter-rater reliability alongside significance tes

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Research-Oriented Human-Centric Evaluation for Foundation Models

arXiv:2506.01793v2 Announce Type: replace Abstract: Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in human-AI collaboration. To address this gap, we propose a research-oriented Human-Centric Evaluation framework. It captures user perceptions across three core dimensions: problem-solving ability, information quality, and interaction experience, providing a structured, fine-grained approach to understanding how users evaluate and respond to model behavior in multi-modal research contexts. We conduct 604 human evaluation sessions across various disciplines, involving recent advanced foundation models. Through open-ended, time-limited collaborative tasks, we gather rich subjective assessments that highlight model capabilities and user preferences. Additionally, we perform an LLM-as-a-judge experiment and find that even sophisticated models struggle to accurately

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Leveraging Few-Shot Learning and Large Language Models for Analyzing Blood Pressure Variations Across Biological Sex from Scientific Literature

arXiv:2402.01826v2 Announce Type: replace Abstract: Current blood pressure (BP) technologies and standards were established decades ago, and these standards are still used worldwide today, often without adjusting BP readings for individual demographic factors such as sex and age. While these standards provide useful guidelines and help identify at-risk patients, they are not fully reliable for diagnosis due to the lack of demographic considerations. This study aims to assess the feasibility of using large language models (LLMs) for the automated extraction of BP-related information from the scientific literature, with a focus on biological sex-based distinctions in BP distributions. We employed natural language processing (NLP) methods to extract the means and standard deviations of BP values from the literature, distinguishing by biological sex. We developed a Solr-based search engine to retrieve scientific articles containing BP-related keywords and biological sex indicators from Pub

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

arXiv:2608.14509v1 Announce Type: cross Abstract: Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. This conflates two operations with different requirements. Interpreting a source rewards capacity and context. Combining interpretations rewards fixed arithmetic, comparability across instances, and the option to return nothing. Once separated, the design problem becomes the interface between them. We propose a four-field evidence tuple (hypothesis, reliability bucket, rationale, provenance) and show that fixing it determines both halves. The separation also reveals a failure mode in how such systems combine, which we call count-scale drift. Thresholding a sum of unnormalized weights is exactly posterior thresholding, but at an operating point that slides with the number of sources consulted. The slide grows with reader reliability. When source reliabilities differ, the vote rule and the posterior order instances different

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

LLMs Don't Pay for the Jump

arXiv:2608.14397v1 Announce Type: cross Abstract: Zahavy [2026] argues that Large Language Models, despite their capabilities in induction and deduction, cannot perform the abductive "Jump" that produced Einstein's equivalence principle, and attributes this limitation to the absence of embodied simulation. Zheng-Xin [2026] and Farmer [2026] question whether embodiment is necessary for abduction, pointing to alternative routes to General Relativity and forms of abduction that require no sensorimotor grounding. Max Planck resolved the blackbody radiation problem in 1900. Planck's move to E = h{\nu} required no embodied simulation. It was motivated by a mathematical consequence of classical theory, an infinite predicted energy for a finite measured quantity, that could not be physically accepted. We show that neither induction nor deduction could have produced the postulate and argue that its adoption required a coupling between epistemic error and physical cost. We formalize this distinc

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

arXiv:2608.14375v1 Announce Type: cross Abstract: Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likely to be correct is also worth keeping. Yet a wrong answer can contain a useful decomposition, constraint, or scientific principle. We test this distinction with Diverse Hypothesis Deliberation (DHD), a controlled measurement protocol that caches five independently generated messages and replays the same downstream solver, called the integrator, with each message available or hidden. The replay comparison measures a message's trajectory value: whether making the message available helps or harms subsequent reasoning. Across five mathematics and science benchmarks and two openly available model families, gpt-oss-120b and gemma-4-31B-it, wrong-helpful messages appear in every benchmark-model combination. Among wrong-answer messages that change final correctness, m

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

arXiv:2608.14320v1 Announce Type: cross Abstract: The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large language models (LLMs) exhibit similar behavior. However, existing work on anchoring in LLMs typically evaluates only a narrow set of anchor pathways and rarely distinguishes irrelevant from plausible anchors. We introduce AnchorBench, a benchmark for the anchoring effect in LLMs that evaluates multiple anchor pathways under an explicit anchor relevance axis. Across fourteen models, including ten open-weight models and four frontier API models, and a large set of controlled prompts, we find that (1) anchoring is strongly pathway-dependent, (2) plausible anchors usually induce larger shifts than irrelevant ones when introduced through stronger pathways, (3) anchor influence generally weakens as the anchor moves farther

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Seeing Red, Thinking Bad: Color Bias in Vision Language Models

arXiv:2608.14286v1 Announce Type: cross Abstract: Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In this work, we study how VLMs interpret text rendered as an image, and investigate the influence of visual styling biases. To this end, we introduce Stealth Visual Prompts, which subtly change visual styling of text, such as color and contrast, while preserving semantic content. Using these prompts, we systematically control the visual styling of words in text and measure their impact on the analysis performed by VLMs. We further analyze how such visual perturbations affect the latent representations of the vision encoder. From our experiments, we observed that coloring positive words in green consistently shifts sentiment predictions toward a positive direction. As a result, VLMs often fail to properly account for

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models

arXiv:2608.14252v1 Announce Type: cross Abstract: Recent work suggests that some large language model representations have content or reference. Grounding can secure either without supplying live routes for correction. This paper asks what follows from that gap. An output is answerable when discrepancies can affect what a target- and task-specific arrangement produces, accepts, or withdraws. The arrangement has corrective control only when live, sufficiently independent routes can detect and repair fresh discrepancies. A route profile records which routes constrain the arrangement and how they are related. Those profiles support analysis of truth-tracking: patterned support for representational success. Language models are the pressure case; text-only arrangements provide a task-relative limiting case. Text-trained models inherit patterns of testimony, coherence, and prior correction. Where target-sensitive correction survives training, these can supply derivative answerability (inheri

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement

arXiv:2608.14221v1 Announce Type: cross Abstract: Autoformalization is commonly framed as translating natural-language mathematical statements into machine-verifiable formal languages such as Lean 4. However, faithful formalization requires more than translation. Models must map mathematical concepts to the complex hierarchy of types and definitions in formal libraries such as Mathlib, while ensuring that generated statements preserve the meaning of the source propositions. Existing approaches struggle because they rely heavily on the model's parametric memory for library-specific knowledge, while common data construction pipelines often resort to filtering single-pass outputs and lack mechanisms for feedback-driven revision. To address these challenges, we introduce MathForm, an autoformalization framework for constructing verified training data through Mathlib knowledge retrieval and verification-guided iterative refinement. Before generation, a retrieval planner gathers relevant def

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

MINT: A Universal Zero-Shot Predictor for Transaction Data

arXiv:2608.14198v1 Announce Type: cross Abstract: Banks analyse sequential financial transaction data to perform many tasks, including fraud prevention, credit risk assessment and offer personalization. To improve the predictive accuracy of these tasks, Payments Foundation Models encode transaction sequence data as rich contextual embeddings, which can then be provided to task-specific models as features. However, these Foundation Models are not designed for flexible zero-shot reasoning across novel downstream prediction tasks, limiting their adaptability and utility. Existing LLM-based approaches to zero-shot prediction often fail to fully exploit the predictive signal within transaction data, while relying on costly text serialization or task-specific architectures that scale poorly. To address these limitations, we present the Multimodal Instruction Network for Transactions (MINT), a framework that connects a pretrained transaction sequence encoder to a decoder-only LLM through ligh

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

KV Cache Compression Through the Lens of Transform Coding

arXiv:2608.14191v1 Announce Type: cross Abstract: The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference. Existing quantization methods address this bottleneck by representing the KV cache uniformly with lower-precision data types and designing quantization schemes to minimize reconstruction error in the cache itself, without accounting for how that error propagates through attention mechanisms. We prove that, under a white-noise quantization model, the expected attention-aware distortion decomposes into additive key and value contributions that factor across tokens and channels. Building on transform coding and reverse water-filling, which are classical tools from signal processing and rate-distortion theory, we introduce Attention-Aware Transform Coding (AATC), which allocates bits over a calibration set to minimize attention-aware distortion. On Llama-3.1-8B-Instruct and Qwen-2.5-7B-Instruct, evaluated across LongBench

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers

arXiv:2608.14089v1 Announce Type: cross Abstract: Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves. We present Regime-Conditional Verification (RCV), a lightweight wrapper that adapts an off-the-shelf safety classifier without retraining it. RCV estimates, from the classifier's internal representations, the probability that each prediction disagrees with the deployer's policy, and selectively corrects predictions likely to be wrong. The same correctness estimates also provide a label-free signal for detecting distribution shift, enabling a maintenance loop that updates the correctness estimation layer and resorts to classifier fine-tuning only when necessary. Across three off-the-shelf safety classifiers and two benchmark datasets, RCV improves adherence to the deployer's policy in every clas

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

arXiv:2608.13966v1 Announce Type: cross Abstract: As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality. However, QAT computes the loss and surrogate gradients using a lossy reconstruction of latent full-precision weights, while applying updates to the latent weights themselves. This mismatch can lead to suboptimal training trajectories and a higher loss floor. Second-order PTQ methods mitigate a similar gap by minimizing loss-aware reconstruction error, but doing it once for a frozen model can take hours; repeating this process throughout QAT as the weights evolve is impractical. We introduce QUASAR, a QAT method that continuously performs lightweight, loss-aware reconstruction in the training loop to lower the loss floor and improve the resulting low-bit model. At each training step, QUASAR uses the exponential moving average of squ

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact

arXiv:2608.13926v1 Announce Type: cross Abstract: Large language models have made natural language interfaces to databases (NLIDB) newly credible, but LLM text-to-SQL systems fail in a way that matters for deployment: a hallucinated column or a mis-aggregated total yields a fluent wrong answer, indistinguishable at the point of use from a right one. Where the consumer cannot inspect the generated query, as in enterprise AI deployments and operational dashboards, and increasingly where the consumer is a tool-using agent rather than a person, accuracy alone is insufficient: nothing marks which answers to distrust. This is a reliability problem before it is an accuracy problem. We propose an architectural pattern for such systems, a trusted kernel with a generative shell, resting on one invariant: a component that can fabricate may influence which question the system answers, never which value it returns. A generative shell interprets underspecified input and phrases replies; a determinis

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

arXiv:2608.13925v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass. However, existing dLLMs can suffer from unreliable predictions in early denoising stages under aggressive parallelism strategies, leading to errors that can propagate to later stages. To tackle this issue, we present Consistency Forcing (CForce) for dLLMs, a distillation method to force the mask predictions of early stages to align with those of later stages. CForce trains the model on pre-collected self-rollout trajectories, thereby improving training-inference alignment. We introduce Confidence Adaptive KL Divergence as a distillation objective to conjoin the merits of forward and reverse KL. We further provide a theoretical analysis for the consistency objective to explain why CForce can approximately minimize the prediction error of early stages. Critically, the same formulation applies to both mask-to-token

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Agentic Transaction: Towards ACID-Compliant Agent Systems

arXiv:2608.13900v1 Announce Type: cross Abstract: Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipulation. As agents increasingly operate over persistent environments and multi-step workflows, they face challenges analogous to those addressed by transactional database systems: reliable execution, consistent outcomes, safe concurrency, and durable state management. We introduce the concept of an agentic transaction and propose an ACID-compliant agent system framework that reinterprets the classical ACID properties for agent execution through four semantic guarantees: Semantic Atomicity, Semantic Consistency, Semantic Isolation, and Semantic Durability. Together, these properties provide a principled foundation for building reliable agent systems despite model uncertainty and dynamic execution environments. To instantiate this framework, w

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification

arXiv:2608.13866v1 Announce Type: cross Abstract: Large language models (LLMs) can generate synthetic training data for text classification, but the quality of generated samples is heterogeneous: some fall in correct class regions of the embedding space while others land in peripheral or cross-class zones. We propose a geometric filtering framework that evaluates each LLM-generated sample by its Euclidean distance to real class examples in a sentence embedding space, selecting only geometrically consistent candidates. A soft weighting mechanism transforms filter scores into sample weights for classifier training. Evaluated across 13 datasets, 5 classifiers, 10 augmentation methods, and over 6,700 configurations, our method achieves +2.61 percentage points (pp) over SMOTE ($p<0.0001$, Cohen's $d=0.95$, 88.9% win rate). The approach generalizes to named entity recognition (+9.26pp, 100% win rate) without filter modification, and is robust across 5 LLMs from 4 providers. A key finding is

Source ↗
Showing 8201–8250 of 11029 signals
← Prev Page 165 of 221 Next →