EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

regulation Mon, 17 Aug 2026 10:30:00 +0000
The 74

Iowa, Deemed Below Basic in Reading, Stays Silent on the ‘Honesty Gap’

Every few years, testing experts at the U.S. Department of Education weigh state assessments against their “gold standard,” the National Assessment of Educational Progress, also known as “the nation’s report card.” In a review released in 2021, two states stood out, but not in a good way. Both Iowa and Virginia established fourth grade reading […]

Source ↗
behavior Mon, 17 Aug 2026 10:00:00 +0000
MindShift (KQED)

Slow Math: Kids May Learn More When AI Makes Them Review Mistakes

A randomized experiment involved more than 6,000 Tennessee middle school students learning fractions.

Source ↗
behavior Mon, 17 Aug 2026 10:00:00 +0000
eSchool News

5 ways AI can strengthen your teaching this school year

A year ago, many educators were still asking whether artificial intelligence belonged in the classroom. Some were cautiously experimenting with AI-generated lesson plans, while others were focused on preventing students from using it altogether.

Source ↗
need Mon, 17 Aug 2026 10:00:00 +0000
Hechinger Report

Slow math: Kids may learn more when AI makes them review mistakes

Artificial intelligence makes many tasks faster and easier. But a new study suggests that students may learn more when AI makes them slow down. In an experiment involving more than 6,000 middle schoolers in Tennessee, students learned slightly more math when an AI tutor walked them through their mistakes and then required them to demonstrate […] The post Slow math: Kids may learn more when AI makes them review mistakes appeared first on The Hechinger Report .

Source ↗
behavior Mon, 17 Aug 2026 09:15:00 +0000
Getting Smart

Coming Full Circle: The Portrait Lives Through Students

The final installment in Norwalk's Portrait of a Graduate series asks the most important question of all: what do students say when you ask them to describe the Portrait? This post makes a compelling case that student co-design is not a nice-to-have but a core implementation lever, and it offers concrete strategies, from design sprints to feedback cycles to student-led reflection, for keeping student voice at the center of the work. For education leaders navigating the gap between vision and practice, this is an essential read. The post Coming Full Circle: The Portrait Lives Through Students appeared first on Getting Smart .

Source ↗
technology Mon, 17 Aug 2026 09:00:00 +0000
Tech & Learning

What is Atlas and How Can I Use It To Teach?

Atlas is an AI-powered learning platform built to meet students where they are.

Source ↗
technology Mon, 17 Aug 2026 09:00:00 +0000
Tech & Learning

8 of the Best AI Teaching Tools

Use these best AI teaching tools to educate across subjects with intelligent digital assistance.

Source ↗
technology Mon, 17 Aug 2026 09:00:00 +0000
eCampus News

Schools are building AI rules before they know the destination

America’s schools are moving quickly to respond to the rise of artificial intelligence, with parents, teachers, administrators, and lawmakers working to wrap their arms around what this means for student education. The post Schools are building AI rules before they know the destination appeared first on eCampus News .

Source ↗
audience Mon, 17 Aug 2026 07:00:00 +0000
Inside Higher Ed

How LGBTQ+ Higher Ed Leaders Are Navigating the Political Moment

How LGBTQ+ Higher Ed Leaders Are Navigating the Political Moment Sara Weissman Mon, 08/17/2026 - 03:00 AM A group for LGBTQ+ higher ed administrators has been hit hard by an increasingly hostile policy landscape. But members are working to support each other. Byline(s) Sara Weissman

Source ↗
audience Mon, 17 Aug 2026 07:00:00 +0000
Inside Higher Ed

Troy Shut Off Professor’s Email After He Contacted State Officials

Troy Shut Off Professor’s Email After He Contacted State Officials Emma Whitford Mon, 08/17/2026 - 03:00 AM Micah Swartz has also received two written disciplinary warnings since he sent a letter to Alabama lawmakers outlining his concerns about an ongoing state syllabi review. Byline(s) Emma Whitford

Source ↗
audience Mon, 17 Aug 2026 07:00:00 +0000
Inside Higher Ed

CUNY’s Computer Science Growing Pains

CUNY’s Computer Science Growing Pains Joshua.Bay Mon, 08/17/2026 - 03:00 AM After a decade-long enrollment boom, CUNY faces faculty shortages and pressure to prepare students in the field for an AI-driven job market. Byline(s) Joshua Bay

Source ↗
audience Mon, 17 Aug 2026 07:00:00 +0000
Inside Higher Ed

The Heated Discourse About Civil Discourse

The Heated Discourse About Civil Discourse Ryan Quinn Mon, 08/17/2026 - 03:00 AM Groups advocating for “dialogue across difference” and finding better ways to discuss and disagree about controversial topics are themselves part of a polarized debate. Byline(s) Ryan Quinn

Source ↗
audience Mon, 17 Aug 2026 07:00:00 +0000
Inside Higher Ed

Foxes in the Higher Ed Henhouse

Foxes in the Higher Ed Henhouse kjohnsonbowles… Mon, 08/17/2026 - 03:00 AM Professional athletics and the entertainment industry are dining and dashing like foxes in a henhouse when it comes to higher education and its students. Byline(s) Kathy Johnson Bowles

Source ↗
audience Mon, 17 Aug 2026 07:00:00 +0000
Inside Higher Ed

The Fatal, Familiar Takedown of Jason Arday

The Fatal, Familiar Takedown of Jason Arday Sara Brady Mon, 08/17/2026 - 03:00 AM I wish the vicious envy and violent despising of Black male intellectual leadership had somehow bypassed him. Byline(s) Shaun Harper

Source ↗
audience Mon, 17 Aug 2026 07:00:00 +0000
Inside Higher Ed

Have You Ever Noticed … ?

Have You Ever Noticed … ? Sara Brady Mon, 08/17/2026 - 03:00 AM Gen ed as cultural sense making. Byline(s) Matt Reed

Source ↗
audience Mon, 17 Aug 2026 07:00:00 +0000
Inside Higher Ed

NIH Proposes Removing Peer Review Scores From Grant Applications

NIH Proposes Removing Peer Review Scores From Grant Applications Susan H. Greenberg Mon, 08/17/2026 - 03:00 AM Byline(s) Susan H. Greenberg

Source ↗
audience Mon, 17 Aug 2026 07:00:00 +0000
Inside Higher Ed

Maestro College Accreditor Discloses Private Student Information

Maestro College Accreditor Discloses Private Student Information Katherine Knott Mon, 08/17/2026 - 03:00 AM Byline(s) Katherine Knott

Source ↗
audience Mon, 17 Aug 2026 07:00:00 +0000
Inside Higher Ed

ED Clarifies How Current Students Can Access Uncapped Loans

ED Clarifies How Current Students Can Access Uncapped Loans jessica.blake@… Mon, 08/17/2026 - 03:00 AM Byline(s) Jessica Blake

Source ↗
need Mon, 17 Aug 2026 05:30:00 +0000
Hechinger Report

Most kids miss recommended screening for autism, speech delays

SYRACUSE, N.Y. — During the first year and a half of her son’s life, Valerie Gregory did everything she could to help him thrive: She read to him, talked to him, showed up for wellness checks, and whenever possible took him to a weekly playgroup, where both mother and son had the chance to socialize […] The post Most kids miss recommended screening for autism, speech delays appeared first on The Hechinger Report .

Source ↗
audience Mon, 17 Aug 2026 05:00:00 -0400
Higher Ed Dive

Beyond the degree: How students are defining value in 2026

New research reveals how student expectations are shifting and how institutions can keep pace.

Source ↗
audience Mon, 17 Aug 2026 05:00:00 -0400
Higher Ed Dive

Week in review: Trump lawsuit against Harvard dismissed

We’re rounding up recent stories, from senators raising alarm about student visa appointment availability to a higher education antitrust case proceeding.

Source ↗
regulation Mon, 17 Aug 2026 05:00:00 -0400
K-12 Dive

How a California school is streamlining its approach to SEL supports

Foussat Language Academy is embracing a multi-tiered social-emotional learning model to improve screening and classroom supports.

Source ↗
regulation Mon, 17 Aug 2026 05:00:00 -0400
K-12 Dive

Week In Review: Unions sue over ‘professional’ degrees

We’re rounding up last week’s news, from the educational impact of extreme weather to chronic absenteeism data.

Source ↗
need Mon, 17 Aug 2026 05:00:00 +0000
Hechinger Report

Once globally preeminent, U.S. universities are sliding into decline

The nation was distracted by other things when a quiet warning sounded: Support was diminishing for American academic research that had led to breakthroughs and innovation. That familiar-sounding worry came not this year, but near the end of World War II, in a pivotal government study by the wartime Office of Scientific Research and Development. […] The post Once globally preeminent, U.S. universities are sliding into decline appeared first on The Hechinger Report .

Source ↗
need Mon, 17 Aug 2026 04:59:00 +0000
Hechinger Report

OPINION: Why investing in rural education can bolster local economies and help students build a future close to home

Morgan County, Tennessee, is a rural community shaped by strong relationships, deep attachment to place and a shared commitment to its future. Its public lands and tourism economy are part of that story, as are the intergenerational ties, civic commitment and local institutions that hold the community together. Like many rural areas, Morgan County also […] The post OPINION: Why investing in rural education can bolster local economies and help students build a future close to home appeared first on The Hechinger Report .

Source ↗
behavior Mon, 17 Aug 2026 00:00:00 GMT
EdSurge

College Admissions Testing is Returning to Mixed Reactions

Supporters argue that reinstating SAT and ACT test requirements better vet student readiness.

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

arXiv:2607.13124v2 Announce Type: replace-cross Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. Two observations trace this gap. First, greedy \textsc{pass}@$1$ nearly vanishes after compression, yet \textsc{pass}@$k$ recovers substantially under repeated sampling: useful generations are demoted, not erased. Second, the recoverable regime fails mainly through suffix repetition. Recovery should therefore train on the compressed model's own on-policy states with dense token-level supervision, which On-Policy Distillation (OPD) provides by reusing the pre-compression model as a frozen teacher. However, long on-policy rollouts spend early recovery budget on low-information repetitive suffixes, delaying loss descent. To mitigate this waste, we propose \textbf{\shortopd}, a short-to-long OP

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

arXiv:2605.29668v2 Announce Type: replace-cross Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-improvement methods accumulate natural-language guidance without checking that each new item preserves previously correct behavior, so a note that fixes one trajectory can silently regress another. We introduce GRASP (Gated Regression-Aware Skill Proposer), which treats agent improvement as a sequence of edits to a bounded skill library, admitting each candidate only if it produces a net improvement on a balanced held-out probe under a hard regression budget. We evaluate GRASP across five base models on two FHIR-based clinical benchmarks, which score procedural reliability against FHIR state rather than clinical correctness or patient outcomes. On MedAgentBench, GRASP lifts gpt-oss-120b from 40.6% to 88.8%, exceeds the strongest of five self-improvement b

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Understanding and Mitigating Over-refusal for Large Language Models via Representation Intervention

arXiv:2511.19009v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate powerful capabilities across various natural language processing tasks,yet their inherent safety vulnerabilities undermine the reliable application of LLMs in real-world scenarios. To enhance LLM safety, various jailbreak defense methods have been proposed to guard against harmful outputs. However, improvements in model safety often come at the cost of severe over-refusal, failing to strike a good balance between safety and usability. This phenomenon is a critical reliability degradation issue in LLM intelligent systems, failing to strike a good balance between safety defense effectiveness and system usability reliability. In this paper, we first analyze the causes of over-refusal from a representation perspective, revealing that LLMs are unable to effectively distinguish between over-refusal samples and malicious samples. Based on this, we propose to mitigate overrefusal by intervening i

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models

arXiv:2510.25577v2 Announce Type: replace-cross Abstract: Recent advances in Speech Foundation Models (SFMs) enable direct processing of raw audio, allowing models to respond to subtle paralinguistic variation. However, how these models interpret non-lexical cues remains largely unstudied. We introduce VQ-Bench, a controlled evaluation suite featuring a parallel dataset of synthesized modal, breathy, creaky, and end-creak phonation types. We evaluate SFM sensitivity through open-ended generation across four ecologically valid domains, alongside speech emotion recognition. Our results reveal performance gaps: while a leading commercial API failed basic biometric sanity checks, other models exhibited systematic shifts in agency, empathy, and leadership based on phonation. Our findings also highlight gender asymmetries in salary and leadership endorsements, demonstrating that SFMs may mirror or amplify human social biases. This work establishes a reproducible framework for ensuring respon

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

BAT: Learning to Reason about Spatial Sounds with Large Language Models

arXiv:2402.01591v4 Announce Type: replace-cross Abstract: Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and dista

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Lost in Historical Time? A Polish History Matura Benchmark for Large Language Models

arXiv:2608.12343v2 Announce Type: replace Abstract: Language models are widely used by students as knowledge sources, yet benchmarks rarely assess their interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exam (Matura) in history - three official papers from 2023-2025, comprising short-answer questions and extended essays - and compare model performance against the human examinee population. Although models score near the ceiling, aggregate scores mask distinct competency profiles: rankings are unstable across task types, source modalities, and geographical scopes, with a consistent penalty for Polish versus Global history content. Qualitative error analysis reveals two recurring failure modes - source decontextualization, when models reason from source content rather than treating it as an object of analysis, and temporal disorientation, when responses are historically misplaced. This study introduces the first LLM history benchmark groun

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration

arXiv:2606.04743v2 Announce Type: replace Abstract: Agents are widely deployed as assistants over documents, tools, and code. However, they typically act only on explicit user requests, which surface only the problems the user has noticed, while many other important problems coexist, hidden in plain sight, within the broader user context, with their total number unknown in advance. We frame this as the task of discovering multiple hidden problems from context, in which coexisting problems should be uncovered, grounded in supporting evidence, and paired with concrete actions. To this end, we introduce TIDE, a template-guided iterative framework with two complementary mechanisms. Specifically, motivated by the observation that single-pass prediction anchors on the most salient cases and yields generic claims, we propose iterative discovery, which surfaces a small batch of candidates per round while conditioning on what has already been found, so subsequent rounds extend coverage; and tho

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

PhoneWorld: Scaling Phone-Use Agent Environments

arXiv:2605.29486v2 Announce Type: replace Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Existing mobile-agent benchmarks have made important progress on evaluation, but they do not by themselves provide a scalable way to construct many new phone-use environments. We present PhoneWorld, a reusable pipeline that converts real GUI trajectories and screenshots into controllable phone-use environments, executable tasks, automatic verifiers, and training rollouts. Rather than hand-building one mobile benchmark at a time, PhoneWorld uses real trajectories to recover which screens matter, how screens connect, which interactions must change environment state, and which user goals admit automatic verification. From these signals, it builds runnable mock Android apps backed by read-only app content and mutable state, then derives executable tasks, rule-based verifiers, and training roll

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Fine-grained Claim-level RAG Benchmark for Law

arXiv:2605.21071v4 Announce Type: replace Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses. In high-stake domains such as law, retrieval-augmented generation (RAG) is commonly used to mitigate hallucinations in generated responses. Nonetheless, prior work shows that RAG systems, whether general-purpose or legal-specific, still hallucinate at varying rates, making fine-grained evaluation essential. Despite the need, existing evaluation frameworks for legal RAG systems lack the granularity required to provide detailed analysis of retrieval and generation performance separately. Moreover, current benchmarks are largely English-only and centered on legal expert queries, overlooking non-expert needs. We introduce ClaimRAG-LAW, a comprehensive dataset for legal RAG that supports French and English, targets both experts and non-experts, and includes diverse quest

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

When Gradient Importance Lies: Adaptive LoRA Rank Allocation Fails Under GRPO

arXiv:2605.07366v2 Announce Type: replace Abstract: Adaptive rank allocation for LoRA - allocating more parameters to important layers and fewer to unimportant ones - consistently improves efficiency under supervised fine-tuning (SFT). We test whether this success transfers to reinforcement learning, specifically Group Relative Policy Optimization (GRPO). Using gradient-magnitude profiling on Qwen 2.5 1.5B with GSM8K, we find that, in our setting, it does not: proportional rank allocation degrades accuracy by 4.5 points compared to uniform allocation (70.0% vs. 74.5%), despite using identical parameter budgets. We identify two mechanisms behind this failure. First, the gradient landscape under GRPO is fundamentally flatter than under SFT: the max-to-min layer importance ratio is only 2.17x, whereas the layer concentration reported by Shi et al. (2024) for SFT (top 30% of layers carrying >80% of the gradient signal) implies a max/min ratio well above 10x. All layers carry meaningful gra

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization

arXiv:2604.26460v2 Announce Type: replace Abstract: Stylistic personalization - making LLMs write in a specific individual's style, rather than merely adapting to task preferences - lacks evaluation grounded in authorship science. We show that grounding evaluation in authorship verification theory transforms what benchmarks can measure. Drawing on three measurement traditions - LUAR (a trained authorship verification model), an LLM-as-judge with decoupled trait matching, and classical function-word stylometrics - we evaluate four inference-time personalization methods across 50 authors and 1,000 generations. The theory-grounded metric (LUAR) provides what ad hoc alternatives cannot: calibrated baselines (human ceiling 0.756, cross-author floor 0.626) that give scores absolute meaning. All methods score below this floor (0.484-0.508), exposing an authorship gap invisible to uncalibrated metrics. The three metrics produce near-zero pairwise correlations (|r| < 0.07), confirming that with

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Learning to Interrupt in Language-based Multi-agent Communication

arXiv:2604.06452v2 Announce Type: replace Abstract: When a colleague starts explaining something you already understand, you interrupt them. This simple act, a listener taking control of the conversation, is natural in human communication but absent in current verbose LLM multi-agent systems. Current approaches address communication efficiency only from the speaker side, compressing messages before they are sent. We flip the perspective: rather than making speakers more concise, we let listeners decide when they have heard enough. We propose a new communication paradigm in which the listener can interrupt the speaker mid-generation. We find that LLMs, given this ability, are overconfident and interrupt too early before receiving sufficient information. This finding motivates HANDRAISER, a learning method that predicts the right moment to interrupt based on estimated future reward and communication cost. We evaluate our framework on three multi-agent tasks: 2-agent text pictionary, 3-ag

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Early Stopping for Large Reasoning Models via Confidence Dynamics

arXiv:2604.04930v2 Announce Type: replace Abstract: Large reasoning models rely on long chain-of-thought generation to solve complex problems, but extended reasoning often incurs substantial computational cost and can even degrade performance due to overthinking. A key challenge is determining when the model should stop reasoning and produce the final answer. In this work, we study the confidence of intermediate answers during reasoning and observe two characteristic behaviors: correct reasoning trajectories often reach high-confidence answers early, while incorrect rollouts tend to produce long, unproductive reasoning traces and exhibit less reliable confidence dynamics. Motivated by these observations, we propose CoDE-Stop (Confidence Dynamics Early Stop), an early stopping method that leverages the dynamics of intermediate answer confidence to decide when to terminate reasoning, requiring no additional training and easily integrating into existing models. We evaluate CoDE-Stop on di

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Adaptive Stopping for Multi-Turn LLM Reasoning

arXiv:2604.01413v3 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly rely on multi-turn reasoning and interaction, such as adaptive retrieval-augmented generation (RAG) and ReAct-style agents, to answer difficult questions. These methods improve accuracy by iteratively retrieving information, reasoning, or acting, but introduce a key challenge: \textbf{When should the model stop?} Existing approaches rely on heuristic stopping rules or fixed turn budgets and provide no formal guarantees that the final prediction still contains the correct answer. This limitation is particularly problematic in high-stakes domains such as finance and healthcare, where unnecessary turns increase cost and latency, while stopping too early risks incorrect decisions. Conformal prediction (CP) provides formal coverage guarantees, but existing LLM-CP methods only apply to a single model output and cannot handle multi-turn pipelines with adaptive stopping. To address this gap, we propos

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

TriageSim: A Conversational Emergency Triage Simulation Framework from Structured Electronic Health Records

arXiv:2603.10035v3 Announce Type: replace Abstract: Research in emergency triage is restricted to structured electronic health records (EHR) due to regulatory constraints on nurse-patient interactions. We introduce TriageSim, a simulation framework for generating persona-conditioned triage conversations from structured records. TriageSim enables multi-turn nurse-patient interactions with explicit control over disfluency and decision behaviour, producing a corpus of ~800 synthetic transcripts and corresponding audio. We use a combination of automated analysis for linguistic, behavioural and acoustic fidelity alongside manual evaluation for medical fidelity using a random subset of 50 conversations. The utility of the generated corpus is examined via conversational triage classification. We observe modest agreement for acuity levels across three modalities: generated synthetic text, ASR transcripts, and direct audio inputs. We provide the code for TriageSim at https://github.com/dipankar

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models

arXiv:2602.09992v2 Announce Type: replace Abstract: Several recent contributions have evaluated the Poverty of the Stimulus Hypothesis (PoSH) using Artificial Neural Networks (ANNs). The results suggest that ANN-based language models can acquire certain structure-dependent generalizations from limited input without structural inductive biases of the type traditionally hypothesized by linguists. However, existing studies have largely focused on individual phenomena and adopted different evaluation protocols, leaving it unclear whether previous findings generalize across phenomena and learning conditions. We introduce \poshbench, a unified benchmark covering four canonical PoS phenomena. Training Transformer, LSTM, and n-gram models, we find that ANN-based models can achieve above-chance generalization from surprisingly limited input (10M words), but they show less efficient learning than children as input scale grows. Moreover, cognitively motivated inductive biases substantially improv

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Seeing is Coding: On the Effectiveness of Vision Language Models in Code Understanding

arXiv:2602.01785v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable success in source code understanding, yet as software systems grow in scale, computational efficiency has become a critical bottleneck. Currently, these models rely on a text-based paradigm that treats source code as a linear sequence of tokens, which leads to a linear increase in context length and associated computational costs. The rapid advancement of Multimodal LLMs (MLLMs) introduces an opportunity to optimize efficiency by representing source code as rendered images. Unlike text, which is difficult to compress without losing semantic meaning, the image modality is inherently suitable for compression. By adjusting resolution, images can be scaled to a fraction of their original token cost while remaining recognizable to vision-capable models. To explore the feasibility of this approach, we conduct the first systematic study on the effectiveness of MLLMs for code understanding

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Decoding Student Minds: Leveraging Conversational Agents for Psychological and Learning Analysis

arXiv:2512.10441v2 Announce Type: replace Abstract: This paper presents a psychologically-aware conversational agent designed to enhance both learning performance and emotional well-being in educational settings. The system combines Large Language Models (LLMs), a knowledge graph-enhanced BERT (KG-BERT), and a bidirectional Long Short-Term Memory (LSTM) network with attention to classify students' cognitive and affective states in real time. Unlike prior chatbots limited to either tutoring or affective support, our approach leverages multimodal data-including textual semantics, prosodic speech features, and temporal behavioral trends-to infer engagement, stress, and conceptual understanding. A pilot study with 45 university students demonstrated improved motivation, reduced stress, and moderate academic gains compared to unimodal baselines. We explicitly discuss the exploratory nature of this small-sample pilot, report effect sizes and inter-rater reliability alongside significance tes

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Research-Oriented Human-Centric Evaluation for Foundation Models

arXiv:2506.01793v2 Announce Type: replace Abstract: Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in human-AI collaboration. To address this gap, we propose a research-oriented Human-Centric Evaluation framework. It captures user perceptions across three core dimensions: problem-solving ability, information quality, and interaction experience, providing a structured, fine-grained approach to understanding how users evaluate and respond to model behavior in multi-modal research contexts. We conduct 604 human evaluation sessions across various disciplines, involving recent advanced foundation models. Through open-ended, time-limited collaborative tasks, we gather rich subjective assessments that highlight model capabilities and user preferences. Additionally, we perform an LLM-as-a-judge experiment and find that even sophisticated models struggle to accurately

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Leveraging Few-Shot Learning and Large Language Models for Analyzing Blood Pressure Variations Across Biological Sex from Scientific Literature

arXiv:2402.01826v2 Announce Type: replace Abstract: Current blood pressure (BP) technologies and standards were established decades ago, and these standards are still used worldwide today, often without adjusting BP readings for individual demographic factors such as sex and age. While these standards provide useful guidelines and help identify at-risk patients, they are not fully reliable for diagnosis due to the lack of demographic considerations. This study aims to assess the feasibility of using large language models (LLMs) for the automated extraction of BP-related information from the scientific literature, with a focus on biological sex-based distinctions in BP distributions. We employed natural language processing (NLP) methods to extract the means and standard deviations of BP values from the literature, distinguishing by biological sex. We developed a Solr-based search engine to retrieve scientific articles containing BP-related keywords and biological sex indicators from Pub

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

arXiv:2608.14509v1 Announce Type: cross Abstract: Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. This conflates two operations with different requirements. Interpreting a source rewards capacity and context. Combining interpretations rewards fixed arithmetic, comparability across instances, and the option to return nothing. Once separated, the design problem becomes the interface between them. We propose a four-field evidence tuple (hypothesis, reliability bucket, rationale, provenance) and show that fixing it determines both halves. The separation also reveals a failure mode in how such systems combine, which we call count-scale drift. Thresholding a sum of unnormalized weights is exactly posterior thresholding, but at an operating point that slides with the number of sources consulted. The slide grows with reader reliability. When source reliabilities differ, the vote rule and the posterior order instances different

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

LLMs Don't Pay for the Jump

arXiv:2608.14397v1 Announce Type: cross Abstract: Zahavy [2026] argues that Large Language Models, despite their capabilities in induction and deduction, cannot perform the abductive "Jump" that produced Einstein's equivalence principle, and attributes this limitation to the absence of embodied simulation. Zheng-Xin [2026] and Farmer [2026] question whether embodiment is necessary for abduction, pointing to alternative routes to General Relativity and forms of abduction that require no sensorimotor grounding. Max Planck resolved the blackbody radiation problem in 1900. Planck's move to E = h{\nu} required no embodied simulation. It was motivated by a mathematical consequence of classical theory, an infinite predicted energy for a finite measured quantity, that could not be physically accepted. We show that neither induction nor deduction could have produced the postulate and argue that its adoption required a coupling between epistemic error and physical cost. We formalize this distinc

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

arXiv:2608.14375v1 Announce Type: cross Abstract: Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likely to be correct is also worth keeping. Yet a wrong answer can contain a useful decomposition, constraint, or scientific principle. We test this distinction with Diverse Hypothesis Deliberation (DHD), a controlled measurement protocol that caches five independently generated messages and replays the same downstream solver, called the integrator, with each message available or hidden. The replay comparison measures a message's trajectory value: whether making the message available helps or harms subsequent reasoning. Across five mathematics and science benchmarks and two openly available model families, gpt-oss-120b and gemma-4-31B-it, wrong-helpful messages appear in every benchmark-model combination. Among wrong-answer messages that change final correctness, m

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CL

AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

arXiv:2608.14320v1 Announce Type: cross Abstract: The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large language models (LLMs) exhibit similar behavior. However, existing work on anchoring in LLMs typically evaluates only a narrow set of anchor pathways and rarely distinguishes irrelevant from plausible anchors. We introduce AnchorBench, a benchmark for the anchoring effect in LLMs that evaluates multiple anchor pathways under an explicit anchor relevance axis. Across fourteen models, including ten open-weight models and four frontier API models, and a large set of controlled prompts, we find that (1) anchoring is strongly pathway-dependent, (2) plausible anchors usually induce larger shifts than irrelevant ones when introduced through stronger pathways, (3) anchor influence generally weakens as the anchor moves farther

Source ↗
Showing 9651–9700 of 18694 signals
← Prev Page 194 of 374 Next →