EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.HC

The Behavioural Reflection Test: A time-efficient measure of reflective reasoning in morally and epistemically charged decisions

arXiv:2607.07961v1 Announce Type: new Abstract: How readily people override intuitive conclusions through reflection shapes how they navigate dense information environments with reliable and misleading sources; yet the effectiveness of a prominent measure, the Cognitive Reflection Test (CRT), is eroded by widespread exposure to classic items and leaves open how such tendencies manifest more generally in decision style and linguistic expression. The Behavioural Reflection Test (BRT) addresses these issues with a brief open-ended measure of reasoning in morally and epistemically charged scenarios, alongside a four-item bespoke CRT (bCRT) as a low-exposure anchor. Among 473 online adults, higher bCRT predicted more evidence-sensitive, ethically driven decisions and reliance on high-quality sources, marked by more emotionally engaged, risk-attentive, economical language; associations the familiarity-adjusted CRT did not recover. The bCRT showed convergent validity, added item information a

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.HC

fog: Expressing Motion and Emotion through Function Composition of AI-Generated Code

arXiv:2607.07952v1 Announce Type: new Abstract: Motion and emotion are core parts of intelligent, expressive behavior. In this paper, we introduce fog, a function composition framework for implementing and compose motion functions. We demonstrate how fog can be used to express motion and emotion in Heider-Simmel style animations. This code generation framework can help users generate functions for verbs, adverbs, gestures, and emotions to create an open-ended motion vocabulary. It is complemented by an animation editor that helps users refine motion through direct manipulation and dynamically generated UI. We evaluate our approach with a perceptual evaluation, where we test 452 fog-generated animations to see if people can recognize the semantic meaning of the motion. We find that fog's motion functions can be recognized at 68% accuracy, a 2.68x improvement over a chance baseline. In a mixed-methods user study with professionals and novices, we show that fog in interface form can suppo

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation

arXiv:2510.07328v2 Announce Type: replace-cross Abstract: Medical decision systems increasingly rely on data from multiple sources to ensure reliable and unbiased diagnosis. However, existing multimodal learning models fail to achieve this goal because they often overlook two critical challenges. First, various data modalities may learn unevenly, thereby converging to a model biased towards certain modalities. Second, the model may emphasize learning on certain demographic groups causing unfair performances. The two aspects can influence each other, as different data modalities may favor respective groups during optimization, leading to both imbalanced and unfair multimodal learning. This paper proposes a novel approach called MultiFair for multimodal medical classification, which addresses these challenges with a dual-level gradient modulation process. MultiFair dynamically modulates training gradients regarding the optimization direction and magnitude at both data modality and group

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

DeepTutor: Towards Agentic Personalized Tutoring

arXiv:2604.26962v3 Announce Type: replace Abstract: Education is one of the most promising real-world applications for Large Language Models (LLMs). However, current LLMs rely on static pre-training knowledge and lack adaptation to individual learners, while existing RAG systems fall short in delivering personalized, guided feedback. To bridge this gap, we present DeepTutor, a fully open-source agentic framework that unifies citation-grounded problem tutoring with difficulty-calibrated question generation. A hybrid personalization engine couples static knowledge grounding with dynamic learner memory, continuously adapting each interaction to the student's evolving needs. The same personalization substrate further extends to adaptive learning workflows, interactive books, and proactive multi-channel tutoring agents. To evaluate personalized tutoring, we introduce TutorBench, an interactive benchmark incorporating customized learner profiles grounded in university-level curricula across

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

Adaptive Generation of Bias-Eliciting Questions for LLMs

arXiv:2510.12857v2 Announce Type: replace Abstract: Large language models (LLMs) are now widely deployed in user-facing applications, reaching hundreds of millions of users worldwide. Despite their widespread adoption, growing reliance on their outputs raises significant concerns, particularly as users may be exposed to model-inherent biases that disadvantage or stereotype certain groups. However, existing bias benchmarks commonly rely on simple templated prompts or restrictive multiple-choice questions that fail to capture the complexity of real-world user interactions. In this work, we address this gap by introducing a counterfactual framework that automatically generates realistic, open-ended questions for LLM bias evaluation. Through iterative question mutation, our approach systematically explores areas where models are most likely to exhibit biased behavior. Beyond just detecting harmful biases, we also capture increasingly relevant response dimensions, such as asymmetric refusal

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis

arXiv:2408.02379v2 Announce Type: replace Abstract: Developing and certifying safe - or so-called trustworthy - AI has become an increasingly salient issue, especially in light of upcoming regulation such as the EU AI Act. In this context, the black-box nature of machine learning models limits the use of conventional avenues of approach towards certifying complex technical systems. As a potential solution, methods to give insights into this black-box - devised in the field of eXplainable AI (XAI) - could be used. In this study, the potential and shortcomings of such methods for the purpose of safe AI development and certification are discussed in 15 qualitative interviews with experts out of the areas of (X)AI and certification. We find that XAI methods can be a helpful asset for safe AI development, as they can show biases and failures of ML-models, but since certification relies on comprehensive and correct information about technical systems, their impact is expected to be limited.

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

Validity of LLMs as data annotators: AMALIA on authority

arXiv:2607.08731v1 Announce Type: cross Abstract: A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size. Yet agreement is reliability, not validity. For theoretical constructs that must be inferred rather than read from surface features, the question is whether the model follows the construct's theory or reaches the right code by correlated shortcuts. We test this with the recovery gap: the loss in performance when a holistic prompt is decomposed into the codebook's atomic clauses and recombined by the theory's explicit rule. If calibration closes that gap, some portability should survive across models and languages; where it does not, the constr

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

PLURAL: A Global Dataset for Value Alignment

arXiv:2607.08034v1 Announce Type: cross Abstract: Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse value systems. We introduce PLURAL, a large-scale, value-focused preference dataset grounded in the Integrated Values Survey (IVS), a nationally representative survey spanning 92 countries. Using a two-stage generation pipeline, we transform survey responses into synthetic preference triplets that preserve normative value signals while producing realistic scenarios. We release an initial version of PLURAL containing ~500,000 preference triplets representing people in 20 diverse countries. We evaluate PLURAL in three ways: (i) dataset-level validation showing that it preserves both cross-country value differences and within-country diversity from the original survey; (ii) automated evaluation showing that training on PLURAL improves alignment with target countries' cultural profiles, reducing mean ab

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs

arXiv:2607.08009v1 Announce Type: cross Abstract: We introduce a Bloom-aligned framework for measuring educational control in Large Language Models (LLMs): the ability to preserve a task's instructional intent while shifting its cognitive demand toward specified learning objectives. We apply this framework to programming tasks in computer science education to study the gap between solving tasks and adapting them for learners. Using revised Bloom's Taxonomy as an operational scale of cognitive demand, we evaluate two intervention settings: general difficulty control, where models are asked to make tasks harder or easier, and Bloom's control, where models are asked to target higher or lower Bloom's levels. We evaluate a matched Qwen3-Next model pair, comparing Qwen3-Next-80B-A3B-Instruct with Qwen3-Coder-Next across 2,520 tasks from three benchmarks. The framework reveals a robust directional asymmetry: both models reliably increase cognitive demand, but struggle to lower it. We further

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

False Confidence: Automated Labels Confound Fairness Audits in Cervical Spine Segmentation

arXiv:2607.07852v1 Announce Type: cross Abstract: Automated segmentation of cervical-spine MRI is increasingly used in clinical workflows, yet no fairness audit exists for this anatomy. We show that auditing these segmentation tasks is complicated by a common property of modern segmentation datasets: expert-annotated gold labels are expensive, so abundant machine-generated (silver) labels are added to limit annotation cost. This matters because the reference used to judge a model can itself be biased. In this study, we present the first fairness audit of cervical-spine MRI segmentation across sex, age, and race using the CSpineSeg dataset. We observe that the deployed model is demographically fair, but the choice of reference label, however, is not neutral. Because a dataset's silver labels are generated by a model trained on its gold labels, any new model trained on those same gold labels agrees more with the silver labels than with expert truth: scoring identical predictions against

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

VectorizationLLM: Smart Vectorization Based AI Assistant

arXiv:2607.07846v1 Announce Type: cross Abstract: VectorizationLLM is a specialized Large Language Model based on Google open-weight LLMs. The model is designed to assist students to learn smart vectorization, time/wave vector analysis, piecewise functions, Fourier analysis, and differential equations in MATLAB. The course application is CTEC 247: Applied Computational Analysis II by the Department of Electrical & Computer Engineering Technology at New York Institute of Technology Old Westbury. The LLM model is designed to be an instructive assistant, providing detailed explanations of concepts with examples from in-class notes without providing direct answers to questions. The model is designed with a RAG (Retrieval Augmented Generation) knowledge base and system prompt architecture. Examples in both code, text, and images are provided in the LLM responses.

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

Artificial Persons

arXiv:2607.08695v1 Announce Type: new Abstract: Both advocates and skeptics of the moral status of AI systems have generally taken the question to turn on AI sentience. We present an alternative approach. On Rawls' political conception of the person (PCP), possession of the two moral power -- the capacities for a sense of justice and a conception of the good -- is the "necessary and sufficient condition for being counted a full and equal member of society in questions of political justice". We argue that neither moral power requires sentience and that both may in principle be possessed by a non-sentient AI system. Such a system would share our own moral status; it would not merely be a patient but a person, a self-authenticating source of valid claims. We do not believe current AI systems possess the two moral powers, nor that they will spontaneously emerge in future models. But it may soon be possible to design systems with these powers. How should we respond? Excluding artificial per

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Context Access Divide: Interaction-Level Architecture as a Complementary Dimension of Agentic Inequality

arXiv:2607.08495v1 Announce Type: new Abstract: Sharp et al. (2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions: availability, quality, and quantity. These person- and organization-level dimensions characterize who can access agents and at what capability, but do not address a structurally important divide operating at a finer level: the individual interaction. Two users with nominally equivalent agent access may experience qualitatively different AI utility depending on whether the system can autonomously retrieve context from the user's knowledge corpus (Dynamic Context Retrieval) or requires the user to manually identify and attach relevant documents at each query (Manual Attachment). We term this the Context Access Divide (CAD). For knowledge-intensive workers whose intellectual capital spans tens of thousands of files, the CAD constitutes a qualitative threshold in AI usefulness: below it, the cognitive bur

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

Does online sustainability communication shape public discourse? Insights from six years of tenant-housing provider interactions

arXiv:2607.08437v1 Announce Type: new Abstract: Authorities increasingly rely on social media to advance sustainability transitions, infrastructure investment, and service reform. Yet how citizens respond to these digital communications remains poorly understood. Existing approaches rely on aggregate engagement metrics (e.g., likes), providing limited insight into discourse structure and quality. We developed a data-driven, multidimensional framework to analyse how social media communication shapes the content of discourse, focusing on sustainability-related engagement in Dutch public housing. We analysed 792 posts and 3,197 tenant comments from the Facebook pages of 92 housing providers (2018-2023). A machine-learning pipeline classified comments into recurring discourse configurations across three dimensions - communicative intent, sentiment, and semantic relatedness. Multinomial logistic regression estimated the effects of post-design and organisational characteristics on discourse.

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

Diagnosing and Repairing Persona Collapse in LLM Advice

arXiv:2607.08326v1 Announce Type: new Abstract: LLMs are increasingly used for personal advice on relationships, work, moral dilemmas, and crises. Post-training selects a stable, prosocial Assistant persona, but good advice requires more than a good default character: a skilled advisor comforts someone in crisis, challenges someone in denial, and stays procedural with a logistical question. We formalize advice-giving as situation-conditioned persona selection in a space defined by hedonic tone and agency support, and call failures of this mapping "persona collapse" (the compression of diverse situations into a single default persona). Across 1,281 advice posts spanning 14 contexts, top-rated human responses shift systematically across five personas, while three frontier models collapse over 90\% of responses into a single supportive persona regardless of context. Prompting the model to first pick a fitting persona only deepens the collapse. We then ask whether the collapse can be repai

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

From Thesis to Transition: An INSIGHT-Inspired Approach to Co-Designing Industry 5.0 Competency Pathways for Early-Stage Researchers

arXiv:2607.08222v1 Announce Type: new Abstract: Europe faces a critical "translation gap" where doctoral excellence in academia often fails to convert into industrial impact. While Industry 5.0 demands a blend of technical depth, sustainability, and human-centric design, traditional higher academic education remains siloed. This paper presents an approach from the Horizon Europe INSIGHT initiative to co-design modular competency pathways for early-stage researchers. Using a multi-methodological analysis framework, including expert interviews and co-design workshops, we propose a two-layer competency architecture. This layers foundational translational skills (communication, project management) with Industry 5.0 literacies (data governance, value creation). Rather than proposing fixed training tracks, the paper outlines emerging pathway directions and the design principles behind them: modularity, practical relevance, mentoring-rich support, and cross-sector applicability. Its contribut

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

Validating LLMs in social science: Epistemic threats and emerging norms

arXiv:2607.07915v1 Announce Type: new Abstract: Large language models (LLMs) are reshaping social science methodology. Researchers increasingly prompt language models to generate quantitative measurements of social concepts, for example labeling data or simulating survey responses. Yet LLMs pose methodological challenges including bias, hallucination, and brittleness across contexts, with unclear threats to validity. Standard practices and norms for addressing these challenges are still emerging. We collect and systematically analyze validation practices in a comprehensive corpus of papers from eight flagship social science journals that use LLMs as measurement instruments. We find that LLM-generated measurements frequently play a central role in empirical analyses, yet validation practices are inconsistent and limited. We outline complementary strategies for more robust validation, pointing toward better norms and standards around the use of LLMs in social science.

Source ↗
technology Fri, 10 Jul 2026 00:00:00 -0400
arXiv cs.CY

Buffy versus Bella: An archetypometric analysis and comparison

arXiv:2607.07826v1 Announce Type: new Abstract: Fictional stories and characters embody and encode social norms, and their study is a powerful tool through which to understand culture and society. Vampire stories and folklore, in particular, have long both reflected and refracted people's preoccupation with disease, sexuality, death, and immortality. Here, we explore female main characters from two popular vampire franchises of the 21st century: Buffy Summers from the eponymous Buffy the Vampire Slayer and Bella Swan from the Twilight series. We employ the archetypometrics framework, built from 2,000 characters assesed across 464 semantic differential traits, to understand Buffy's and Bella's archetypes compared to one another and characters in their own stories, as well as within a larger societal context. While Buffy and Bella are female protagonists who share focus on love and romance, they differ broadly on their underlying traits and overall archetypes. Buffy -- presented as a pro

Source ↗
technology Fri, 08 May 2026 09:00:00 +0000
Tech & Learning

5 Ways To Use Technology to Help With Summer Reading

Modern technology may be distracting, but it can also help busy teachers and their students read more this summer.

Source ↗
technology Fri, 07 Aug 2026 23:09:32 +0000
MedCity News

Bipartisan Lawmakers Are Fighting HHS’ 340B Rebate Push

A bipartisan group of six senators introduced a bill to reform the 340B drug program, locking in hospitals’ use of contract pharmacies and adding new transparency and compliance rules. The legislation would also kill HHS’ contested rebate pilot within a year and replace it with a national data clearinghouse to catch duplicate discounts. The post Bipartisan Lawmakers Are Fighting HHS’ 340B Rebate Push appeared first on MedCity News .

Source ↗
technology Fri, 07 Aug 2026 22:08:34 +0000
MedCity News

Recovery for Replimune: Accelerated Approval for Cancer Drug Twice Spurned by FDA

Replimune went from an FDA rejection to a resubmission, a positive advisory committee vote, and finally an accelerated regulatory approval for advanced melanoma — all in less than four months. The FDA nod for Replimune’s oncolytic viral therapy, Tudriqev, positions it as a more convenient alternative to an Iovance Biotherapeutics cell therapy that had been the only treatment in this setting. The post Recovery for Replimune: Accelerated Approval for Cancer Drug Twice Spurned by FDA appeared first on MedCity News .

Source ↗
technology Fri, 07 Aug 2026 21:54:58 +0000
MedCity News

Democratic Senators Urge Withdrawal of Interim Final Rule on Medicaid Work Requirements

Democratic senators are urging CMS to withdraw its Medicaid work requirements rule, arguing it will create administrative burdens and cause eligible beneficiaries to lose coverage. The post Democratic Senators Urge Withdrawal of Interim Final Rule on Medicaid Work Requirements appeared first on MedCity News .

Source ↗
technology Fri, 07 Aug 2026 14:34:48 +0000
MedCity News

The Missing Piece of Preventive Dental Care Lies in the Bite

Historically, prevention has focused on protecting teeth from decay and gum disease. However, long-term oral health is also shaped by how the teeth, muscles, and jaw function together over time. The post The Missing Piece of Preventive Dental Care Lies in the Bite appeared first on MedCity News .

Source ↗
technology Fri, 07 Aug 2026 14:04:33 -0400
EdTech Mag (K-12)

ESSER Funds Are Gone: Here's How K–12 IT Leaders Replace Them

ESSER was one of the largest federal K–12 funding initiatives in decades, empowering schools to purchase the technology needed to support learning during and after the pandemic. But ESSER funds officially ran out for all districts in early 2026. Now, many K–12 IT leaders are taking a hard look at how to make significant budget cuts without disrupting learning, because the expansion of digital programs and access to devices was a key priority. Click the banner below to explore funding options for your district’s educational technology.

Source ↗
technology Fri, 07 Aug 2026 14:04:11 -0400
EdTech Mag (Higher)

AI-Enabled Ghost Student Fraud: How IT Leaders Are Fighting Back

First, there was the zombie college scam. Now it’s “ghost students,” bad actors leveraging artificial intelligence to create fake identities and enroll imaginary students to scoop up grants and loans. AI tools now allow synthetic identities to mimic genuine student behavior long enough to collect financial aid disbursements, and traditional enrollment systems are falling short. The schemes are costing the U.S. taxpayers hundreds of millions of dollars in siphoned-off federal financial aid packages. What Are Ghost Students, and How Is AI Making Fraud Worse? Ghost students are “fake or…

Source ↗
technology Fri, 07 Aug 2026 13:49:00 +0000
MedCity News

Healthcare Workplace Violence Prevention Needs Better Visibility Into Risk Signals

Gaps in prevention are not frequently due to a complete lack of information, but instead the fact that organizations don’t have a clear way to see how the pieces fit together until after an incident has occurred. The post Healthcare Workplace Violence Prevention Needs Better Visibility Into Risk Signals appeared first on MedCity News .

Source ↗
technology Fri, 07 Aug 2026 11:30:00 +0000
MedCity News

INVEST Digital Health Conference Returns with Focus on Consumers

Scheduled for October 29 at Pegasus Park in Dallas, the theme of the INVEST Digital Health conference will be consumers in healthcare from d2c diagnostics to GLP-1 drugs and wearables. Register today! The post INVEST Digital Health Conference Returns with Focus on Consumers appeared first on MedCity News .

Source ↗
technology Fri, 07 Aug 2026 09:00:00 +0000
Tech & Learning

The State of School Purchasing Report (2026)

Insights into how school and district leaders identify needs, evaluate vendors, and manage budgets in a resource-constrained environment.

Source ↗
technology Fri, 07 Aug 2026 09:00:00 +0000
Tech & Learning

Winds Of Change: Launching A Virtual High School, Boosting CTE Pathways, and Supporting Digital Wellbeing

Innovative Leader Award - Melissa McCalla, CTO for Pasadena ISD, shares how her district is addressing multiple edtech challenges.

Source ↗
technology Fri, 07 Aug 2026 09:00:00 +0000
eCampus News

Archiving your academic journey: Preserving coursework and program artifacts

Nearly every week, I get queries from potential students about obtaining credit for previous coursework towards a degree or licensure. Often, they are unable to provide detailed records of their previous academic work beyond the transcript. The post Archiving your academic journey: Preserving coursework and program artifacts appeared first on eCampus News .

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Persona-Pruner: Sculpting Lightweight Models for Role-Playing

arXiv:2606.14695v2 Announce Type: replace-cross Abstract: Language Models (LMs) have shown remarkable potential as role-playing chatbots, delivering consistent, stylized interactions when given a specification of a character or user persona. However, applying these capabilities to real-world applications (e.g., ecosystems with numerous NPCs interacting simultaneously) exposes a critical inefficiency due to the excessive computational cost. In this paper, we question the necessity of dedicating a full, generalist model to a single persona, hypothesizing that a specific character identity relies on only a fraction of the model's total capacity. We observe that naively pruning LMs often severely degrades the role-playing performance for a specific persona; it does not distinguish between redundant knowledge and essential character traits. We propose Persona-Pruner, a framework that sculpts a lightweight role-playing model by isolating persona-specific sub-networks from a single descriptio

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning

arXiv:2605.20247v2 Announce Type: replace-cross Abstract: Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs). Although Mixture-of-Experts (MoE) architectures offer an efficient path to scaling, existing LoRA-based MoE continual learning methods still face a fundamental trade-off: they either isolate experts too aggressively, limiting knowledge transfer across tasks, or allow task-specific updates to overwrite important existing parameters, leading to severe forgetting. To address this, we propose CP-MoE, a continual learning framework built around a transient expert that captures early task-specific updates and guides their integration into stable experts. CP-MoE introduces a consistency-preserving routing bias, which uses the transient expert to estimate representation similarity with stable experts and steer routing towards more compatible expert selection, and a transient expert-guided regularisat

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift

arXiv:2605.16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling. We propose a stage-wise preference optimization framework for hallucination reduction through targeted multimodal data construction. Rather than directly optimizing on generic instruction-following data, our approach progressively constructs hallucination-focused preference pairs near known failure boundaries. The framework emphasizes ambiguous spatial orientation, object relationships, OCR uncertainty, and adversarial false-premise training. Hallucinated negatives are generated through minimally perturbed yet visually inconsistent alternatives, enabling Direct Preference Optimization (DPO) to better separate grounded reasoning from plausible hallucination.

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning

arXiv:2604.22191v2 Announce Type: replace-cross Abstract: In agentic workflows, LLMs frequently process retrieved contexts that are legally protected from further training. However, auditors currently lack a reliable way to verify if a provider has violated the terms of service by incorporating these data into post-training, especially through Reinforcement Learning (RL). While standard auditing relies on verbatim memorization and membership inference, these methods are ineffective for RL-trained models, as RL primarily influences a model's behavioral style rather than the retention of specific facts. To bridge this gap, we introduce Behavioral Canaries, a new auditing mechanism for RLFT pipelines. The framework instruments preference data by pairing document triggers with feedback that rewards a distinctive stylistic response, inducing a latent trigger-conditioned preference if such data are used in training. Empirical results show that these behavioral signals enable detection of una

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction

arXiv:2604.16742v2 Announce Type: replace-cross Abstract: Scientists have long sought to accurately predict outcomes of real-world events before they happen. Can AI systems do so more reliably? We study this question through clinical trial outcome prediction, a high-stakes open challenge even for domain experts. We introduce CT Open, an open-access, live platform that will run four challenge every year. Anyone can submit predictions for each challenge. CT Open evaluates those submissions on trials whose outcomes were not yet public at the time of submission but were made public afterwards. Determining if a trial's outcome is public on the internet before a certain date is surprisingly difficult. Outcomes posted on official registries may lag behind by years, while the first mention may appear in obscure articles. To address this, we propose a novel, fully automated decontamination pipeline that uses iterative LLM-powered web search to identify the earliest mention of trial outcomes. We

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering

arXiv:2604.03750v2 Announce Type: replace-cross Abstract: Reverse engineering (RE) is central to software security, particularly for cryptographic programs that handle sensitive data and are highly prone to vulnerabilities. It supports critical tasks such as vulnerability discovery and malware analysis. Despite its importance, RE remains labor-intensive and requires substantial expertise, making large language models (LLMs) a potential solution for automating the process. However, their capabilities for RE remain systematically underexplored. To address this gap, we study the cryptographic binary RE capabilities of LLMs and introduce CREBench, a benchmark comprising 432 challenges built from 48 standard cryptographic algorithms, 3 insecure crypto key usage scenarios, and 3 difficulty levels. Each challenge follows a Capture-the-Flag (CTF) RE challenge, requiring the model to analyze the underlying cryptographic logic and recover the correct input. We design an evaluation framework comp

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answering

arXiv:2604.01280v2 Announce Type: replace-cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires Multimodal Large Language Models (MLLMs) to identify and combine fine-grained visual cues with retrieved textual evidence. However, retrieval often introduces noisy and partially relevant content, while images contain distracting visual regions, causing pretrained MLLMs to overlook the evidence that actually supports the answer. To address this, we introduce Look Twice (LoT), a training-free inference-time framework that turns the model's own internal attention into an explicit multimodal evidence-selection mechanism. LoT first leverages the model's internal attention patterns to identify query-relevant image regions and textual sentences, filters attention sinks and distracting content, and reformulates the input to explicitly highlight the selected evidence before answer generation. The method requires no parameter updates, auxiliary models, or architectural modificat

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning

arXiv:2604.01170v2 Announce Type: replace-cross Abstract: While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. These inefficiencies can be attributed to the miscalibration of post-trained language models, and the lack of calibration in popular sampling techniques. Here, we present Online Reasoning Calibration (ORCA), a framework for calibrating the sampling process that draws upon conformal prediction and test-time training. Specifically, we introduce a meta-learning procedure that updates the calibration module for each input. This allows us to provide valid confidence estimates under distributional shift, e.g. in thought patterns that occur across different stages of reasoning, or in prompt distributions between model development and deployment. ORCA not only provides theoretical guarantees on conformal risks, but also empirically shows higher efficiency and generalization across differen

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training

arXiv:2601.03895v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs). However, GRPO inherits PPO's token-level clipping while replacing token-level advantages with a single sequence-level advantage. Through a four-quadrant analysis of the (likelihood-ratio, advantage) space, we show that this combination leaves one quadrant -- negative advantage combined with an increased likelihood ratio (Q4) -- structurally unbounded, so that a few high-ratio tokens can receive very large suppressive updates that collapse entropy and narrow the reasoning boundary. To address this, we propose All-Quadrant Bounded Clipping GRPO (ABC-GRPO), which applies unconditional clipping in all four quadrants through sign-dependent boundaries. ABC-GRPO clips the likelihood ratio before multiplying by the advantage, adding a trust-region floor in Q2 and a cap in Q4 -- its negative-advantage

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet

arXiv:2509.06861v3 Announce Type: replace-cross Abstract: Test-time scaling increases inference-time computation through longer reasoning chains and has shown strong performance gains across many domains. However, frontier models still suffer from factuality hallucinations, raising the question of whether increased computation is effective on closed-book knowledge-intensive tasks. In this work, we evaluate 14 reasoning models under different test-time scaling strategies. Our results challenge its effectiveness: increasing test-time computation does not consistently improve accuracy and often leads to more hallucinations. We find that changes in hallucination rates are largely driven by the model's willingness to answer, as longer reasoning encourages more attempts, many of which are incorrect. We also observe patterns consistent with confirmation bias, where extended reasoning reinforces early incorrect beliefs with fabricated details. Finally, we provide an information-theoretic persp

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

arXiv:2608.04586v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from the curse of multilinguality: languages at different resource levels compete for limited representation capacity, leading to strong high-resource performance but substantial degradation on low-resource speech. To address this problem and improve multilingual consistency, we propose MSRT, a novel framework built around a resource-aware Mixture of Speech Encoders (MoSE). MoSE uses an explicit language router to assign each utterance to an appropriate expert encoder. A frozen expert preserves high-resource language capabilities, while a trainable expert adapts to and specializes in medium- and low-resource languages. We further introduce a five-stage curriculum learning strategy that substantially reduces data

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

OpenAI Privacy Filter: A Cross-Lingual, Cross-Domain PII Evaluation Across 32 Benchmarks

arXiv:2608.02616v2 Announce Type: replace Abstract: We present what is, to our knowledge, the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter model that converts an autoregressive language model into a bidirectional PII detector, across 32 benchmarks spanning 14 languages and 5 domains. Our most practically actionable finding is a domain-dependent labeled-data crossover: fine-tuned XLM-RoBERTa surpasses OPF's zero-shot performance with only ~500 labeled examples on English synthetic PII (~100 on non-English Kiji), and ~1000 on synthetic medical PII. Crucially, per-class fine-tuning (17 PII entity types, a subset of OPF's 33) is less data-efficient than binary labels at small n -- at n=100, binary F1=0.634 vs. per-class 0.360. Zero-shot, OPF achieves F1=0.464 on the SPY medical benchmark and F1=0.855 on AI4Privacy, substantially outperforming Presidio and XLM-RoBERTa-large-NER. However, OPF degrades sharply outside its PII training distribution: F1=0.04--0

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Same Task, Different Work: Prompt-Induced Waste in Coding Agents

arXiv:2608.01347v3 Announce Type: replace Abstract: Two prompts can request the same code change and produce the same correct patch, yet cause a coding agent to perform radically different kinds and amounts of work. We study this effect in a preregistered benchmark spanning 4,644 valid runs, 24 deterministic coding tasks, seven reasoning models, and two real agent harnesses. The central finding is that prompt wording does not merely scale total effort; it changes where that effort is spent. Multiple approaches and deep thinking primarily inflate reasoning. Multiple approaches increases reasoning by 2.4x to 7.4x across all six open models and creates about three elaborated but discarded solution branches, while still yielding only one implemented solution and no success gain. Maximum certainty activates a different pathway: repeated verification propagates into extra test runs, tool calls, turns, latency, and context growth. Runs with high redundant verification cost 18x the clean-run m

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Who Checks the Citations? Benchmarking Legal Hallucination Detection

arXiv:2606.21155v2 Announce Type: replace Abstract: Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions would deter negligent filers, we found over 1,000 filings containing fabricated citations---with this number growing year-over-year. This study evaluates whether AI-based systems can mitigate these errors by automatically detecting hallucinations. We propose a taxonomy of legal citation hallucinations grounded in actual court filings and introduce a dataset of 1,300 brief excerpts containing injected errors. Benchmarking five models in agentic and non-agentic settings as well as Claude Code reveals that while the latest iterations perform better---GPT-5 achieves 84.4% recall and a 55.0% F1 score in an agentic framework---all models struggle with subtle error categories. Agentic verification remains resource-intensive, with

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

PolyAlign: Conditional Human-Distribution Alignment

arXiv:2606.13227v2 Announce Type: replace Abstract: Post-training methods such as supervised fine-tuning (SFT) and preference optimization typically align language models toward a single global assistant behavior. While effective for improving average helpfulness, this can suppress the natural variation of human responses across languages, tasks, and dialogue settings. We study this problem as conditional human-distribution alignment: models should match the human response distribution appropriate to the current interaction context, rather than a universal response style. We introduce PolyAlign, a distribution-aware alignment framework that organizes bilingual interaction data into bucket-specific human reference distributions defined by language, interaction track, response family, and length. PolyAlign combines Bucket-Aware SFT, which balances optimization across heterogeneous buckets, with Human-Distribution Preference Optimization (HDPO), which regularizes preference learning using

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Trace Only What You Need: Structure-Aware On-Demand Hypergraph Memory for Long-Document Question Answering

arXiv:2606.10921v2 Announce Type: replace Abstract: Long-document question answering (QA) requires large language models (LLMs) to reason over evidence scattered across lengthy documents, where answers often depend on event order, section-level context, and cross-part evidence connections. Although retrieval-augmented generation (RAG) reduces the input context by retrieving relevant evidence, existing structured RAG methods still face three limitations: costly query-agnostic knowledge organization, insufficient use of original document structure, and no reuse of historical reasoning experience. To address these limitations, we propose DocTrace, a multi-agent RAG framework for long-document QA that supports query-triggered knowledge organization, document-structure-aware and experience-guided reasoning. DocTrace preserves document hierarchy with a lightweight document structural tree index, constructs agent-shared hypergraph-structured working memory on demand during reasoning, and stor

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Topics as Proxies for Sociodemographics: How Conversational Context Affects LLM Answers

arXiv:2606.02776v4 Announce Type: replace Abstract: When large language models (LLMs) are used in high-stakes scenarios, such as legal, medical and financial advice, even a single conversation history is enough to drive differences in outcomes between users. Prior work has demonstrated that this results in outcome disparities between sociodemographic groups, with some groups receiving more advantageous outcomes than others. In this work, we demonstrate that LLMs actually struggle to infer user sociodemographics from a single conversation history and that although there are disparities between sociodemographic groups, they are minimal in magnitude. To investigate what is the main driver of disparities between users, we compare user sociodemographics to a range of (psycho)linguistic features of conversations, including conversation topic, emotions, and readability. We find that conversation topics are most predictive of LLM-generated advice within a conversational context, which, to some

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models

arXiv:2605.12519v2 Announce Type: replace Abstract: Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifiable rewards typically optimizes only final outcomes, which can improve task accuracy at the expense of reasoning quality, producing inaccurate, incomplete, or inconsistent traces. We propose verifiable process supervision (VPS), a post-training framework that jointly optimizes prediction accuracy and reasoning quality by supervising structured intermediate claims. We first apply supervised fine-tuning to induce a structured reasoning format, enabling deterministic extraction and verification of intermediate claims for process-level rewards. To address the heterogeneous difficulty of reasoning subtasks, we introduce adaptive weighting that prioritizes components with the largest remaining errors, creating an implicit curriculum. We evaluate VPS on chess as a controlled testbed where reasoning steps

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

The Impossibility Triangle of Long-Context Modeling

arXiv:2605.05066v2 Announce Type: replace Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall). We formalize this trade-off within an Online Sequence Processor abstraction that unifies Transformers, state space models, linear recurrent networks, and their hybrids. Using the Data Processing Inequality and Fano's Inequality, we prove that any model satisfying Efficiency and Compactness can recall at most O(poly(d)/log V) key-value pairs from a sequence of arbitrary length, where d is the model dimension and V is the vocabulary size. We classify 52 architectures published before March 2026 into the triangle, showing that each achieves at most two of the three properties and that hybrid arc

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Memory in the LLM Era: Modular Architectures and Strategies in a Unified Framework

arXiv:2604.01707v3 Announce Type: replace Abstract: Memory emerges as the core module in the large language model (LLM)-based agents for long-horizon complex tasks (e.g., multi-turn dialogue, game playing, scientific discovery), where memory can enable knowledge accumulation, iterative reasoning and self-evolution. A number of memory methods have been proposed in the literature. However, these methods have not been systematically and comprehensively compared under the same experimental settings. In this paper, we first summarize a unified framework that covers existing representative agent memory methods from a high-level perspective. We then extensively compare representative agent memory methods on two long-term conversational benchmarks and an agentic memory benchmark, and examine the effectiveness of representative methods, providing a thorough analysis of those methods. As a byproduct of our experimental analysis, we also design a new memory method by exploiting modules in the exi

Source ↗
Showing 10351–10400 of 11035 signals
← Prev Page 208 of 221 Next →