EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Wed, 29 Jul 2026 22:29:27 +0000
MedCity News

Doctronic Acquires Summer Health to Expand into Pediatric Primary Care

Doctronic acquired Summer Health to expand into pediatric care and strengthen its AI-powered primary care platform. The post Doctronic Acquires Summer Health to Expand into Pediatric Primary Care appeared first on MedCity News .

Source ↗
technology Wed, 29 Jul 2026 22:12:43 +0000
MedCity News

Processa Acquisition Lands Immunology Drug That Could Rival a New Novartis Product

With the merger, Processa’s new lead asset is a Vidya Therapeutics immunology drug that offers pipeline-in-a-product potential addressing a target that has clinical and regulatory validation from a recently approved Novartis drug. Vidya’s molecule is a next-generation drug that could have advantages over Novartis’s pill. The post Processa Acquisition Lands Immunology Drug That Could Rival a New Novartis Product appeared first on MedCity News .

Source ↗
technology Wed, 29 Jul 2026 21:14:17 +0000
HN: education

AI in Education: Using ChatGPT Without Losing Critical Thinking

Article URL: https://www.notion4teachers.com/blog/the-ai-learning-divide-when-technology-supports-thinking-and-when-it-replaces-it Comments URL: https://news.ycombinator.com/item?id=49103173 Points: 3 # Comments: 1

Source ↗
technology Wed, 29 Jul 2026 20:46:13 +0000
MedCity News

Epic-First Doesn’t Mean Epic-Only: Where Hospitals Are Still Buying from Startups

A survey of 112 health system executives finds that Epic’s hold on hospital IT keeps getting stronger — but startups still have a shot, provided they solve a specific problem, prove strong ROI or offer a product that Epic is lagging behind on. The post Epic-First Doesn’t Mean Epic-Only: Where Hospitals Are Still Buying from Startups appeared first on MedCity News .

Source ↗
technology Wed, 29 Jul 2026 20:41:29 +0000
HN: education

The Miasma Theory of Education

Article URL: https://hollisrobbinsanecdotal.substack.com/p/the-miasma-theory-of-education Comments URL: https://news.ycombinator.com/item?id=49102810 Points: 4 # Comments: 1

Source ↗
technology Wed, 29 Jul 2026 13:45:41 +0000
HN: education

The Canvas Breach Exposed the Risk of One-Platform Education

Article URL: https://teachers-blog.com/the-canvas-breach-exposed-the-risk-of-one-platform-education/ Comments URL: https://news.ycombinator.com/item?id=49097460 Points: 4 # Comments: 2

Source ↗
technology Wed, 29 Jul 2026 13:39:00 +0000
MedCity News

A New Era of Interoperability, But Who Gets Left Behind?

CMS’s new framework sounds like a gentle opt-in. There’s no requirement to meet every standard on day one, just a commitment to work toward doing the right thing. That sounds great. But to actually participate in the network, you have to meet a baseline level of interoperability. The post A New Era of Interoperability, But Who Gets Left Behind? appeared first on MedCity News .

Source ↗
technology Wed, 29 Jul 2026 13:37:00 +0000
MedCity News

The Healthcare Cloud Maturity Model: The Five Stages Most Health Systems Skip

The most common and most expensive misunderstanding in healthcare cloud strategy is treating the cloud as a destination you either reach or you do not, when it is really a progression through five distinct stages, each one unlocking capabilities the stage before it could not support. The post The Healthcare Cloud Maturity Model: The Five Stages Most Health Systems Skip appeared first on MedCity News .

Source ↗
technology Wed, 29 Jul 2026 13:18:14 -0400
EdTech Mag (Higher)

Why AI‑Driven Observability Is Rising to the Top of Higher Ed IT Priorities

While much of the executive conversation around artificial intelligence focuses on transformation within business units, more organizations are looking inward. For IT teams, that means turning to AI not for strategic reinvention but to solve long‑standing operational challenges that impact uptime, user experience and cost efficiency. AI‑enabled observability and AI for operations (AIOps) is emerging as one of the most practical, high‑value starting points for institutional AI adoption. Across networking, infrastructure and security operations, IT leaders are looking to AI to help prevent…

Source ↗
technology Wed, 29 Jul 2026 11:24:57 -0400
EdTech Mag (K-12)

Networking Modernization Supports the Classrooms of Tomorrow

Technology defines the modern learning landscape. In the K–12 classroom, one-to-one device programs, BYOD policies, cloud-based learning platforms, digital displays and Internet of Things devices all are driving educational outcomes. In this environment, the network has quietly become the most critical piece of infrastructure. Success today is defined by the ability to support both current technology demands and tomorrow’s needs. With more devices coming onto the network all the time at New Trier Township High School District 203 in Illinois, CTO Michael Marassa saw the writing on the wall. “…

Source ↗
technology Wed, 29 Jul 2026 09:00:00 +0000
Tech & Learning

Tips For Finding and Keeping Talent: A Conversation with Nate Lichte

Nate Lichte, VP of K-12 Solutions and COO of Lebra and a presenter at the upcoming EdExec Summit, shares how to build strong school culture

Source ↗
technology Wed, 29 Jul 2026 09:00:00 +0000
eCampus News

Accreditation’s AI reckoning: Can colleges still prove students are learning?

Higher education is approaching a credibility crisis: Colleges may still be awarding degrees, but can they prove that students—not their AI tools—have actually learned? Accreditation has long depended on evidence that institutions define outcomes, assess student achievement, and use results to improve. The post Accreditation’s AI reckoning: Can colleges still prove students are learning? appeared first on eCampus News .

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models

arXiv:2607.18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable. Across three families, three benchmarks and six rungs, including in-the-wild chat logs, scaling closes the start-of-response knowledge gap up to $7\times$ while within-response knowledge degradation grows up to $39\times$. We trace that residual to one variable, the per-position disagreement $\delta = \log p_M - \log p_O$ against a stronger oracle, whose second moment splits exactly into bias$^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ and decoding risk $\mathrm{Var}[\delta]$. That split is an interpretability statement before it is a statistical one: the model's self-readable uncertainty $H(p_M)$ enters only the bias term, so the risk term has no model-readable component. Risk also takes a growing share of the squared error with scale, $31\%$ to $49\%$ from $1.7$B to $14$B. At a fabrication $H(p_M)$ relaxes within one token while risk persists up to $23\times$ longer, leaving a confident-but-pr

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection

arXiv:2604.18248v3 Announce Type: replace-cross Abstract: Current open-source prompt-injection detectors converge on two architectural choices: regular-expression pattern matching and fine-tuned transformer classifiers. Both share failure modes that recent work has made concrete. Regular expressions miss paraphrased attacks. Fine-tuned classifiers are vulnerable to adaptive adversaries: a 2025 NAACL Findings study reported that eight published indirect-injection defenses were bypassed with greater than fifty percent attack success rates. This work proposes seven detection techniques that each port a mechanism from a discipline outside LLM security: forensic linguistics, materials-science fatigue analysis, deception technology, local-sequence alignment from bioinformatics, mechanism design, spectral signal analysis, and taint tracking. Three are implemented in the prompt-shield v0.4.1 release (Apache 2.0) and evaluated across six datasets (deepset/prompt-injections, NotInject, LLMail-In

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

FlashEvaluator: Expanding Search Space with Parallel Sequence-Level Evaluation

arXiv:2603.02565v2 Announce Type: replace-cross Abstract: The Generator-Evaluator (G-E) framework generates K candidate sequences and uses an evaluator to select the highest-scoring one, which is widely used in recommender systems (RecSys) and natural language processing (NLP). Existing evaluators commonly score candidates independently. Although such evaluations can be batched, independent scoring neither models interactions among candidates nor eliminates repeated computation of request-level context and recurring candidate elements, causing the total evaluation work to grow approximately linearly with K. To handle with, we propose FlashEvaluator, a joint evaluator that scores all candidate sequences in a single forward pass. FlashEvaluator factorizes evaluation into shared request-level encoding, reusable candidate-side computation, sequence assembly by indexing, and cross-sequence interaction for setwise comparison. We call this request-local reuse scheme QKV-Cache: inspired by aut

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Controllable LLM Reasoning via Sparse Autoencoder-Based Steering

arXiv:2601.03595v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) exhibit human-like cognitive reasoning strategies (\eg backtracking, cross-verification) during the reasoning process, which improves their performance on complex tasks. Currently, reasoning strategies are autonomously selected by LRMs themselves. However, such autonomous selection often produces inefficient or even erroneous reasoning paths. To make reasoning more reliable and flexible, it is important to develop methods for controlling reasoning strategies. Existing methods struggle to control fine-grained reasoning strategies due to conceptual entanglement in LRMs' hidden states. To address this, we leverage Sparse Autoencoders (SAEs) to decompose strategy-entangled hidden states into a disentangled feature space. To identify the few strategy-specific features from the vast pool of SAE features, we propose SAE-Steering, an efficient two-stage feature identification pipeline. SAE-Steering first re

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Deep Delta Learning

arXiv:2601.00417v4 Announce Type: replace-cross Abstract: Transformer residual streams evolve through additive updates. Although a sufficiently expressive residual block can represent content replacement, standard architectures do not parameterize reading, comparison, and replacement as an explicit residual operation. We introduce Deep Delta Learning (DDL), a structured residual update that preserves the identity path while enabling target-seeking edits to the residual state. Each layer reads the current state along a learned direction, compares the resulting readout with a learned target, and writes back a gated rank-1 correction along the same direction. Closing the gate recovers the identity map, while fully opening it exactly overwrites the selected residual readout. We instantiate DDL with both scalar and expanded residual states. The expanded formulation provides multiple persistent value channels while keeping attention and MLP computation at the original model width, thereby se

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

SAGE: Stochastic Prompt Optimization via Agent-Guided Exploration

arXiv:2606.18902v2 Announce Type: replace Abstract: Context engineering has emerged as a primary lever for improving AI systems without parameter updates. Recent work showing that textual gradients do not function as real gradients motivates treating automatic prompt optimization (APO) as black-box search. We introduce SPO (Stochastic Prompt Optimization), a framework for stochastic search over prompt space, and compare three strategies of increasing sophistication: error-informed random search, a genetic algorithm with evolutionary operators, and SAGE (SPO via Agent-Guided Exploration), a multi-agent pipeline with diagnostic code execution. Across three benchmarks, no single strategy dominates; effectiveness depends on the interaction of landscape structure with error type. We further deploy SAGE on a mental-health chatbot under a continuous optimization paradigm, where it compounds eight cycles of individually-noisy A/B tests into a statistically robust gain in next-day retention. We

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report

arXiv:2606.03773v2 Announce Type: replace Abstract: High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English counterparts: they are often smaller, less carefully curated, weakly documented, and rarely validated through controlled training experiments. We introduce KletterMix, a high-quality German corpus for language model pretraining and annealing, designed as a reusable dataset artifact for the natural language processing and modeling community. KletterMix is built by translating a state-of-the-art English pretraining corpus into German while preserving document boundaries, metadata, source structure, and topical diversity. This construction yields a German corpus with the scale and diversity of a modern pretraining dataset, while enabling direct comparison to its English source. We document the dataset through a broad set of corpus-level analyses, including translation quality, documen

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments

arXiv:2605.30880v4 Announce Type: replace Abstract: World models for interactive text agents must typically be learned from observation-action trajectories alone. Specifically, the environment returns text observations after each action, but does not expose a ground-truth latent state nor an inspectable transition model.A research gap remains in how to induce executable code as a world model in this black-box setting for prediction and agent decision making. We introduce PatchWorld, a gradient-free framework that turns offline trajectories into executable Python world models through counterexample-guided code repair.Instead of predicting the next observation with a black-box model, PatchWorld induces symbolic belief-state programs whose action updates can be inspected, replayed, and locally patched. Across seven AgentGym environments, PatchWorld-Simple achieves the highest code-based decision-making score among evaluated methods (76.4% macro success in live one-step lookahead), matchin

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

A Study of Crosslinguistic Influence in Language Models

arXiv:2601.21587v2 Announce Type: replace Abstract: The sequential acquisition of languages inevitably leads to Crosslinguistic Influence (CLI), where the syntactic properties of a first language (L1) impact the processing of a second language (L2). While modern language models exhibit robust cross-lingual transfer, the exact mechanisms governing how language dominance, relative proficiency, and typological distance dictate structural interference warrant deeper investigation. In this work, we systematically investigate CLI in artificial learners by training simultaneous and sequential bilingual models across 15 typologically diverse L1s and varying the Step of Exposure (SoE), defined as the specific training step at which the L2 is introduced. Utilizing crosslinguistic structural priming, we decouple latent CLI into distinct positive and negative transfer rates. Our evaluations reveal a critical computational tradeoff: while increased L1 dominance (higher SoE) strongly amplifies the c

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Beyond Factual Accuracy: Evaluating Global Reasoning Integrity in RAG Systems with LogicScore

arXiv:2601.15050v5 Announce Type: replace Abstract: Current evaluation methods for Retrieval Augmented Generation (RAG) suffer from \textit{factual myopia}: they relentlessly emphasize factual accuracy yet neglect global logical integrity in long-form answer generation. This drives models to force unnatural connections, producing factually grounded yet logically incoherent responses with unaddressed gaps, ambiguous links, or redundant premises. To mitigate this, we present \textsc{LogicScore}, shifting from local, fact-by-fact assessment to rigorous global reasoning scrutiny. Grounded in Horn Rules, our approach integrates a backward verification mechanism to systematically evaluate three key reasoning dimensions: \textit{Completeness} (logically sound deduction), \textit{Essentiality} (non-redundancy), and \textit{Determinateness} (consistent answer entailment). Extensive experiments across three multi-hop QA datasets (HotpotQA, MusiQue, and 2WikiMultiHopQA) and over 20 LLMs (includin

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers

arXiv:2512.17351v2 Announce Type: replace Abstract: Understanding architectural differences in language models is challenging, especially at academic-scale pretraining (e.g., 1.3B parameters, 100B tokens), where results are often dominated by noise and randomness. To overcome this, we introduce controlled synthetic pretraining tasks that isolate and evaluate core model capabilities. Within this framework, we discover CANON LAYERS: lightweight architectural components -- named after the musical term "canon" -- that promote horizontal information flow across neighboring tokens. Canon layers compute weighted sums of nearby token representations and integrate seamlessly into Transformers, linear attention, state-space models, or any sequence architecture. We present 12 key results. This includes how Canon layers enhance reasoning depth (e.g., by $2\times$), reasoning breadth, knowledge manipulation, etc. They lift weak architectures like NoPE to match RoPE, and linear attention to rival SO

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

$M^2PO$: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation

arXiv:2510.13434v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human preferences is pivotal for Machine Translation (MT), yet current approaches are often hindered by misleading reward signals. Our analysis reveals that prevailing Quality Estimation (QE) models exhibit a systematic blind spot toward partial errors, specifically partial hallucinations and omissions, often favoring superficially fluent but unfaithful translations. To address this issue, we propose $M^2PO$ (Multi-Perspective Multi-Pair Preference Optimization), a data-centric framework for preference optimization in machine translation. First, to correct the bias toward fluency, $M^2PO$ uses a dual-perspective mechanism that decouples semantic fidelity from fluency and prioritizes faithfulness through a curriculum strategy. Second, after correcting this bias, partial errors fall between perfect and severely incorrect translations, making them difficult to learn through standard best-versus-

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation

arXiv:2510.09733v2 Announce Type: replace Abstract: Visual Retrieval-Augmented Generation (VRAG) has emerged as a promising paradigm for equipping Vision-Language Models (VLMs) with external visual evidence, enabling them to go beyond parametric knowledge when answering visually grounded questions. However, in such multi-image settings, VLMs still often suffer from visual hallucinations and struggle to accurately identify the question-relevant evidence needed for reliable reasoning. Existing methods usually lack an explicit cross-image evidence collection process, and also provide limited credit assignment when jointly optimizing perception and reasoning. To address this issue, we propose EVisRAG, an evidence-guided visual retrieval-augmented framework for multi-image reasoning. EVisRAG first observes the retrieved images, records question-relevant visual evidence from each image, and then performs reasoning and answer generation based on the aggregated evidence. We further introduce R

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Building Large-Scale English-Romanian Literary Translation Resources with Open Models

arXiv:2509.07829v4 Announce Type: replace Abstract: Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small open models remains an open problem, particularly for low-resource languages such as Romanian. We introduce the TinyFabulist Translation Framework (TF2), a unified framework for dataset creation, fine-tuning, and evaluation in English $\to$ Romanian literary translation. Building on DS-TF1-EN-3M, the largest collection of synthetic English fables to date, our pipeline first generates 15k high-quality Romanian references from the TF1 pool using a high-performing large language model (LLM). We then apply a two-stage fine-tuning process to a 12B-parameter open-weight model: (i) instruction tuning to capture genre-specific narrative style, and (ii) adapter compression for efficient deployment. Evaluation combines a five-dimension LLM-based rubric (accuracy, fluency, coherence, style, cultural adaptati

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning

arXiv:2508.00344v5 Announce Type: replace Abstract: Large Language Models (LLMs) have shown remarkable advancements in tackling agent-oriented tasks. Despite their potential, existing work faces challenges when deploying LLMs in agent-based environments. The widely adopted agent paradigm ReAct centers on integrating single-step reasoning with immediate action execution, which limits its effectiveness in complex tasks requiring long-term strategic planning. Furthermore, the coordination between the planner and executor during problem-solving is also a critical factor to consider in agent design. Additionally, current approaches predominantly rely on supervised fine-tuning, which often leads models to memorize established task completion trajectories, thereby restricting their generalization ability when confronted with novel problem contexts. To address these challenges, we introduce an adaptive global plan-based agent paradigm AdaPlan, aiming to synergize high-level explicit guidance w

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning

arXiv:2507.23541v5 Announce Type: replace Abstract: In medical scenarios, effectively retrieving external knowledge and leveraging it for rigorous logical reasoning is of significant importance. Despite their potential, existing work has predominantly focused on enhancing either retrieval or reasoning capabilities of the models in isolation, with little attention given to their joint optimization, which leads to limited coordination between the two processes. Additionally, current methods rely heavily on supervised fine-tuning (SFT), which can cause models to memorize existing problem-solving pathways, thereby restricting their generalization ability when confronted with novel problem contexts. Furthermore, while some studies have explored to improve retrieval-augmented reasoning in general domains via reinforcement learning, their reward function designs do not adequately capture the specific demands of the medical domain. To address these challenges, we introduce **Med-R$^3$**, a **M

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Towards Understanding the Cognitive Habits of Large Reasoning Models

arXiv:2506.21571v3 Announce Type: replace Abstract: Large Reasoning Models (LRMs), which autonomously produce a reasoning Chain of Thought (CoT) before producing final responses, offer a promising approach to interpreting and monitoring model behaviors. Inspired by the observation that certain CoT patterns -- e.g., ``Wait, did I miss anything?'' -- consistently emerge across tasks, we explore whether LRMs exhibit human-like cognitive habits. Building on Habits of Mind, a well-established framework of cognitive habits associated with successful human problem-solving, we introduce CogTest, a principled benchmark designed to evaluate LRMs' cognitive habits. CogTest includes 16 cognitive habits, each instantiated with 25 diverse tasks, and employs an evidence-first extraction method to ensure reliable habit identification. With CogTest, we conduct a comprehensive evaluation of 16 widely used LLMs (13 LRMs and 3 non-reasoning ones). Our findings reveal that LRMs, unlike conventional LLMs, n

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Localizing Persona Representations in LLMs

arXiv:2505.24539v4 Announce Type: replace Abstract: We present a study on how and where personas -- defined by distinct sets of human characteristics, values, and beliefs -- are encoded in the representation space of large language models (LLMs). Using a range of dimension reduction and pattern recognition methods, we first identify the model layers that show the greatest divergence in encoding these representations. We then analyze the activations within a selected layer to examine how specific personas are encoded relative to others, including their shared and distinct embedding spaces. We find that, across multiple pre-trained decoder-only LLMs, the analyzed personas show large differences in representation space only within the final third of the decoder layers. We observe overlapping activations for specific ethical perspectives -- such as moral nihilism and utilitarianism -- suggesting a degree of polysemy. In contrast, political ideologies like conservatism and liberalism appear

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data

arXiv:2503.07303v3 Announce Type: replace Abstract: Texts, whether literary or historical, exhibit structural and stylistic patterns shaped by their purpose, authorship, and cultural context. Formulaic texts, which are characterized by repetition and constrained expression, tend to differ in their \textit{information content} (as defined by Shannon) compared to more dynamic compositions. Identifying such patterns in historical documents, particularly multi-author texts like the Hebrew Bible, provides insights into their origins, purpose, and transmission. This study aims to identify formulaic clusters: sections exhibiting systematic repetition and structural constraints, by analyzing recurring phrases, syntactic structures, and stylistic markers. However, distinguishing formulaic from non-formulaic elements in an unsupervised manner poses a computational challenge, especially in high-dimensional, sample-poor data sets where patterns must be inferred without predefined labels. To addres

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Eye Tracking Based Cognitive Evaluation of Automatic Readability Assessment Methods

arXiv:2502.11150v5 Announce Type: replace Abstract: Automatic methods for scoring text readability have been studied for over a century, and are widely used in research and in user-facing applications in many domains. Thus far, the development and evaluation of such methods have primarily relied on two types of offline human behavioral data, performance on reading comprehension tests and ratings of text readability levels. In this work, we instead focus on a fundamental and understudied aspect of readability, real-time reading ease, captured with online reading measures using eye tracking. We introduce a new cognitive evaluation framework for readability scoring methods that quantifies their ability to account for reading ease, while controlling for content variation across texts. Applying this evaluation to prominent traditional readability formulas, NLP-based methods, commercial systems used in education, and frontier LLMs suggests that they are all poor predictors of English reading

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series

arXiv:2607.25947v1 Announce Type: cross Abstract: Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable promise in general-purpose time-series QA, they remain poorly equipped to model the sparsity, asynchrony, and irregular sampling patterns of clinical observations. To fill this gap, we propose ClinPRISM, a cost-effective multimodal LLM reasoning framework for question answering over ICTS data. First, we devise an irregularity-aware multi-scale encoder to capture sparse clinical evidence at diverse temporal scales. Then, we propose a temporal evidence distiller to integrate representations across these scales and compress them into a small number of LLM-compatible tokens. Moreover, we introduce a progressive alignment strategy that sequentially aligns the irregular trajectories with the LLM's textual embedding space.

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

arXiv:2607.25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no inference-time model access. Our target is an "evaluation-awareness" latent-linearly readable and steerable in recent work-whose control would threaten the validity of safety evaluations if models behave differently when they detect being tested. Adapting Fluent Dreaming / EPO with a negated feature term (GCG-style token optimization plus a self-cross-entropy fluency regularizer, swept over a fluency weight), we suppress the latent under five target constructions-a CAA direction, a subspace norm, an SAE feature, a single MLP neuron, and a behavioral logit-on Llama-3.2-3B and Llama-3.1-8B. The latent is robustly suppressible ($z\approx-7$), and a causally-validated Llama Scope SAE feature can be fully and selectively turned o

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

arXiv:2607.25886v1 Announce Type: cross Abstract: Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capability gaps, designing and validating training-data strategies, and learning from checkpoint feedback. Can LLM agents automate this loop? Existing benchmarks entangle research decisions with optimization, serving, evaluation, and systems implementation, obscuring agents' research capability. We introduce RSIBench-Data, a controlled benchmark of LLM agents as data-centric researchers with a fixed post-training stack. Agents iteratively revise training-data strategies for a fixed target model; training and serving use Tinker-backed services, official evaluation runs through Harbor and E2B sandboxes, and budgets are fixed across agents. We evaluate four frontier agents on six benchmarks across software engineering, terminal use, scientific question answering, and mathematics. Agents demonstra

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Stemma: Induced Decision Regions Reveal LLM Provenance

arXiv:2607.25880v1 Announce Type: cross Abstract: LLM provenance testing asks whether a suspect LLM belongs to the same lineage as a source. Existing black-box methods largely infer this relationship from response-level characteristics, but these characteristics may shift under adaptation or deployment even when the underlying meaning remains unchanged, weakening the reliability of provenance evidence. To address this limitation, we introduce induced decision regions by mapping open-ended outputs into a finite decision space, thereby abstracting away surface-form variation and reframing provenance testing as measuring the inheritance of decision regions. Empirical analysis shows that the source's induced regions are preserved more strongly in related models than in unrelated models. Building on this signal, we propose Stemma, a practical black-box LLM fingerprinting method that operationalises stability, robustness, and specificity as complementary probe-selection principles for reliab

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

arXiv:2607.25672v1 Announce Type: cross Abstract: We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Localized Adaptation Reveals Distinct Learning Signatures in Transformers

arXiv:2607.25663v1 Announce Type: cross Abstract: Transformer adaptation is typically distributed across model depth, even when the intended change is narrow. We investigate how adaptation site shapes what a model learns, how well that learning generalizes, and how selectively it is applied. We introduce a controlled benchmark spanning five objectives (lexical binding, factual association, behavioral policy learning, causal mapping, and procedural reasoning) and define each objective's "adaptation geometry" as its profile of acquisition, transfer, and boundedness under full-stack and early-, middle-, or late-layer LoRA. The objectives exhibit distinct geometries. Lexical binding favors early-layer adaptation for acquisition and boundedness but requires broader updates for transfer; factual association favors later layers among localized adapters; behavioral learning separates late-layer action acquisition from middle-layer policy gating; and causal and procedural transfer benefit most

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

arXiv:2607.25642v1 Announce Type: cross Abstract: Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have accelerated progress toward practical ``one-sentence image editing" systems. This survey presents a systematic taxonomy and comprehensive review of IIE research, structured around five core dimensions: (1) task definition and hierarchical categorization of editing operations, (2) methodologies for training data construction, (3) architectural evolution from GAN-based to diffusion and autoregressive paradigms, (4) standardized evaluation metrics and benchmark development, and (5) introduction of commercial solutions. Our analysis shows critical technological milestones across model generations. We further propose a Comprehensive, in-Depth, and Diagnostic benchmark for IIE task (CDD-IIE Bench), which can rigorously assess the multiple aspects of

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations

arXiv:2607.25634v1 Announce Type: cross Abstract: We present AIriskEval-edu Demo, a platform that audits the pedagogical quality of instructional explanations and provides explainable audit results. The platform evaluates an explanation against a rubric covering five dimensions of pedagogical risk: factual accuracy, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias. For each dimension, it returns a binary decision and a confidence score. Detected risks also include a natural-language rationale and, except for Depth and Completeness, a localized evidence span. The platform integrates GPT-5.5 through an external API and a self-hosted Llama 3.1 8B evaluator that runs on consumer-grade GPUs. The local evaluator is fine-tuned on AIriskEval-edu, a dataset of K-12 instructional explanations with risk and explainability annotations. The platform operates in two modes: in AI mode, both evaluators assess stored explanations generated under six simul

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

arXiv:2607.25614v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantially degrade performance on general tasks. We propose MemSFT, which mitigates the alignment tax by decoupling domain specialization from backbone parameter updates through a plug-and-play parametric memory. The memory is trained to imitate the behavior of a non-parametric retriever operating over domain data, thereby memorizing knowledge and patterns that would otherwise be accessed through retrieval. Once trained on a specific domain, the memory can be reused across LLMs of different sizes. During generation, a learned router dynamically fuses the output distributions of the memory and backbone at each decoding step, allowing domain expertise to be invoked selectively. Across biology, geoscience, and law, evaluations with models ranging from Qwen3-8B to Qwen3-

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs

arXiv:2607.25600v1 Announce Type: cross Abstract: Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evidence and unnecessary computation. We investigate whether verbalized confidence from black-box language models can serve as an actionable signal for retrieval routing. Our method, BeyondUncertainty, first elicits a structured provisional answer and confidence estimate, then applies a model-specific threshold selected on held-out validation data and frozen before test evaluation. Low-confidence questions receive top-$5$ TF--IDF retrieval followed by a second answer call, whereas high-confidence questions return the provisional answer directly. We evaluate 27,000 policy instances across six QA benchmarks, three model families, and three retrieval policies. BeyondUncertainty achieves $0.483$ mean token-level F1, compared with $0.467$ for always retrieval and $0.401$ for no retrieval, while reducing retrie

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

arXiv:2607.25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than tested. We performed a retrospective forensic reproducibility audit of a preserved chest-radiograph vision-language model (VLM) pilot; no model was called again and no image or report was newly annotated. We traced prompt bindings, DICOM metadata, output completeness, label extraction, matched analyses, and release propagation. Of 300 planned model-prompt calls, 297 yielded nonempty reports. Sixty Claude calls labeled A/B were executed with the same C prompt. The 30 studies represented 28 patients. Four MONOCHROME1 images were rendered without required polarity inversion, dataset split membership was not retained, and the unvalidated extractor truncated five reports to 4000 characters. Reconstructing one common co

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

arXiv:2607.25485v1 Announce Type: cross Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarke

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

arXiv:2607.25398v1 Announce Type: cross Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use horizon. We present HANDBOOK.md, a benchmark of 65 agentic tasks modeled on how enterprise employees follow company handbooks. Each task places an agent in a self-contained company environment, a file workspace together with mock email, chat, calendar, issue-tracking, and commerce services exposed over the Model Context Protocol, and instructs it to carry out routine professional work governed by an expert-written standard operating procedure of 20 to 124 pages. Tasks span five domains (finance, medical

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Data Quality Profiling at Scale with Progressive Sampling: A Benchmark for Data-Centric AI Pipelines

arXiv:2607.25356v1 Announce Type: cross Abstract: Data quality profiling -- computing missing-value rates, duplicate fractions, outlier densities, and functional-dependency violations -- is foundational for data-centric AI pipelines, yet exhaustive scans over millions of rows are prohibitively slow for near-real-time monitoring. Progressive sampling is the standard alternative; the open question is which strategy best preserves profile fidelity at scale. We benchmark nine sampling strategies -- blind (random uniform, geometric, Yamane, cluster) and proxy-guided (Metropolis-Hastings, DAG, stratified by column type or quality score, importance-weighted) -- on three real-world datasets (NYC 311, NYPD arrests, UCI Adult; up to 500K rows), an IoT sensor stream (2.3M rows), two ultra-large real datasets including Ultra-Marathon Running (up to 7.4M rows), and synthetic data scaled to 5x10^6 rows. Contrary to the assumption sharpens estimates, blind representative samplers dominate uniformly.

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management

arXiv:2607.25340v1 Announce Type: cross Abstract: The same episode of atrial fibrillation is a minor finding in a healthy adult and grounds for anticoagulation in an elderly patient with hypertension: identical signal, opposite decision. Naming the rhythm is only the start; what determines a patient's outcome is the judgement that follows -- what the arrhythmia is across the whole record, what it means for this patient, and what should be done about it. Recent work pairing large language models with the ECG stops short of this, reading one recording without assembling a patient-level finding; and agentic systems built around it either receive the arrhythmia a device has already detected or target a different diagnostic task, stopping before the decision this task requires. We formulate patient-level arrhythmia decision support as a task and present Cardiologent, a multi-agent system that spans it from detection to decision. An agent for each signal -- a single ECG lead and the photople

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

arXiv:2607.25294v1 Announce Type: cross Abstract: Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes, or web pages. We introduce CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning. CLBench-V combines converted public benchmarks with newly constructed datasets spanning domains such as science, finance, long-document understanding, spatial reasoning, a

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation

arXiv:2607.25140v1 Announce Type: cross Abstract: This paper studies the behavior of language models in a multi-agent crowd simulation, focusing on how affect propagates among agents that perceive and appraise one another. Each agent perceives its neighbors through visual, auditory, and tactile channels, then appraises these perceptions in light of its prompted personality profile, memory, current affective state, and situational context. Appraisal is carried out by an LLM, which updates the agent's internal affective state and selects its outward expression. The architecture contains no hand-authored mechanism for directly transferring affective state between agents; instead, inter-agent influence arises through the perception-appraisal-expression loop. The agent representation draws on the Big Five personality model and Russell's circumplex model of affect. To limit latency, low-level steering and navigation are handled by a conventional crowd simulator operating independently of the

Source ↗
technology Wed, 29 Jul 2026 00:00:00 -0400
arXiv cs.CL

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

arXiv:2607.25091v1 Announce Type: cross Abstract: The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen (model, corpus) configurations were trained using Proximal Policy Optimization (PPO). The experiments included Pythia-70M, 160M, 410M and SmolLM2-135M, 360M on the TinyStories, CNN/DailyMail, and Wikitext-103 corpora. Three reproducible failure modes were identified in small-scale language models: silent LoRA parameter freezing in standard PEFT/TRL pipelines, numerical overflow in importance ratios when using bfloat16, and catastrophic policy collapse due to reward-model error. These issues were addressed using a merge-and-reinitialize adapter technique, float32 precision during PPO updates, and a three-layer safety mechanism comprising reward whitening, importance-ratio

Source ↗
Showing 1–50 of 10876 signals
← Prev Page 1 of 218 Next →