EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

regulation Thu, 20 Aug 2026 15:48:25 -0400
K-12 Dive

Judge deals blow to Trump administration’s gutting of Teen Pregnancy Prevention Program

An "abstinence-only approach" led to "unreasonable or unexplained" grant cancellations, said the judge, but he stopped short of reinstating funds.

Source ↗
regulation Thu, 20 Aug 2026 14:30:00 +0000
The 74

Opinion: Head Start Changed My Life. The Proposed Reforms Would Gut the Program

I am a product of Head Start. Before I understood federal funding, program performance standards or the complicated systems that make Head Start possible, I understood what it meant to be a Head Start child. Years later, I returned, first as a center director and later as a regional director of a Head Start program […]

Source ↗
technology Thu, 20 Aug 2026 13:54:00 +0000
MedCity News

AI in Pharmacovigilance: Why Governance Will Define Success

As adoption accelerates, organizations must ensure that the use of AI strengthens and not weakens accountability and patient safety. The post AI in Pharmacovigilance: Why Governance Will Define Success appeared first on MedCity News .

Source ↗
behavior Thu, 20 Aug 2026 13:20:57 +0000
District Admin

Utah’s largest school district just welcomed students back for the final time. Here’s how it will soon split into 3.

By next fall, the Alpine School District—the state’s largest district—will split into three: the newly named Lake Mountain, Aspen Peaks and Timpanogos school districts. The post Utah’s largest school district just welcomed students back for the final time. Here’s how it will soon split into 3. appeared first on District Administration .

Source ↗
behavior Thu, 20 Aug 2026 13:17:53 +0000
District Admin

Florida school board at the center of culture wars loses conservative majority

Democrats say the flip in Sarasota County indicates voters want education to move on from political fights over social issues, while conservatives point to their wins. The post Florida school board at the center of culture wars loses conservative majority appeared first on District Administration .

Source ↗
regulation Thu, 20 Aug 2026 12:30:00 +0000
The 74

Opinion: Dismantling the Dyslexia Industrial Complex

The rapid rise of dyslexia in public discourse reflects a long-overdue shift in education. For decades, students with persistent reading difficulties were misunderstood, mislabeled or overlooked entirely. Increased awareness has helped dispel harmful myths, fueled the Science of Reading movement, and prompted schools to take literacy instruction more seriously. The urgency could not be greater. […]

Source ↗
technology Thu, 20 Aug 2026 11:58:27 -0400
EdTech Mag (K-12)

When In-House Break/Fix Works (and When To Call for Backup)

Virtually every school district has it: A closet, back room or office corner cluttered with broken devices that nobody has time to fix. What felt manageable when a device fleet was a year or two old can quickly become overwhelming as Chromebooks, laptops and tablets age and begin needing maintenance or repairs. When K–12 IT teams are already working at maximum capacity and wearing multiple hats, device repairs and hardware maintenance can be one of the tasks that gets pushed to the backburner. But hardware downtime can hinder classroom learning and teacher efficiency, and district budgets can…

Source ↗
regulation Thu, 20 Aug 2026 10:30:00 +0000
The 74

Student Journalists: AI Is Changing Our Work — And Not For the Better

In the nearly four years since generative artificial intelligence began colonizing the academic lives of teens, it has changed the experience of school for millions of young people. Perhaps no group has watched their reality shift more than student journalists. Just as the technology has rewired the relationship between young writers-in-training and their teachers, it […]

Source ↗
behavior Thu, 20 Aug 2026 10:00:00 +0000
eSchool News

What it really takes to accelerate adolescent literacy across our district

When people hear about the literacy gains we’ve seen in our middle schools, they often ask a simple question: What program did you use?

Source ↗
behavior Thu, 20 Aug 2026 09:15:00 +0000
Getting Smart

How Empathy Interviews Transform Stakeholder Engagement into Shared Ownership

What if the most powerful tool in school improvement is a single, unhurried conversation? In this post, Jennifer Poon and Doannie Tran make the case for empathy interviews as the essential first step in genuine community co-creation, and explain why it matters deeply who conducts them. Drawing on their work with Burlington School District, they show how shifting the interviewer from consultant to community member transforms not just the data collected, but the relationships and shared ownership that follow. This is a must-read for any education leader serious about building strategies that reflect the full truth of their community. The post How Empathy Interviews Transform Stakeholder Engagement into Shared Ownership appeared first on Getting Smart .

Source ↗
technology Thu, 20 Aug 2026 09:00:00 +0000
Tech & Learning

8 of the Best Tools To Teach Math

These are the best tool to teach math across a range of ages and abilities.

Source ↗
behavior Thu, 20 Aug 2026 08:07:58 +0000
HN: online learning

LinkVault – A local-first archive for online learning

Article URL: https://github.com/Howard-Starfield/LinkVault-Linkedin-Learning-Courses-Downloader Comments URL: https://news.ycombinator.com/item?id=49371769 Points: 1 # Comments: 0

Source ↗
audience Thu, 20 Aug 2026 07:00:00 +0000
Inside Higher Ed

2 Milligan University Cyclists Killed in Crash

2 Milligan University Cyclists Killed in Crash Johanna Alonso Thu, 08/20/2026 - 03:00 AM Byline(s) Johanna Alonso

Source ↗
audience Thu, 20 Aug 2026 07:00:00 +0000
Inside Higher Ed

College Leaders Get a Boost From Corporate Board Service

College Leaders Get a Boost From Corporate Board Service Josh Moody Thu, 08/20/2026 - 03:00 AM Multiple college presidents are paid handsomely to sit on corporate boards in addition to the demands of their job. Is the practice beneficial to institutions or a personal windfall? Byline(s) Josh Moody

Source ↗
audience Thu, 20 Aug 2026 07:00:00 +0000
Inside Higher Ed

Cybersecurity Threat Delays Start of Classes at UT San Antonio

Cybersecurity Threat Delays Start of Classes at UT San Antonio kathryn.palmer… Thu, 08/20/2026 - 03:00 AM Although the university says no data was compromised, experts say the disruption still has consequences. Byline(s) Kathryn Palmer

Source ↗
audience Thu, 20 Aug 2026 07:00:00 +0000
Inside Higher Ed

Charges Filed in Penn State Cocaine Trafficking Case

Charges Filed in Penn State Cocaine Trafficking Case Olivia.sanchez Thu, 08/20/2026 - 03:00 AM Byline(s) Olivia Sanchez

Source ↗
audience Thu, 20 Aug 2026 07:00:00 +0000
Inside Higher Ed

With New AI Requirements and Courses, Colleges Eye AI Fluency

With New AI Requirements and Courses, Colleges Eye AI Fluency Emma Whitford Thu, 08/20/2026 - 03:00 AM Employers increasingly seek graduates who know how to use artificial intelligence. Through new courses and curricula, several institutions are ensuring that their students will fit the bill. Byline(s) Emma Whitford

Source ↗
audience Thu, 20 Aug 2026 07:00:00 +0000
Inside Higher Ed

The Key Podcast: Newsroom Federal Policy Analysis

The Key Podcast: Newsroom Federal Policy Analysis sara.custer@in… Thu, 08/20/2026 - 03:00 AM Byline(s) IHE Staff

Source ↗
audience Thu, 20 Aug 2026 07:00:00 +0000
Inside Higher Ed

Broward College President Fends Off Firing Attempt

Broward College President Fends Off Firing Attempt Josh Moody Thu, 08/20/2026 - 03:00 AM Byline(s) Josh Moody

Source ↗
audience Thu, 20 Aug 2026 07:00:00 +0000
Inside Higher Ed

UNC System Answers McMahon’s ‘Call to Action’

UNC System Answers McMahon’s ‘Call to Action’ Katherine Knott Thu, 08/20/2026 - 03:00 AM Byline(s) Katherine Knott

Source ↗
audience Thu, 20 Aug 2026 07:00:00 +0000
Inside Higher Ed

Rethinking Calculus for Future Engineers

Rethinking Calculus for Future Engineers Joshua.Bay Thu, 08/20/2026 - 03:00 AM The University of Michigan is redesigning foundational math to connect calculus to real-world engineering—and get students to apply it earlier. Byline(s) Joshua Bay

Source ↗
audience Thu, 20 Aug 2026 05:00:00 -0400
Higher Ed Dive

DOJ charges 17 in Iran-backed hacking campaign against US colleges, others

Officials allege the Mabna Institute was behind an effort to steal research from American universities, companies and government agencies.

Source ↗
regulation Thu, 20 Aug 2026 05:00:00 -0400
K-12 Dive

Efforts to keep ICE off of school grounds ramp up

Immigration enforcement activities on or around schools remain a concern for many educators at the start of the 2026-27 school year.

Source ↗
need Thu, 20 Aug 2026 05:00:00 +0000
Hechinger Report

How Florida quietly removed climate change content from textbooks

Had Florida’s high school marine biology students read the original draft of their textbook, they would have learned about research that found climate change is killing at least 150,000 people per year, and another 1 degree Celsius of warming would double that. Instead, they learned that “human health will also be affected by the increase […] The post How Florida quietly removed climate change content from textbooks appeared first on The Hechinger Report .

Source ↗
behavior Thu, 20 Aug 2026 00:00:00 GMT
EdSurge

Department of Education Issues Long-Awaited Edtech Guidance for States and Districts

The letter emphasizes outcomes and evidence but stops short of issuing federal regulations.

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks

arXiv:2606.21654v2 Announce Type: replace-cross Abstract: Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across multiple objectives. We study this gap with ChainWorld, which composes atomic OSWorld tasks into long horizon desktop workloads through directional compatibility search while preserving the source evaluators. The resulting workload contains 347 chains of length two to four and compares two renderings of the same task sequence. In single turn evaluation, all tasks are presented together in one prompt. In multi turn evaluation, tasks are revealed one at a time. Across four current computer use agents, maximum chain completion is 31%. Multi turn evaluation improves completion for three models, but both protocols remain challenging. The two protocols also expose different failure profiles. Single turn failures concentrate on artifact precision, while multi turn failures more often reflect session

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining

arXiv:2606.16246v3 Announce Type: replace-cross Abstract: As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shifting toward a data-constrained, compute-abundant regime that demands productive multi-epoch training on fixed corpora. Standard autoregressive (AR) pretraining overfits severely in this setting, reaching its optimum early and then continuously deteriorating. We investigate training-time data augmentation as a regularizer to mitigate this overfitting and enable productive training for hundreds of epochs on the same data. We introduce three orthogonal categories of augmentation for AR pretraining: token-level noise (masking, random replacement), sequence permutations (right-to-left prediction, Fill-in-the-Middle), and target offset prediction ($x_{t+i}$ for $i > 1$). Through systematic ablations, we find that individual augmentations delay overfitting and lower validation loss relative

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

arXiv:2605.09874v2 Announce Type: replace-cross Abstract: Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long videos, relevant information is sparsely distributed across hours or days, making memory a fundamental challenge: models must accumulate information over time, recall prior states, track temporal order, and abstract recurring patterns. However, existing week-long video benchmarks are primarily designed for perception and recognition, such as moment localization or global summarization, rather than reasoning that requires integrating evidence across multiple days. To address this gap, we introduce EgoMemReason, a comprehensive benchmark for week-long egocentric video understanding through memory-driven reasoning. EgoMemReason evaluates three complementary memory types: entity memory, tracking how object states evolve and change across d

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition

arXiv:2605.02782v2 Announce Type: replace-cross Abstract: Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility of improving performance by conditioning on additional clinical context at inference time, but it is unclear whether these models can make use of such information. We introduce a benchmark built on the Speech Accessibility Project (SAP) dataset that tests whether diagnosis labels, clinician-derived speech ratings, and progressively richer clinical descriptions improve transcription accuracy for dysarthric speech. Across matched comparisons on nine models, we find that current models do not meaningfully use this context: diagnosis-informed and clinically detailed prompts yield negligible improvements and often degrade word error rate. We complement the prompting analysis with context-dependent fine-tuning, showing that LoRA adaptation with a mixture of clinical prompt formats achiev

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

SkillNet: Create, Evaluate, and Connect AI Skills

arXiv:2603.04448v2 Announce Type: replace-cross Abstract: Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systematic accumulation and transfer of skills. Without a unified mechanism for skill consolidation, agents frequently ``reinvent the wheel'', rediscovering solutions in isolated contexts without leveraging prior strategies. To address this challenge, we introduce SkillNet, an open infrastructure for creating, evaluating, and organizing AI skills at scale. SkillNet structures skills within a unified ontology that supports creating skills from heterogeneous sources, establishing rich relational connections, and performing multi-dimensional evaluation across Safety, Completeness, Executability, Maintainability, and Cost-awareness. Our infrastructure integrates a repository of over 600,000 skills, an interactive platform, and a versatile Python toolkit. Experiments on ALFWorld, WebShop, and ScienceWorld

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum

arXiv:2601.01684v2 Announce Type: replace-cross Abstract: While dense retrieval models have been the standard for state-of-the-art information retrieval, their deployment is often constrained by high memory requirements and reliance on GPU accelerators for vector similarity search at scale. Learned sparse retrieval offers a compelling alternative by enabling efficient search via inverted indices, yet it has historically received less attention than dense approaches. In this paper, we introduce LACONIC, a family of learned sparse retrievers based on the Llama3 architecture (1B, 3B, and 8B). We propose a streamlined two-phase training curriculum consisting of (1) weakly supervised pre-finetuning to adapt causal LLMs for bidirectional contextualization and (2) high-signal finetuning using curated hard negatives. Our results demonstrate that LACONIC effectively bridges the performance gap with dense models: the 8B variant achieves a state-of-the-art 60.2 nDCG@10 on the MTEB Retrieval bench

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Jailbreaking in the Haystack

arXiv:2511.04707v2 Announce Type: replace-cross Abstract: Recent advances in long-context language models (LMs) have enabled million-token inputs, expanding their capabilities across complex tasks like computer-use agents. Yet, the safety implications of these extended contexts remain unclear. To bridge this gap, we introduce NINJA (short for Needle-in-haystack jailbreak attack), a method that jailbreaks aligned LMs by appending benign, model-generated content to harmful user goals. Critical to our method is the observation that the position of harmful goals play an important role in safety. Experiments on standard safety benchmark, HarmBench, show that NINJA significantly increases attack success rates across state-of-the-art open and proprietary models, including LLaMA, Qwen, Mistral, and Gemini. Unlike prior jailbreaking methods, our approach is low-resource, transferable, and less detectable. Moreover, we show that NINJA is compute-optimal -- under a fixed compute budget, increasin

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers

arXiv:2606.22361v2 Announce Type: replace Abstract: Why do multilingual language models sometimes generate in the wrong language, and why is this so hard to fix? We introduce Language Identity Head Ablation (LIHA), a causal intervention that zeros each attention head individually and measures the resulting language switch rate across a parallel dataset of 2,700 prompt-language pairs spanning seven languages. Applied to GPT-2, LIHA identifies a small set of first-token broadcaster heads - led by L6H1 (switch rate 0.32, 3.23 $\sigma$ above the population mean) - that attend persistently to the first prompt token, propagating its language signal throughout generation. Compensatory redistribution when heads are ablated is statistically significant (p < $10^{-5}$) and follows a directional, hierarchical pattern: compensation always recruits heads in layers above the ablated head, suggesting a feedforward cascade rather than global diffusion. To probe how training regime shapes these circuit

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models

arXiv:2606.00919v2 Announce Type: replace Abstract: Large language models (LLMs) have seen widespread adoption across various domains, yet their reliability is frequently undermined by hallucinations - responses that are plausible-sounding but factually incorrect. In high-stakes domains, these errors can reduce trust and introduce real-world risk. To address this challenge, we present a parameter-efficient approach that uses soft prompts to mitigate hallucinated content and promote responsible abstention in generative question-answering (QA) tasks. Our method, called Responsible Contrastive Soft Prompting (RCSP), uses a composite loss to train soft prompts that balance three goals: suppressing hallucinatory content, encouraging abstention under uncertainty, and preserving or improving factual recall. To achieve these goals, we incorporate contrastive loss, curriculum learning, and KL regularization into our training mechanism. We evaluate our approach on five diverse generative QA data

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law

arXiv:2605.30497v2 Announce Type: replace Abstract: RAG-based legal assistants have been growing in popularity, but LLM hallucinations remain a key issue and potentially undermines justice. While benchmarks have been developed to evaluate progress, many rely on synthetic queries rather than realistic legal scenarios. Moreover, Canadian law remains underrepresented in existing evaluations. To address this gap, we introduce CanLegalRAGBench, a Canadian legal QA benchmark based on realistic queries and expert-annotated answers grounded in case law. Our evaluation shows that retrieval performance is sensitive to design choices and that open-source embedding models are competitive with closed source models. However, it also reveals the limitation of automatic evaluations that penalize systems for retrieving alternative relevant documents. We also find that generated answers often diverge from gold responses, either with hallucinations or by producing overly detailed or irrelevant content, w

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Key Coverage Matters: Semi-Structured Extraction of OCR Clinical Reports

arXiv:2605.09440v2 Announce Type: replace Abstract: Clinical reports are often fragmented across healthcare institutions because privacy regulations and data silos limit direct information sharing. When patients seek care at a different hospital, they often carry paper or scanned reports from prior visits. This hinders EHR integration and longitudinal review, and downstream applications that depend on more complete patient records, such as patient management, follow-up care, real-world studies, and clinical-trial matching. Although OCR can digitize such reports, reliable extraction remains challenging because clinical documents are heterogeneous, OCR text is noisy, and many healthcare settings require low-cost on-premise deployment. We formulate this problem as canonical key-conditioned extractive question answering over OCR-derived clinical reports. Because the key fields are neither fixed nor known in advance, the key space is open. We maintain a canonical key inventory through itera

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering

arXiv:2605.05962v2 Announce Type: replace Abstract: This paper addresses end-to-end geospatial question answering over multilingual toponymic data. We introduce a bilingual (Russian-Tatar) dataset of 9,688 toponyms with linguistic, etymological, and coordinate information (93.1 percent georeferenced). Based on this, we construct about 39,000 question-context-answer triples with guaranteed answer localization. Our architecture combines a hybrid retriever (dense semantic indexing with multilingual-e5-large plus geospatial filtering/ranking using KD-trees and haversine distance) and an extractive reader fine-tuned on transformer models. On 500 test queries, hybrid search achieves Recall@1 = 0.988, Recall@5 = 1.000, MRR = 0.994, significantly outperforming BM25 and spatial-only methods. Among readers (RuBERT, XLM-RoBERTa-large, T5-RUS), XLM-RoBERTa-large gives best results: EM = 0.992, F1 = 0.994. RuBERT models fail on coordinate questions due to tokenization artifacts, but simple post-pro

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports

arXiv:2605.03103v2 Announce Type: replace Abstract: Semi-structured information extraction (IE) from OCR-derived clinical reports is crucial for efficiently reconstructing patients' longitudinal medical histories. In practice, this scenario commonly involves three tasks: (i) field-header (key) discovery, (ii) key-conditioned question answering (QA), and (iii) end-to-end key-value pair extraction. However, existing evaluations often under-model two factors: heterogeneous and incompletely known key representations, and OCR-induced noise. This makes it difficult to assess model robustness in real-world settings. We present MedStruct-S, a benchmark specifically designed to evaluate these tasks under unknown keys and OCR noise. MedStruct-S contains 3,582 annotated real-world clinical report pages. Using MedStruct-S, we benchmark two representative paradigms: encoder-only sequence labeling with post-processing and decoder-only structured generation, covering four encoder-only and five decode

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

An Information-theoretic Propagation Denoising and Fusion Framework for Fake News Detection

arXiv:2605.02259v2 Announce Type: replace Abstract: Incomplete propagation data significantly hinders robust fake news detection. Recent approaches leverage large language models to simulate missing user interactions via role-playing, thereby enriching propagation with synthetic signals. However, such propagation data is intrinsically unreliable, and directly fusing it can lead to biased representations and limited detection performance. In this paper, we alleviate the unreliability of synthetic propagation from the mutual information perspective and propose a novel information-theoretic propagation denoising and fusion (InfoPDF) framework to learn effective representations from both real and synthetic propagation. Specifically, we first generate attribute-specific synthetic propagation using large language models. Then we model each synthetic propagation graph as a probabilistic latent distribution to guide reliability-aware adaptive fusion with real propagation. During training, we d

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Predicting the Benefit of Retrieval Augmentation in Open-Domain Question Answering

arXiv:2604.07985v3 Announce Type: replace Abstract: While retrieval augmented generation has become a common approach for enhancing question answering systems, retrieval is not universally advantageous. We study the problem of predicting whether incorporating external retrieved information is likely to improve response quality for a given question. To this end, we evaluate a range of prediction methods that are based on retrieval signals, answer characteristics, and semantic consistency between generated responses and retrieved passages. We further devise a predictor that probes the LLM's internal state. Its prediction performance significantly narrows the performance gap between post-generation methods which are computationally demanding and pre-generation (post-retrieval) methods. We use the prediction methods to devise a selective retrieval framework that dynamically chooses between retrieval and non-retrieval generation modes per question. Experimental results demonstrate that sele

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't

arXiv:2604.06422v2 Announce Type: replace Abstract: Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if models adhere to their introspective reasoning are central challenges for trustworthy deployment. To study this, we introduce the Graded Color Attribution (GCA) dataset, a controlled benchmark designed to elicit decision rules and evaluate participant faithfulness to these rules. GCA consists of line drawings that vary pixel-level color coverage across three conditions: world-knowledge recolorings, counterfactual recolorings, and shapes with no color priors. Using GCA, we ask both VLMs and human participants to state a threshold rule: the share of an object's pixels that must be a given color for the object to receive that color label. We then compare these rules with their subsequent color attribution decisions. Our findings reveal that models systematically violate their own introspective rules. F

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Self-Improvement of Large Language Models: A Technical Overview and Future Outlook

arXiv:2603.25681v2 Announce Type: replace Abstract: As large language models (LLMs) continue to advance, improving them solely through human supervision is becoming increasingly costly and limited in scalability. As models approach human-level capabilities in certain domains, human feedback may no longer provide sufficiently informative signals for further improvement. At the same time, the growing ability of models to make autonomous decisions and execute complex actions naturally enables abstractions in which components of the model development process can be progressively automated. Together, these challenges and opportunities have driven increasing interest in self-improvement, where models autonomously generate data, evaluate outputs, and iteratively refine their own capabilities. In this paper, we present a system-level perspective on self-improving language models and introduce a unified framework that organizes existing techniques. We conceptualize the self-improvement system a

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

KA2L: A Knowledge-Aware Active Learning Framework for LLMs

arXiv:2603.17566v2 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) with high-quality knowledge has been shown to enhance their performance effectively. However, there is a paucity of research on the depth of domain-specific knowledge comprehension by LLMs and the application of targeted active learning to improve their expertise. To address this gap, we introduce the Knowledge-Aware Active Learning (KA2L) framework. This framework assesses LLMs' mastery of specific knowledge points to aid in constructing unanswerable or unknowable questions through latent space analysis. This active learning strategy enhances training efficiency by focusing on knowledge the model has yet to master, thereby minimizing redundancy in learning already acquired information. This study innovatively employs a knowledge distribution probing technique to examine the hidden states of specific Transformer layers and identify the distribution of known and unknown knowledge within the LLM.

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

AI Can Learn Scientific Taste

arXiv:2603.14473v3 Announce Type: replace Abstract: Scientific discovery depends on expert judgement and foresight, which we call scientific taste: the ability to judge and propose research ideas with the potential for long-term scientific impact. Scientific taste is largely concentrated among highly experienced researchers, whose expertise is usually limited to a few specialised fields. If AI could learn scientific taste, it could reduce reliance on human experts and accelerate scientific discovery. Whether AI can learn this ability remains an open question. We introduce Reinforcement Learning from Community Feedback (RLCF) to learn judgement and ideation. Scientific Judge learns from community feedback, such as citations. Scientific Thinker learns to propose research ideas with high potential impact. Experiments show that Scientific Judge outperforms strong LLM baselines and that learned judgement generalises to future-year papers, other community metrics, and unseen fields. Furtherm

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Repeatability is not recovery: Quantifying algorithmic stability and topic recovery in Latent Dirichlet Allocation

arXiv:2511.12850v2 Announce Type: replace Abstract: Topic models are often judged by the consistency of their outputs across repeated runs, implicitly assuming that repeatable topic output is a successful recovery of the underlying topics. We show that this assumption is false: repeatability is not recovery. We introduce a stability framework that jointly measures consistency among repeated runs and accuracy relative to known ground truth. Because real-world corpora lack known topic structures, we generate synthetic corpora using the Latent Dirichlet Allocation (LDA) generative process, enabling direct evaluation of topic recovery. Across 50 repeated LDA runs on each corpus, we find that LDA reliably identifies the correct number of topics and frequently converges to highly consistent topic solutions. However, these repeatable solutions frequently fail to recover the true generating topics. Thus, internal stability should not be interpreted as evidence of correctness. Our results illus

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation

arXiv:2510.03115v2 Announce Type: replace Abstract: Speech-to-Text Translation (S2TT) systems built from Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) modules face two major limitations: error propagation and the inability to exploit prosodic or other acoustic cues. Chain-of-Thought (CoT) prompting has recently been introduced, with the expectation that jointly accessing speech and transcription will overcome these issues. Analyzing CoT through attribution methods, robustness evaluations with corrupted transcripts, and prosody-awareness, we find that it largely mirrors cascaded behavior, relying mainly on transcripts while barely leveraging speech. Simple training interventions, such as adding Direct S2TT data or noisy transcript injection, enhance robustness and increase speech attribution. These findings challenge the assumed advantages of CoT and highlight the need for architectures that explicitly integrate acoustic information into translation.

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Future Policy Approximation for Offline Reinforcement Learning in LLM Reasoning

arXiv:2509.19893v3 Announce Type: replace Abstract: Reinforcement learning (RL) has emerged as a key driver of post-training for complex reasoning in large language models (LLMs), yet online RL introduces substantial instability and computational overhead. Offline RL offers a compelling alternative by decoupling generation from training; however, offline algorithms for reasoning remain under-optimized relative to their online counterparts. We revisit the potential of policy-gradient-style offline RL and address a central challenge in offline learning: gradient entanglement. In long-horizon reasoning trajectories, correct and incorrect solutions share substantial token overlap, causing gradient updates from incorrect trajectories to suppress tokens that are also critical for correct ones. We propose Future Policy Approximation (FPA), a simple offline policy-gradient method that weights gradients using an estimate of the future policy rather than the current policy, enabling proactive gr

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety

arXiv:2508.20468v2 Announce Type: replace Abstract: Conspiracy theories erode public trust in science and institutions while resisting debunking by evolving and absorbing counter-evidence. As AI-generated misinformation becomes increasingly sophisticated, understanding the rhetorical patterns in conspiratorial content is important for developing interventions such as targeted prebunking and assessing AI vulnerabilities. We introduce CONSPIRED (CONSPIR Evaluation Dataset), which captures the cognitive traits of conspiratorial ideation in multi-sentence excerpts (80-120 words) from online conspiracy articles, annotated using the CONSPIR cognitive framework. CONSPIRED is the first dataset of conspiratorial content annotated for general cognitive traits. Using CONSPIRED, we (i) develop computational models that identify conspiratorial traits and the dominant trait in text excerpts, and (ii) evaluate LLM robustness to conspiratorial inputs. We find that LLMs are readily misaligned by conspi

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

arXiv:2506.01367v4 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly integrated into agentic AI systems, yet their propensity to generate hallucinations remains a critical safety concern. Detecting these factual errors at test-time, particularly without ground-truth labels, is essential for building trustworthy autonomous agents. We propose MMD-Flagger, an hallucination detection method that utilizes Maximum Mean Discrepancy (MMD) and monitors the stability of LLM outputs across varying decoding temperatures. Our method tracks the MMD trajectory between a LLM's response at a certain decoding configuration and a set of stochastic samples, identifying hallucinations based on the trajectory's characteristic shape. We evaluate MMDFlagger on multi-lingual claim verification benchmarks (MUCH) using modern LLMs like Llama-3 families and Gemma-3.

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CL

Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models

arXiv:2307.12896v5 Announce Type: replace Abstract: The article introduces corrections to Zipf's and Heaps' laws based on systematic models of the proportion of hapaxes, i.e., words that occur once. The derivation rests on two assumptions. The first one is the standard urn model which predicts that marginal frequency distributions for shorter texts look as if word tokens were sampled blindly from a given longer text. The second assumption posits that the hapax rate is a simple function of the text length. Four such functions are discussed: the constant model, the cancelation model, the linear model, and the logistic model. As a simple illustration, it is shown that the logistic model yields the best fit for a sample of 14 texts in English. The need and the availability of more complex mixture models that reflect two-regime vocabularies for larger corpora is also discussed.

Source ↗
Showing 7301–7350 of 18402 signals
← Prev Page 147 of 369 Next →