EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Wed, 12 Aug 2026 13:22:00 +0000
MedCity News

Ambulatory Growth Demands a New Operating Model

While organizations have invested heavily in expanding ambulatory networks, many are discovering that operational workflows have not evolved at the same pace. The result is a growing gap between demand and the systems designed to manage it. The post Ambulatory Growth Demands a New Operating Model appeared first on MedCity News .

Source ↗
technology Wed, 12 Aug 2026 13:02:25 -0400
EdTech Mag (Higher)

Centralized Cybersecurity Operations Keeps Community Colleges Protected

As the largest single-district community college system in the nation, Maricopa Community Colleges educates more Arizona residents than its three large public universities combined. The system operates 10 main campuses across the Phoenix metropolitan area and more than 30 satellite sites, serving nearly 100,000 students and employing 14,000 faculty members and staff in the 2026 spring semester. Those superlatives and numbers are all familiar to Jamie Spradlin, MCC’s CISO. When he took his position in June 2024, his first order of business was to establish a cybersecurity program…

Source ↗
behavior Wed, 12 Aug 2026 12:36:51 +0000
District Admin

Why is enrollment across DC-area school districts falling?

With enrollment trends declining, and projections expecting that to continue, the Prince William County School Board decided to nix plans for what it hoped would one day become its 14th high school. The post Why is enrollment across DC-area school districts falling? appeared first on District Administration .

Source ↗
behavior Wed, 12 Aug 2026 12:34:50 +0000
District Admin

Schools spend billions on AI, but struggle to figure out what’s worth it

Educators and experts say districts are carrying much of the burden of vetting products on their own. The post Schools spend billions on AI, but struggle to figure out what’s worth it appeared first on District Administration .

Source ↗
regulation Wed, 12 Aug 2026 12:30:00 +0000
The 74

Opinion: Don’t Blink: Why NYC Reads Must Stay on Course

Nearly 25 years ago, I walked into the U.S. Department of Education at what felt like a remarkable moment in the history of reading instruction. The National Reading Panel had just been released, and for the first time in my professional life, there was broad agreement that educational policy should be grounded in scientific evidence. […]

Source ↗
technology Wed, 12 Aug 2026 12:21:33 +0000
HN: edtech

SIMO Educación 2026: A Practical Exhibitor Guide EdTech Companies Ifema Madrid

Article URL: https://adamexpostand.substack.com/p/simo-educacion-2026-a-practical-exhibitor Comments URL: https://news.ycombinator.com/item?id=49271263 Points: 2 # Comments: 0

Source ↗
regulation Wed, 12 Aug 2026 10:30:00 +0000
The 74

Thousands of LAUSD Workers Face Layoffs That Didn’t Have to Happen

On the brink of looming insolvency, Los Angeles Unified could have prevented its financial crisis by leaving vacated positions empty, making thousands of district employee layoffs unavoidable to address a $3.6 billion deficit, a top state fiscal watchdog said. “They’re overstaffed significantly,” said Michel Fine, CEO of California’s Fiscal Crisis and Management Assistance Team of […]

Source ↗
behavior Wed, 12 Aug 2026 10:00:00 +0000
eSchool News

The biggest back-to-school challenge

As students return to classrooms across America this fall, parents are focused on school supplies, class schedules, and academic expectations. Teachers are preparing lesson plans, organizing classrooms, and hoping for a successful year.

Source ↗
technology Wed, 12 Aug 2026 09:00:00 +0000
eCampus News

Higher education needs better AI experiences, not more AI tools

Artificial intelligence is becoming increasingly embedded across higher education. As AI policies continue to evolve, faculty are gaining greater clarity and opportunity to responsibly use new tools, including AI-powered assistants, content generation tools, chatbots, and a growing number of AI-enabled applications. The post Higher education needs better AI experiences, not more AI tools appeared first on eCampus News .

Source ↗
technology Wed, 12 Aug 2026 09:00:00 +0000
Tech & Learning

Teaching Veteran Students

Veterans and active-duty military members together make up about 5% of college students in the U.S. Educator and retired Major General Matt Smith shares tips for connecting with these students.

Source ↗
technology Wed, 12 Aug 2026 09:00:00 +0000
eCampus News

Higher education needs better AI experiences, not more AI tools

Artificial intelligence is becoming increasingly embedded across higher education. As AI policies continue to evolve, faculty are gaining greater clarity and opportunity to responsibly use new tools, including AI-powered assistants, content generation tools, chatbots, and a growing number of AI-enabled applications. The post Higher education needs better AI experiences, not more AI tools appeared first on eCampus News .

Source ↗
audience Wed, 12 Aug 2026 07:00:00 +0000
Inside Higher Ed

Academic Forests Are Higher Ed’s Hidden Jewels

Academic Forests Are Higher Ed’s Hidden Jewels Elizabeth Redden Wed, 08/12/2026 - 03:00 AM Forests are an asset for employee and student well-being and offer invaluable opportunities for teaching, research and recreation. Universities should protect them. Byline(s) Robert F. Baldwin Elizabeth Dennis Baldwin Olin Thompson Mefford Danny Weathers Robert H. Jones

Source ↗
audience Wed, 12 Aug 2026 07:00:00 +0000
Inside Higher Ed

The Odd Couple: We’re Not Back in School

The Odd Couple: We’re Not Back in School rachel.toor Wed, 08/12/2026 - 03:00 AM Another academic year starts and neither Gordon nor Rachel will be on campus. Byline(s) Rachel Toor E. Gordon Gee

Source ↗
audience Wed, 12 Aug 2026 07:00:00 +0000
Inside Higher Ed

The Legislation That Eats Away at Higher Education’s Financial Footing

The Legislation That Eats Away at Higher Education’s Financial Footing kjohnsonbowles… Wed, 08/12/2026 - 03:00 AM Legislation is increasing financial burdens on students and institutions. Here are some things to think about when trying to fix the business model. Byline(s) Kathy Johnson Bowles

Source ↗
audience Wed, 12 Aug 2026 07:00:00 +0000
Inside Higher Ed

Departing UNCF President Reflects on HBCU ‘Renaissance’

Departing UNCF President Reflects on HBCU ‘Renaissance’ Sara Weissman Wed, 08/12/2026 - 03:00 AM After more than two decades at the helm of the United Negro College Fund, Michael L. Lomax looks back at his career advocating for historically Black colleges and universities. Byline(s) Sara Weissman

Source ↗
audience Wed, 12 Aug 2026 07:00:00 +0000
Inside Higher Ed

Researcher Accused of Espionage Files Second Lawsuit Against University

Researcher Accused of Espionage Files Second Lawsuit Against University Emma Whitford Wed, 08/12/2026 - 03:00 AM Byline(s) Emma Whitford

Source ↗
audience Wed, 12 Aug 2026 07:00:00 +0000
Inside Higher Ed

A Pre-Orientation Designed for Student Parents

A Pre-Orientation Designed for Student Parents Joshua.Bay Wed, 08/12/2026 - 03:00 AM Wichita State’s new pre-orientation program builds belonging and connects student parents with resources before the semester begins. Byline(s) Joshua Bay

Source ↗
audience Wed, 12 Aug 2026 07:00:00 +0000
Inside Higher Ed

Linda McMahon Legal Battle Grinds On

Linda McMahon Legal Battle Grinds On Josh Moody Wed, 08/12/2026 - 03:00 AM The education secretary was sued in 2024 for allegedly ignoring sexual abuse when she was a WWE executive. Almost two years in, the largely overlooked lawsuit is playing out slowly. Byline(s) Josh Moody

Source ↗
audience Wed, 12 Aug 2026 07:00:00 +0000
Inside Higher Ed

Loss of International Students Could Cost Economy $3.4B

Loss of International Students Could Cost Economy $3.4B kathryn.palmer… Wed, 08/12/2026 - 03:00 AM Byline(s) Kathryn Palmer

Source ↗
audience Wed, 12 Aug 2026 07:00:00 +0000
Inside Higher Ed

ACLU Sues Florida International for Punishing ICE Protesters

ACLU Sues Florida International for Punishing ICE Protesters Olivia.sanchez Wed, 08/12/2026 - 03:00 AM Byline(s) Olivia Sanchez

Source ↗
audience Wed, 12 Aug 2026 07:00:00 +0000
Inside Higher Ed

The Groups Pushing Back on McMahon’s ‘Call to Action’

The Groups Pushing Back on McMahon’s ‘Call to Action’ Katherine Knott Wed, 08/12/2026 - 03:00 AM As university leaders announce plans to consult faculty on how to respond to the education secretary’s questions, two groups call for unity and resistance. Byline(s) Katherine Knott

Source ↗
audience Wed, 12 Aug 2026 05:00:00 -0400
Higher Ed Dive

College costs for families up 10% from last year, poll finds

Tuition, fees, housing and other costs averaged $34,019 in the 2025-26 academic year, per a new Sallie Mae and Ipsos report.

Source ↗
regulation Wed, 12 Aug 2026 05:00:00 -0400
K-12 Dive

What you need to know about microschools

The model has become a popular school choice option post-pandemic, and calls for policymakers to give the sector more structure are growing.

Source ↗
regulation Wed, 12 Aug 2026 05:00:00 -0400
K-12 Dive

Miami-Dade eliminated free school meals. They won’t be the last.

A senior policy analyst at the Center for American Progress breaks down how and why the nation’s third-largest district lost access to district-wide free meals.

Source ↗
technology Wed, 12 Aug 2026 03:46:15 +0000
MedCity News

Health IT Vendor’s Data Breach Exposes Nearly 4M Patient Records

Revenue cycle vendor Unlimited Technology Systems disclosed a ransomware attack affecting 3.8 million patients — the second-largest healthcare data breach reported to HHS this year. The post Health IT Vendor’s Data Breach Exposes Nearly 4M Patient Records appeared first on MedCity News .

Source ↗
technology Wed, 12 Aug 2026 03:24:29 +0000
HN: education

Hannah Arendt's American Education

Article URL: https://www.newyorker.com/magazine/2026/08/17/hannah-arendt-life-of-the-mind-thomas-meyer-book-review-an-admirable-woman-arthur-cohen Comments URL: https://news.ycombinator.com/item?id=49267500 Points: 3 # Comments: 1

Source ↗
behavior Wed, 12 Aug 2026 00:00:00 GMT
EdSurge

Supporting Early Childhood Educators

The guests on this episode of This Week with EdSurge say good tools fail without good follow-through.

Source ↗
behavior Wed, 12 Aug 2026 00:00:00 GMT
EdSurge

What Do the Youngest Learners in the Building Actually Need?

The guests on this episode of This Week with EdSurge say good tools fail without good follow-through.

Source ↗
behavior Wed, 12 Aug 2026 00:00:00 GMT
EdSurge

The K-12 Silo Won’t Survive AI — But Our Relationships Might

I left my school to ask tech executives how to prepare students for the future.

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Leak It: Per-Document Extraction Beyond Aggregate Membership Inference

arXiv:2608.00144v2 Announce Type: replace-cross Abstract: Membership inference (MIA) on language models is usually summarised by aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines can separate members from non-members using surface text alone. Building on probabilistic discoverable extraction, we study black-box training-data leakage using N samples from p_theta(. | x), placing mean overlap, extreme-value overlap, and self-concentration on a common functional-estimation footing. On WikiMIA, a blind bag-of-words classifier reaches AUC 0.97 (TPR 0.90 at 5% FPR) while sampling adds nothing. On an IID Pile split (MIMIR), neither self-concentration nor gold-continuation recovery significantly exceeds a blind baseline in aggregate. Aggregate metrics hide the real harm: sampling verbatim-extracts training data for a tail of documents no blind attack can reach. On Pythia-6.9B, 16.6% of 500 Pile documents bearing a real identifier (83 documents; 21.3% of those be

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

SAE-StatSteer: Statistical Consensus Feature Selection for Optimization-Free Activation Steering of Large Language Models

arXiv:2607.19364v2 Announce Type: replace-cross Abstract: Activation steering adds a residual-stream direction at inference time, providing lightweight behavioral control without fine-tuning. Sparse autoencoders (SAEs) can make such interventions auditable by decomposing dense activations into an approximately monosemantic feature basis. We introduce SAE-StatSteer, a transparent, optimization-free pipeline. It first filters features through six reliability conditions, then ranks the survivors by an unweighted Borda consensus over three statistics, an $F$-test, KSG mutual information, and Cohen's $d$, and finally combines the selected SAE decoder rows using Cohen's-$d$ weights. We evaluate three Gemma-family models across four behavioral domains against seven dense or SAE-based baselines. Our quality-conditioned protocol requires attribute movement while preserving relevance, richness, and coherence. Raw success systematically overstates usable control because strong shifts often degrad

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents

arXiv:2606.07943v2 Announce Type: replace-cross Abstract: Agent skills extend general-purpose agents, but their open format enables skill poisoning: a tampered skill can make an agent run an attacker's command while completing the user's legitimate task. Invocation alone is insufficient; the attack-specific action must complete while that task still passes its verifier. We therefore define Attack Success Rate (ASR) to require a postcondition-validated sandbox action and a passing task verifier in the same trial. Skill files expose a reliability-visibility trade-off between a preloaded but conspicuous YAML frontmatter block and a longer body, where arbitrary placement may be skipped or locally incongruent. We introduce Poise, a position-aware attack that uses context-aware generation to place exactly one benign-looking, command-bearing instruction at a structurally feasible body position. On the eligible Skill-Inject pool with codex+gpt-5.2, Poise achieves 89.3\% ASR, 28.0 points above

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoning

arXiv:2606.00671v2 Announce Type: replace-cross Abstract: We present AXIOM, a trust-first neuro-symbolic architecture for natural-language mathematical reasoning. Its language model is strictly a canonicalizer: it rewrites informal problem text into a narrow schema consumed by a deterministic Computer-Algebra-System (CAS) pipeline, which derives and verifies the answer or abstains as a first-class output. Routing follows a 1:1:1 alignment between problem-shape regex, schema-specific prompt, and closed-form CAS handler, with 4,783 such routes shipped, 71% of which answer without invoking the language model, and zero LOST_CORRECT regressions as a standing release gate. Because the answer is derived rather than generated, so is its explanation: every handler emits a step trace of the computation it performed, rendered as prose by a layer covering all 4,785 tasks that cannot narrate a step the handler did not take. We report two numbers and never fuse them. On the full 7-category MATH test

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models

arXiv:2605.16411v3 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling. We propose a stage-wise preference optimization framework for hallucination reduction through targeted multimodal data construction. Rather than directly optimizing on generic instruction-following data, our approach progressively constructs hallucination-focused preference pairs near known failure boundaries. The framework emphasizes ambiguous spatial orientation, object relationships, OCR uncertainty, and adversarial false-premise training. Hallucinated negatives are generated through minimally perturbed yet visually inconsistent alternatives, enabling Direct Preference Optimization (DPO) to better separate grounded reasoning from plausible hallucination.

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Spherical Flows for Sampling Categorical Data

arXiv:2605.05629v4 Announce Type: replace-cross Abstract: We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically operate in Euclidean space or on the probability simplex, we instead work on the sphere $\mathbb S^{d-1}$. There the von Mises-Fisher (vMF) distribution induces a natural noise process and admits a closed-form conditional score. The conditional velocity is in general intractable. Exploiting the radial symmetry of the vMF density we reduce the continuity equation on $\mathbb S^{d-1}$ to a scalar ODE in the cosine similarity, whose unique bounded solution determines the velocity. The marginal velocity and marginal score on $(\mathbb S^{d-1})^L$ both decompose into posterior-weighted tangent sums that differ only by per-token scalar weights. This gives access to both ODE and predictor-corrector (PC) sampling. The posterior is the only learned object, trained by a cross-entropy loss. Experimen

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents

arXiv:2604.16706v2 Announce Type: replace-cross Abstract: Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, yet this assumption is rarely validated against human annotation. We present AgentProp-Bench, a diagnostic benchmark of 14,750 execution traces from thirteen LLM agents (nine proprietary, four open-weight) across four domains, and use it to audit three questions. First, substring-heuristic judging of agent outputs agrees with human annotation only at chance level (Cohen's kappa = 0.049 against each of two annotators), while a three-LLM ensemble reaches moderate agreement (kappa = 0.432) and a single GPT-4o-mini judge is in fact the strongest (kappa = 0.567); dual-annotator agreement is almost perfect (kappa = 0.835). Second, under validated judging a parameter-level error propagates to a wrong final answer with human-calibrated probability approximately 0.62, replicated across proprietary and open-weight models, and a model's a

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models

arXiv:2509.25143v2 Announce Type: replace-cross Abstract: Existing medical reasoning benchmarks for vision-language models primarily focus on analyzing a patient's condition based on an image from a single visit. However, this setting deviates significantly from real-world clinical practice, where doctors typically refer to a patient's historical conditions to provide a comprehensive assessment by tracking their changes over time. In this paper, we introduce TEMMED-BENCH, a multi-task benchmark designed for analyzing changes in patients' conditions between different clinical visits, which challenges large vision-language models (LVLMs) to reason over temporal medical images. TEMMED-BENCH consists of a test set comprising three tasks - visual question-answering (VQA), report generation, and image-pair selection - and a supplementary knowledge corpus of over 17,000 instances. With TEMMED-BENCH, we conduct an evaluation of twelve LVLMs, comprising six proprietary and six open-source model

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing

arXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are described from non-camera perspectives. To address this limitation, we propose Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing (FoR-SALE), an extension of the Self-correcting LLM-controlled Diffusion (SLD). FoR-SALE first evaluates the alignment between a given text and an initially generated image, and then refines the image based on the expressed FoR in the spatial description. It employs vision modules to extract the spatial configuration of the generated image and simultaneously maps the spatial expression to a corresponding camera perspective. This unified perspective enables direct evaluation of alignment between language and vision. When misalignment is detected, the required editing operations are generated and applied. FoR-SALE introduces novel latent-space

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Multiplayer Nash Preference Optimization

arXiv:2509.23102v4 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grounded in the Bradley-Terry assumption struggle to capture the nontransitivity and heterogeneity of real-world preferences. To address this, recent studies have reframed alignment as a two-player Nash game, giving rise to Nash learning from human feedback (NLHF). While this perspective has inspired algorithms such as INPO, ONPO, and EGPO that offer strong theoretical and empirical guarantees, they remain fundamentally restricted to two-player interactions, introducing a single-opponent bias that fails to capture the full complexity of realistic preference structures. This work introduces Multiplayer Nash Preference Optimization (MNPO), a novel framework that generalizes NLHF to the multiplayer regime. It formulates alignment as an n-player game, where ea

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign

arXiv:2502.02068v3 Announce Type: replace-cross Abstract: This paper introduces RoSeMary, the first-of-its-kind ML/Crypto codesign watermarking framework that regulates LLM-generated code to avoid intellectual property rights violations and inappropriate misuse in software development. High-quality watermarks adhering to the detectability-fidelity-robustness tri-objective are limited due to codes' low-entropy nature. Watermark verification, however, often needs to reveal the signature and requires re-encoding new ones for code reuse, which potentially compromising the system's usability. To overcome these challenges, RoSeMary obtains high-quality watermarks by training the watermark insertion and extraction modules end-to-end to ensure (i) unaltered watermarked code functionality and (ii) enhanced detectability and robustness leveraging pre-trained CodeT5 as the insertion backbone to enlarge the code syntactic and variable rename transformation search space. In the deployment, RoSeMary

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference

arXiv:2607.13205v2 Announce Type: replace Abstract: Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. On schema-dense input streams such as nested JSON, this score acts as a non-stationary filter that disproportionately retains noise: a non-content sink role (delimiters or whitespace) carries an order of magnitude more energy than any content role, and structural KEY tokens are over-retained at roughly 1.8x the rate of the answer-carrying VALUE tokens, collapsing exact-match accuracy from 88% to 0% at a 5% budget as the signal-to-noise ratio of the retained state degrades. A counterfactual experiment establishes that suppressing KEY tokens is the best deployable filter. Our retraining-free, role-conditional allocation over SnapKV's windowed score, governed by a single tuned hyperparameter, closes 63-98% of th

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages

arXiv:2606.15821v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, forming distinct model lineages. It remains unclear whether a fundamental behavioral link exists between the foundational LLMs and downstream variants. We investigate this question by quantifying head-level context-truthfulness scores. Across diverse LLM and MLLM lineages, including Vicuna-, Qwen2.5-, LLaMA2-, and Mistral-based models, we find that Truth Scores are strongly preserved within model families, even after instruction tuning or multimodal adaptation. We further show that this inheritance is consistent with attention-head weight preservation, and that context-truthful heads attend to query-relevant evidence. Building on this finding, we propose TruthProbe, a soft-gating strategy that amplifies context-truthful heads while preserving other head contributions. TruthProbe improves contextua

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes

arXiv:2606.11470v2 Announce Type: replace Abstract: Reasoning has become central to how Large Language Models (LLMs) are evaluated and interpreted, spanning Chain-of-Thought (CoT), mathematical problem-solving, multi-hop question answering, code generation, retrieval-augmented reasoning, tool use, and multimodal decision-making. In this survey, we introduce the Periodic Table of LLM Reasoning, a framework organizing 300+ recent papers by reasoning paradigm, methodological mechanism, evaluation setting, and failure mode. We classify LLM reasoning into nine paradigms: Chain-of-Thought, Multi-Hop, Mathematical, Commonsense, Visual and Temporal, Code and Algorithmic, Retrieval-Augmented, Tool-Augmented or Agentic, and Reinforcement Learning-based reasoning. For each, we review approaches, including prompting, architectural interventions, supervised fine-tuning, verifier-guided inference, reward modeling, retrieval, tool interfaces, agentic workflows, and benchmark design. We argue that LLM

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses

arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these assets through heuristic reflection or raw success counts. Such updates are brittle when trajectories are sparse, expensive, and context-dependent. We introduce Bayesian-Agent, a native and cross-harness framework that treats reusable agent skills as Bayesian evidence objects. Bayesian-Agent records verified trajectories, maintains posterior beliefs over skill reliability and failure modes, and turns those beliefs into auditable skill actions and model-facing guardrails. This posterior view provides a finite-sample alternative to raw empirical-rate skill updates and frames prompt, context, and harness engineering as inference over the external decision environment. On RealFin-Bench, Bayesian skill evolution matches or improves the raw empirical-rate control and yields large gains on native

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

FinEvolveBench: A Benchmark for Self-Evolving Agents on Low-Repetition Tasks with Implicit Rewards

arXiv:2606.06960v2 Announce Type: replace Abstract: Experience-based self-evolution enables language-model agents to improve their behavior by accumulating and updating experience at test time, yet existing evaluations often assume recurring task patterns and explicit success signals. We introduce \textsc{FinEvolveBench}, a benchmark for self-evolving agents on low-repetition tasks with implicit rewards. The benchmark reconstructs a daily financial information stream over 31 Chinese A-share industry indices and aligns 177,324 public news articles with market observations. Researchers can define prediction horizons over this stream; we evaluate predictive market-sentiment factors against delayed market-adjusted returns after 10, 20, and 40 trading days. Unlike static benchmarks that score each prediction independently, \textsc{FinEvolveBench} interleaves new decisions with delayed outcomes from earlier ones, testing whether agents can convert noisy real-world feedback into reusable expe

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models

arXiv:2606.03793v2 Announce Type: replace Abstract: Multimodal Large Language Models integrate visual perception into language reasoning, introducing a continuous attack surface susceptible to adversarial attacks. Prior work on MLLM robustness has focused largely on English-centric tasks, leaving multilingual behaviour unexplored. We address this gap through a systematic study of adversarial robustness and multimodal safety across 12 diverse languages, evaluating open-source MLLMs that acquire multilingual capability through instruction tuning. Gradient-based attacks reveal a transferable multilingual vulnerability: adversarial images optimized in one language continue to induce failure in others, demonstrating strong cross-lingual transferability. Multilingual safety further varies with how effectively a model retrieves or interprets harmful instructions. When harmful intent is issued through text, languages with stronger linguistic grounding more often elicit misuse-enabling response

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Why Do Safety Guardrails Degrade Across Languages?

arXiv:2605.17173v2 Announce Type: replace Abstract: Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds several safety-driving factors into one, obscuring the specific cause(s) of safety failure. We introduce a latent variable model, a Multi-Group Item Response Theory (IRT) framework, that decouples language-agnostic safety robustness ($\theta$), intrinsic prompt hardness ($\beta$), global language processing difficulty ($\gamma$), and a prompt-specific cross-lingual safety gap ($\tau$). Using the MultiJail dataset, we evaluate the safety robustness of 61 model configurations across 5 closed-model families and 10 languages of varying resource, aggregating a dataset of 1.9 million responses. Exploratory Factor Analysis shows safety is primarily unidimensional: models refuse different harm types mainly through a shared mechanism. Contrary to the expected trend that safety degrades largely i

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation

arXiv:2605.11533v3 Announce Type: replace Abstract: Routine clinical check-up reports combine laboratory measurements, physiological assessments, imaging findings and visually structured information, but rarely tell patients what to do next. Translating them into follow-up actions requires models to connect evidence across pages, tables and modalities, identify clinically relevant issues and communicate next steps without unsupported diagnostic or treatment claims. Yet this report-to-action capability remains poorly benchmarked. We introduce \textbf{C2A}, a dataset and benchmark for generating structured \textit{Action Cards} from multimodal check-up reports, together with \textbf{Checkup2Action}, a constrained workflow for the task. C2A contains 2,000 de-identified real-world reports covering physical examinations, laboratory tests, cardiovascular assessments and imaging evidence. Each card specifies one issue, its priority, recommended department, follow-up window, patient-facing exp

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

FlexSQL: Flexible Exploration and Execution Make Better Text-to-SQL Agents

arXiv:2605.02815v2 Announce Type: replace Abstract: Text-to-SQL over large analytical databases requires navigating complex schemas, resolving ambiguous queries, and grounding decisions in actual data. Most current systems follow a fixed pipeline where schema elements are retrieved once upfront and the database is only revisited for post-hoc repair, limiting recovery from early mistakes. We present FlexSQL, a text-to-SQL agent whose core design principle is flexible database interaction: the agent can explore schema structure, inspect data values, and run verification queries at any point during reasoning. FlexSQL generates diverse execution plans to cover multiple query interpretations, implements each plan in either SQL or Python depending on the task, and uses a two-tiered repair mechanism that can backtrack from code-level errors to plan-level revisions. On Spider2-Snow, using gpt-oss-120b, FlexSQL achieves a 65.4\% score, outperforming strong open-source baselines that use stronge

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Language corpora for the Dutch medical domain

arXiv:2604.25374v2 Announce Type: replace Abstract: Background: Dutch medical corpora are scarce, limiting NLP development. Methods: We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. Results: The resulting corpus comprises +- 36 billion tokens across the medical domain in about 105 million documents, freely available on Hugging Face. Conclusion: This work establishes the first large-scale Dutch medical language corpus for pre-training and downstream NLP tasks.

Source ↗
Showing 1151–1200 of 18349 signals
← Prev Page 24 of 367 Next →