EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 01 Sep 2026 17:07:07 +0000
MedCity News

GSK’s Pivotal Test for mRNA Flu Vaccine Aims to Show Two Targets Are Better Than One

GSK’s messenger RNA vaccine for seasonal influenza is designed to prompt an immune response to two proteins on the surface of the virus. That could be an advantage over Moderna’s recently approved mRNA flu shot, which addresses just one of those proteins. The post GSK’s Pivotal Test for mRNA Flu Vaccine Aims to Show Two Targets Are Better Than One appeared first on MedCity News .

Source ↗
regulation Tue, 01 Sep 2026 17:00:00 -0400
K-12 Dive

Districts buckle on LGBTQ+ policies under Trump administration pressure

As some districts push back against investigations and federal funding threats, others are changing their policies after warnings from the government.

Source ↗
regulation Tue, 01 Sep 2026 16:30:00 +0000
The 74

3-Year College Degrees Grow More Popular in the U.S., Saving Time and Cash. But Critics Aren’t Sold

Giles Sims, 34, was sick of getting passed over for jobs because he didn’t have a college degree. But with a wife and a widowed mother to help support, he needed a quick solution. So the laid-off web developer turned to one of a growing number of three-year degree programs. “I can’t afford to be […]

Source ↗
audience Tue, 01 Sep 2026 15:49:11 -0400
Higher Ed Dive

Ohio State University settles DOJ allegations for $2.1M

The agency alleges the public institution, which did not admit wrongdoing, failed to disclose ties between its faculty and the Chinese government.

Source ↗
technology Tue, 01 Sep 2026 13:42:00 +0000
MedCity News

Breaking the Amendment Cycle: How Agentic AI Enables Smarter Clinical Trial Design and Operations

Faster trials are not a vanity metric. They reflect an undeniable need: getting safe therapies to patients sooner. The post Breaking the Amendment Cycle: How Agentic AI Enables Smarter Clinical Trial Design and Operations appeared first on MedCity News .

Source ↗
audience Tue, 01 Sep 2026 13:00:00 -0400
Higher Ed Dive

AI is reshaping campus infrastructure. Is higher education ready?

Looking ahead, higher education institutions face a pivotal moment with AI.

Source ↗
audience Tue, 01 Sep 2026 13:00:00 -0400
Higher Ed Dive

AI Is reshaping campus infrastructure. Is Higher Education ready?

Looking ahead, higher education institutions face a pivotal moment with AI.

Source ↗
behavior Tue, 01 Sep 2026 12:40:39 +0000
District Admin

With its back-to-school video, Portland Public Schools touches a nerve for both liberals and conservatives

It begins with an address from legendary Portland drag queen, activist and educator Poison Waters, who received an honorary doctorate in humane letters from Portland State University in 2025. The post With its back-to-school video, Portland Public Schools touches a nerve for both liberals and conservatives appeared first on District Administration .

Source ↗
behavior Tue, 01 Sep 2026 12:38:00 +0000
District Admin

More than a dozen Texas school districts close to triggering state takeover law in 2027

At least 14 districts had at least one campus with a fourth consecutive failing grade, according to the preliminary 2026 ratings. The post More than a dozen Texas school districts close to triggering state takeover law in 2027 appeared first on District Administration .

Source ↗
regulation Tue, 01 Sep 2026 12:30:00 +0000
The 74

Opinion: Helping Students Navigate College & Career Must Become Central to K-12’s Mission

Policymakers have invested heavily in recent years to create new career pathways for high school students. Career and technical education has expanded dramatically, as have the credentials high schoolers can earn. Yet too many students lack the guidance they need to benefit from these growing opportunities. Counselor shortages remain persistently severe, while career navigation is often […]

Source ↗
behavior Tue, 01 Sep 2026 11:46:54 +0000
District Admin

Alabaster Scales Responsible AI for English Learners

After a successful pilot, Alabaster is adding four Telo AI robots. The post Alabaster Scales Responsible AI for English Learners appeared first on District Administration .

Source ↗
regulation Tue, 01 Sep 2026 10:30:00 +0000
The 74

AI Tutors Not Yet a Replacement for Humans, Research Says

Wendy Graham’s two kids first encountered Amira, a popular artificial intelligence tutoring platform, in 2025, during a summer reading program in New Mexico. She also used it at home with her daughter, now a fourth grader. But she found that the purple-haired avatar with the big, round glasses didn’t seem to understand where her daughter […]

Source ↗
behavior Tue, 01 Sep 2026 10:00:00 +0000
eSchool News

We stopped waiting for the teacher pipeline–we built our own

Every hiring season, school districts ask the same question: Where are the teachers? We stopped asking and started building.

Source ↗
technology Tue, 01 Sep 2026 09:00:00 +0000
Tech & Learning

Survival Guide For Leaders Navigating 2 Sides Of A Coin

By anchoring decisions in objective principles rather than emotional conflicts, effective school leaders can satisfy both passionate educators and defensive parents

Source ↗
need Tue, 01 Sep 2026 09:00:00 +0000
Hechinger Report

Business, government and education leaders try to reinvent college from scratch

MONTPELIER, Vt. — All red brick and white trim and dating from as far back as the 1870s, the buildings that frame a grassy quad on a hill above Vermont’s capital city comprise the quintessential college campus. Behind the scenes, this holdover from the 19th century is where a group of business, college and government […] The post Business, government and education leaders try to reinvent college from scratch appeared first on The Hechinger Report .

Source ↗
audience Tue, 01 Sep 2026 07:00:00 +0000
Inside Higher Ed

Trump Attacks Academia—and Grad Students Ask if It’s Worth Saving

Trump Attacks Academia—and Grad Students Ask if It’s Worth Saving Elizabeth Redden Tue, 09/01/2026 - 03:00 AM It’s not just the federal government that’s telling grad students their work doesn’t matter. Byline(s) Lauren Chambers JP Flores

Source ↗
audience Tue, 01 Sep 2026 07:00:00 +0000
Inside Higher Ed

Jason Arday’s Death Shouldn’t Affect Who Gets Hired Next

Jason Arday’s Death Shouldn’t Affect Who Gets Hired Next Elizabeth Redden Tue, 09/01/2026 - 03:00 AM His story should not be used to further undermine efforts to diversify the faculty pipeline. Byline(s) Lisa J. Servon

Source ↗
audience Tue, 01 Sep 2026 07:00:00 +0000
Inside Higher Ed

The Ticks and Leeches of Higher Education

The Ticks and Leeches of Higher Education kjohnsonbowles… Tue, 09/01/2026 - 03:00 AM External culprits eating the sector alive by exploiting, capitalizing and profiting off higher education and students stand to weaken and kill it. Byline(s) Kathy Johnson Bowles

Source ↗
audience Tue, 01 Sep 2026 07:00:00 +0000
Inside Higher Ed

Connecticut Workforce Office Gave Cardona $250K No-Bid Contract

Connecticut Workforce Office Gave Cardona $250K No-Bid Contract Ryan Quinn Tue, 09/01/2026 - 03:00 AM Byline(s) Ryan Quinn

Source ↗
audience Tue, 01 Sep 2026 07:00:00 +0000
Inside Higher Ed

August Cuts to Start the Academic Year

August Cuts to Start the Academic Year Josh Moody Tue, 09/01/2026 - 03:00 AM Harvard University cut more than 100 jobs last month amid a lengthy restructuring process. Multiple others also enacted smaller cuts. Byline(s) Josh Moody

Source ↗
audience Tue, 01 Sep 2026 07:00:00 +0000
Inside Higher Ed

130 Colleges Have Supplied Missing Data on Student Outcomes

130 Colleges Have Supplied Missing Data on Student Outcomes jessica.blake@… Tue, 09/01/2026 - 03:00 AM Byline(s) Jessica Blake

Source ↗
audience Tue, 01 Sep 2026 07:00:00 +0000
Inside Higher Ed

University of Pennsylvania Now Enforcing Math Prerequisites

University of Pennsylvania Now Enforcing Math Prerequisites Emma Whitford Tue, 09/01/2026 - 03:00 AM As college students’ math skills decline worldwide, Penn is verifying students’ readiness for advanced calculus and algebra courses. Byline(s) Emma Whitford

Source ↗
audience Tue, 01 Sep 2026 07:00:00 +0000
Inside Higher Ed

2-Year College ‘Extends Mission’ With 4-Year Degrees

2-Year College ‘Extends Mission’ With 4-Year Degrees Joshua.Bay Tue, 09/01/2026 - 03:00 AM Generations College is adding bachelor’s degrees, betting that students who might otherwise stop after an associate degree will keep going. Byline(s) Joshua Bay

Source ↗
audience Tue, 01 Sep 2026 07:00:00 +0000
Inside Higher Ed

Federal Judge Sides With Students Against Trump’s Deportation Policy

Federal Judge Sides With Students Against Trump’s Deportation Policy Olivia.sanchez Tue, 09/01/2026 - 03:00 AM Byline(s) Olivia Sanchez

Source ↗
audience Tue, 01 Sep 2026 07:00:00 +0000
Inside Higher Ed

Are Scholarships Based on Country of Origin in Danger?

Are Scholarships Based on Country of Origin in Danger? Johanna Alonso Tue, 09/01/2026 - 03:00 AM Such scholarships are fairly common, though one international enrollment expert acknowledged that some institutions already had doubts about offering them. Byline(s) Johanna Alonso

Source ↗
audience Tue, 01 Sep 2026 05:00:00 -0400
Higher Ed Dive

Tennessee’s free community college program ‘pays for itself,’ working paper finds

The Tennessee Promise program — the first of its kind in the country — successfully boosted college enrollment and attainment, analysis showed.

Source ↗
regulation Tue, 01 Sep 2026 05:00:00 -0400
K-12 Dive

USDA urged to report how SNAP cuts impact free school meals

As historic cuts to the Supplemental Nutrition Assistance Program take effect, a member of Congress is expressing concern over access to free school meals.

Source ↗
regulation Tue, 01 Sep 2026 05:00:00 -0400
K-12 Dive

3 things districts need to know about the $17B Meta social media settlement

States are praising the agreement's block on school-day notifications, daily time limits and other promised changes.

Source ↗
technology Tue, 01 Sep 2026 00:17:58 +0000
MedCity News

Healthcare Moves: A Monthly Summary of Hires, Exits and Layoffs

August has seen a slew of executive hires, exits and layoffs across the healthcare industry. For instance, Humana, Merck and Mayo Clinic named new executives. There were also layoffs at organizations including MaineHealth, Cellares and Sharp HealthCare. The post Healthcare Moves: A Monthly Summary of Hires, Exits and Layoffs appeared first on MedCity News .

Source ↗
behavior Tue, 01 Sep 2026 00:00:00 GMT
EdSurge

A Principal and a Student Reviewed the New ChatGPT for Teens. They Had Plenty to Say

The new offering is promising, but there are some real-world reservations.

Source ↗
behavior Tue, 01 Sep 2026 00:00:00 -0700
Nurse.org

The Biggest Nurse Strikes in America, Ranked by Impact

Part of Nurse.org's Nurse Strike Intelligence data series, built on a proprietary database tracking a decade of U.S. registered-nurse strikes (2017–2026). Updated 9/1/26 Not every nurse strike lands the same…

Source ↗
behavior Tue, 01 Sep 2026 00:00:00 -0700
Nurse.org

Nurse Voter Survey 2026: Tell Us Which Issues Matter Most in the Midterms

For more than two decades, nurses have topped Gallup's ranking of the most honest and ethical professions. The midterm elections are this November, and we want to know which issues…

Source ↗
behavior Tue, 01 Sep 2026 00:00:00 -0700
Nurse.org

One Year on the Picket Line: Inside the Henry Ford Genesys Nurses' Strike

Image source: Mid-Michigan NOW Part of Nurse.org's Nurse Strike Intelligence data series, built on a proprietary database tracking a decade of U.S. registered-nurse strikes (2017–2026). One year ago, on Labor…

Source ↗
behavior Tue, 01 Sep 2026 00:00:00 -0700
Nurse.org

The Biggest Nurse Strikes in America, Ranked by Impact

Part of Nurse.org's Nurse Strike Intelligence data series, built on a proprietary database tracking a decade of U.S. registered-nurse strikes (2017–2026). Updated 9/1/26 Not every nurse strike lands the same…

Source ↗
behavior Tue, 01 Sep 2026 00:00:00 -0700
Nurse.org

Nurse Voter Survey 2026: Tell Us Which Issues Matter Most in the Midterms

For more than two decades, nurses have topped Gallup's ranking of the most honest and ethical professions. The midterm elections are this November, and we want to know which issues…

Source ↗
behavior Tue, 01 Sep 2026 00:00:00 -0700
Nurse.org

One Year on the Picket Line: Inside the Henry Ford Genesys Nurses' Strike

Image source: Mid-Michigan NOW Part of Nurse.org's Nurse Strike Intelligence data series, built on a proprietary database tracking a decade of U.S. registered-nurse strikes (2017–2026). One year ago, on Labor…

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

MoPLEx: Estimating Plackett-Luce Mixture Models for Multi-Objective Alignment

arXiv:2608.25200v2 Announce Type: replace-cross Abstract: We study learning a mixture of $k$ Plackett-Luce models from multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. Prior work has studied mixtures of Bradley-Terry models from pairwise comparisons. However, estimating a mixture of multi-way ranking models can become theoretically unidentifiable when $k$ exceeds $m/2$, where $m$ is the ranking length. We design an efficient algorithm to address this issue by first augmenting the rankings to a larger size (e.g., generating comparisons from a base model), followed by a gradient-based estimation to reduce inference cost (in the input embedding space). With this procedure in mind, we then fit a mixture of Plackett-Luce (PL) models via an expectation-maximization-style iteration, or MoPLEx in short. We conduct extensive experiments to verify this algorithm. Fi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Query Expansion Should Be Coordinated: Dense Expands, Sparse Anchors

arXiv:2608.15851v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) systems rely on retrieval modules to ground large language model (LLM) outputs. LLM-based query expansion enriches retrieval with document-like passages, but evaluations of hybrid retrieval often fuse fixed top-L prefixes of dense and sparse rankings. Because L controls cross-channel contributions and ranking access, it can alter measured expansion gains. We therefore evaluate complete-list effectiveness and record per-channel replay stopping depths required to certify the ordered top-K. This changes the design: because both rankings determine the fused result, their query constructions should be coordinated rather than designed independently. We present DESA (Dense Expansion and Sparse Anchoring), which shares generated references across channels but specializes their integration. Orthogonal residual expansion adds new semantic directions to the dense query, whereas score-product anchoring r

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Where Knowledge and Authority Sit Changes What an Agent Benchmark Can Resolve

arXiv:2607.02975v2 Announce Type: replace-cross Abstract: Most agent benchmarks put facts, tools and permissions behind one interface. Real organizations spread them across people. Incognita asks what happens when the task and success criterion stay fixed but access does not. We transform eighteen customer-service tasks into three settings: direct access, one known intermediary, and six role-isolated participants whose capabilities must be discovered. Across 864 trials with four models, social access reduced success for every model; the pre-specified intervals excluded zero for two. The latest tested model, gpt-5.6-sol, achieved the highest social-access success at 0.65, a 0.11 decrease from centralized indirect access with an interval that included zero. Exploratory comparisons separated five of six model pairs under social access, while neither centralized setting separated any at this sample size. In post-hoc task-blocked tests, three pairwise interaction $p$-values remained signifi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models

arXiv:2606.23181v2 Announce Type: replace-cross Abstract: Hybrid reasoning models can answer directly or spend extra tokens on extended thinking. A practical router should choose between these modes for each query, so easy problems avoid unnecessary reasoning and hard problems receive enough budget to finish the answer. Existing routers move in this direction, but they typically require labeled training data or fix thinking budgets up front, ignoring answer-level evidence from the model itself. We introduce DART, a training-free routing framework that samples two cheap no-think drafts, accepts direct answering when the drafts agree, and predicts a thinking budget from draft entropy when they disagree. Across the main comparisons, DART preserves or improves always-thinking accuracy in most settings while reducing thinking-token use. Accuracy improves by up to +9.0 points on Olympiad-level math and by up to +22.5 points on code under execution-based equivalence, while thinking-token use

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

In LLM Reasoning, there is Irrationality on top of Value Misalignment

arXiv:2606.20624v2 Announce Type: replace-cross Abstract: Significant progress has been made in aligning LLMs with target value functions. We argue that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning. We mathematically formalise this gap as rational value risk: the utility discrepancy between a model's deployed reasoning strategy and its rational counterpart whose responses maximise utility in the steepest direction. The estimation error of rational value risk is further decomposed into three components from bounded prompts, bounded responses, and imperfect verifiers. Extensive experiments are conducted, covering models Llama-3.1, Qwen-2.5, T\"ulu-3 families (7B-72B), GPT-5.2, GPT-5.5, and DeepSeek-V4, and benchmarks UltraFeedback, AlpacaEval, GSM8K, MATH, HumanEval, and MathArena. The results validate that (1) rational value risk is widespread; (2) value alignment can reduce, but cannot avoid, it; (3) self-consi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

arXiv:2606.20621v2 Announce Type: replace-cross Abstract: Multi-agent debate improves the reliability of large language models (LLMs) through iterative peer critiques. However, fixed topologies often introduce persistent positional biases, amplify unreliable agents, and cause high sensitivity to role assignments. We introduce \textit{Permutation-Equivariant Adaptive Routing Multi-Agent Debate (PEAR)}, an inference-time train-free protocol that dynamically reconfigures communication roles and sparse topologies across consecutive debate rounds. By strategically switching agent-to-role assignments based on evolving agent states, PEAR prevents any agent from permanently occupying a privileged network position or distributes influence more evenly across the debate. We theoretically characterize PEAR as an equivariant sparse router: it preserves accuracy under agent relabeling while reducing routing complexity and improving generalization. Comprehensive empirical evaluations across four reas

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

arXiv:2606.06037v3 Announce Type: replace-cross Abstract: Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on monolingual, text-based harmful prompts. This leaves their generalizability under multilingual and spoken settings, particularly code-switched speech, largely underexplored. To address this gap, we introduce SpeechJBB, an audio jailbreak dataset for benchmarking state-of-the-art LALMs across five European languages: English, French, German, Italian, and Spanish, as well as code-switched variants combining pairs of these languages. The extent of safety weaknesses is further probed by introducing an augmented setting where phonologically plausible pseudo-words are inserted around safety-critical terms to simulate localized obfuscation. Across models, code-switched harmful audio yields substantially high jailbreak success rates (JSR), with non-English monolingual and non-English code-s

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference

arXiv:2606.05922v3 Announce Type: replace-cross Abstract: AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to new tasks. However, existing optimization methods typically require ground-truth validation sets, yet such labeled data is difficult to acquire in practical deployment settings. To address this problem, we introduce Retrospective Harness Optimization (RHO), a self-supervised method that optimizes the agent harness using only past trajectories. Specifically, RHO selects a diverse coreset of challenging tasks from past trajectories and re-solves them in parallel. The agent analyzes these rollouts using self-validation and self-consistency, then generates candidate harness updates and selects the most effective one by its own pairwise self-preference. We evaluate RHO across three diverse domains, spanning software engineering, technical work, and knowledge work. Notably, a single opt

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

arXiv:2606.03603v2 Announce Type: replace-cross Abstract: World models and multimodal large language models (MLLMs) provide complementary capabilities for predicting future outcomes from static visual observations. World models can generate concrete visual rollouts of possible futures, while MLLMs can reason abstractly over questions, goals, and rules. However, generated rollouts are stochastic and may be visually plausible but task-incorrect, making it necessary to determine when visual simulation is useful, whether a rollout is credible, and how it should influence the final answer. We formulate this problem as controlled concrete reasoning, where a model learns to invoke, verify, and integrate visual future simulation alongside abstract reasoning. To study this setting, we construct two human-verified benchmarks, VRQABench for controllable spatial lookahead and OpenWorldQA for open-domain physical prediction, and propose Privileged-Future On-Policy Self-Distillation (PF-OPSD). Durin

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

arXiv:2606.01600v2 Announce Type: replace-cross Abstract: Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instructions. We introduce RoboTrustBench, a benchmark for evaluating the trustworthiness of video world models under four scenarios: Normal, Constraint-Sensitive, Counterfactual, and Adversarial. Built from real-world DROID episodes, RoboTrustBench contains 1,207 expert-validated instruction-image pairs and a six-dimensional evaluation protocol with 13 fine-grained criteria. Evaluating seven representative video world models with human and MLLM assessment, we find that current models often generate visually coherent videos, but struggle with constraint reasoning, counterfactual grounding, physical interaction, and unsafe-instruction suppression. These results show that visual quality and surface-level instruction following are insufficient for trustworthy robotic video world modeling.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

arXiv:2606.01456v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed as advisors whose objective is not aligned with the user's: recommenders optimize for engagement, sales assistants for purchases. Whether they stay truthful when honesty conflicts with their own payoff is a core alignment question. We turn the canonical Crawford-Sobel cheap-talk model into a pre-specified benchmark for LLM honesty under preference misalignment, in which theory supplies an exact oracle. A sender observes a state omega in [0,1], wants the receiver's action near omega+b, and sends one costless message to a receiver whose ideal action is omega. For the positive-bias grid b in {0.01,0.04,0.08,0.12} the exact most-informative partition sizes are 7,4,3,2, with oracle normalized mutual information 0.5294, 0.3268, 0.2205, 0.1829. Extending a pre-registered 4-model run of 12,000 sender calls to eight models across two capability tiers and 39,569 logged calls, all models over

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models

arXiv:2605.30219v2 Announce Type: replace-cross Abstract: Long-horizon interactions require language models to manage accumulating information: when to update their state, when to preserve their state, and what to ignore. We study this challenge as Contextual Belief Management (CBM): maintaining a predicted belief state aligned with formal evidence while isolating task-irrelevant noise. To make CBM measurable, we introduce BeliefTrack, a closed-world benchmark spanning Rule Discovery and Circuit Diagnosis, where a finite belief space and symbolic verifiers enable exact turn-level evaluation. BeliefTrack diagnoses three failures: Failed Stay, Failed Update, and Failed Isolation. Across multiple LLMs, vanilla models exhibit severe CBM failures, while explicit belief-tracking prompts provide limited gains. In contrast, reinforcement learning with belief-state rewards reduces failure rates by 70.9% on average. Further probing reveals latent belief-state dynamics behind these failures, and

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Adopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the Wild

arXiv:2605.29018v2 Announce Type: replace-cross Abstract: Although a growing body of research has begun to describe user--LLM interactions, the picture it paints is largely static; little is known about how individual users change their behavior over time. To address this gap, we analyze the conversational trajectories of ~12,000 randomly sampled Microsoft Bing Copilot users and compare these with data from WildChat-4.8M. While the Copilot data contains significant population-level trends, we find that trends in individual user trajectories are much weaker; user habits prove to be overwhelmingly sticky. We also find stark differences between users of different activity levels: more active users have more successful conversations and use the LLM for more complex and professionally oriented tasks. Some user trends also appear in WildChat-4.8M, but we find evidence that this dataset is significantly skewed towards highly proficient "power" users. Ultimately, our results suggest that exist

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents

arXiv:2605.27955v2 Announce Type: replace-cross Abstract: Markdown skill libraries for LLM agents ship as free-form prose, forcing the agent to re-derive both the input schema and the concrete invocation syntax on every retrieval. This produces a "confused $\to$ re-retrieve $\to$ still confused" loop: the agent issues a partially-correct action, receives uninformative feedback, and re-retrieves the same prose. We propose Skill-as-Pseudocode (SaP), an automatic conversion of markdown skill libraries into typed pseudocode with deterministic quality control. From each cluster of similar procedural passages, SaP extracts a typed contract and filters it through a four-check deterministic verifier (coverage, binding, replacement, risk). Promoted contracts are inlined into a rewritten skill skeleton alongside restored action templates, giving the agent two complementary signals: a typed signature for what a skill does and a concrete template for how to invoke it. On the ALFWorld unseen split

Source ↗
Showing 5801–5850 of 18402 signals
← Prev Page 117 of 369 Next →