Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
Democratic senators are urging CMS to withdraw its Medicaid work requirements rule, arguing it will create administrative burdens and cause eligible beneficiaries to lose coverage. The post Democratic Senators Urge Withdrawal of Interim Final Rule on Medicaid Work Requirements appeared first on MedCity News .
Temporary anxiety can often surface before the first day of school. Sometimes, that can help kids get in the right mindset for a big day. But when anxiety becomes a constant feature of children’s and teenagers’ days, that stress doesn’t switch off when they leave the classroom or finish an exam. It can become the […]
New Mexico First Judicial District Judge Bryan Biedscheid at the end of the day Thursday issued a ruling in the state’s pending case against social media giant Meta, ordering the company to pay $567 million for an abatement fund to address the harms it has done to the state’s youth. The case, which ended with […]
Historically, prevention has focused on protecting teeth from decay and gum disease. However, long-term oral health is also shaped by how the teeth, muscles, and jaw function together over time. The post The Missing Piece of Preventive Dental Care Lies in the Bite appeared first on MedCity News .
For nearly two years, teacher Isabelle Ouyang and her colleagues have dealt with the presence of construction workers, shuffling between classrooms, stress over schedules, and a host of disruptive renovations. Ouyang has, essentially, pardoned all that dust, because Philadelphia’s Kensington High School now has air conditioning. With a $17 million estimated price tag, it hasn’t […]
ESSER was one of the largest federal K–12 funding initiatives in decades, empowering schools to purchase the technology needed to support learning during and after the pandemic. But ESSER funds officially ran out for all districts in early 2026. Now, many K–12 IT leaders are taking a hard look at how to make significant budget cuts without disrupting learning, because the expansion of digital programs and access to devices was a key priority. Click the banner below to explore funding options for your district’s educational technology.
First, there was the zombie college scam. Now it’s “ghost students,” bad actors leveraging artificial intelligence to create fake identities and enroll imaginary students to scoop up grants and loans. AI tools now allow synthetic identities to mimic genuine student behavior long enough to collect financial aid disbursements, and traditional enrollment systems are falling short. The schemes are costing the U.S. taxpayers hundreds of millions of dollars in siphoned-off federal financial aid packages. What Are Ghost Students, and How Is AI Making Fraud Worse? Ghost students are “fake or…
Gaps in prevention are not frequently due to a complete lack of information, but instead the fact that organizations don’t have a clear way to see how the pieces fit together until after an incident has occurred. The post Healthcare Workplace Violence Prevention Needs Better Visibility Into Risk Signals appeared first on MedCity News .
Shortly after being named America’s unhealthiest city in 2009, Huntington, West Virginia, gained national attention when British celebrity chef Jamie Oliver arrived to film the first season of his reality TV series Food Revolution. Episodes showed Oliver ridiculing the frozen food served in school cafeterias, curating new menus of made-from-scratch meals, clashing with school officials […]
The Trump administration is proposing significant changes that would diminish the program's federal standards and give states and parents more control. The post Trump administration moves to deregulate Head Start, opening door for sweeping change appeared first on District Administration .
Scheduled for October 29 at Pegasus Park in Dallas, the theme of the INVEST Digital Health conference will be consumers in healthcare from d2c diagnostics to GLP-1 drugs and wearables. Register today! The post INVEST Digital Health Conference Returns with Focus on Consumers appeared first on MedCity News .
After nearly two years out of the limelight, the Biden family has returned to the political conversation this summer. Joe Biden announced last month that he will soon release a memoir of his time in office. That volume follows on the heels of Jill Biden’s own book, which triggered controversy when some Democrats grumbled that […]
America’s schools are moving quickly to respond to the rise of artificial intelligence, with parents, teachers, administrators, and lawmakers working to wrap their arms around what this means for student education.
Insights into how school and district leaders identify needs, evaluate vendors, and manage budgets in a resource-constrained environment.
Innovative Leader Award - Melissa McCalla, CTO for Pasadena ISD, shares how her district is addressing multiple edtech challenges.
Nearly every week, I get queries from potential students about obtaining credit for previous coursework towards a degree or licensure. Often, they are unable to provide detailed records of their previous academic work beyond the transcript. The post Archiving your academic journey: Preserving coursework and program artifacts appeared first on eCampus News .
What the Tuskegee Dress Code Debate Reveals About Institutional Power Elizabeth Redden Fri, 08/07/2026 - 03:00 AM There’s a key distinction between rules and standards. Byline(s) Branden D. Elmore
NIH Has Spent 69 Percent of Its Grants Budget Katherine Knott Fri, 08/07/2026 - 03:00 AM Byline(s) Katherine Knott
U of Minnesota Pays Pro-Palestinian Holocaust Scholar After Revoked Job Offer Sara Weissman Fri, 08/07/2026 - 03:00 AM Byline(s) Sara Weissman
Star Sociology Professor Resigns After Cambridge Opens Investigation Susan H. Greenberg Fri, 08/07/2026 - 03:00 AM Jason Arday, the youngest Black professor ever appointed at Cambridge, has quit amid allegations of plagiarism and academic misconduct. Byline(s) Tom Williams for Times Higher Education
Everything Is Not Fine sara.custer@in… Fri, 08/07/2026 - 03:00 AM People hear “wellness” and think yoga, meditation and deep breathing. They say, “Everything is fine!” and go on pretending. It’s time to get honest about both. Byline(s) Jessi Gold
How HBCUs Are Redefining Student Success Joshua.Bay Fri, 08/07/2026 - 03:00 AM From helping stopped-out students return to creating faster routes to graduate school, HBCUs are removing barriers at every stage of college. Byline(s) Joshua Bay
GAO Report: Athletic Programs Bleed Money Josh Moody Fri, 08/07/2026 - 03:00 AM A new government report shows that the vast majority of colleges lose money on athletics even at the highest levels, with students and the academic enterprise subsidizing sports. Byline(s) Josh Moody
DOJ Finds Duke Law Discriminated in Admissions Johanna Alonso Fri, 08/07/2026 - 03:00 AM Byline(s) Johanna Alonso
States Wrestle With Data and Program Tweaks to Access Workforce Pell Ryan Quinn Fri, 08/07/2026 - 03:00 AM At a meeting of state higher ed and workforce officers this week, how to meet federal requirements to unlock Workforce Pell funds dominated many discussions. Byline(s) Ryan Quinn
Will a Data Center Bring Risk or Reward to Fisk? kathryn.palmer… Fri, 08/07/2026 - 03:00 AM The center is part of a $900 million plan to carry the historically Black university into the future. Now, it’s caught up in a wave of backlash against the proliferation of new data centers. Byline(s) Kathryn Palmer
It would slim regulation of the 60-year-old anti-poverty program and aim to save money. Advocates worry the changes could undermine the early education program's mission.
The conservative think tank framed the proposal as a way for state lawmakers to “codify the principles” of the Trump administration’s higher ed compact.
Despite strong enrollment, the public university said it’s facing a large budget hole due to stagnant state funding.
From one district ending universal school meals to concerns about prepping for potential disease outbreaks, what did you learn from our recent stories?
A Center on Reinventing Public Education report outlines ways to successfully transform school systems for improvement.
I've always felt that AI tutoring as we see it now is going in the wrong direction. I made Knowable as a way to see if it's possible to have a real AI tutor. Because of hardware limitations it can only be used on Macbooks 2023+. Let me know your thoughts! Comments URL: https://news.ycombinator.com/item?id=49204343 Points: 8 # Comments: 8
arXiv:2606.14695v2 Announce Type: replace-cross Abstract: Language Models (LMs) have shown remarkable potential as role-playing chatbots, delivering consistent, stylized interactions when given a specification of a character or user persona. However, applying these capabilities to real-world applications (e.g., ecosystems with numerous NPCs interacting simultaneously) exposes a critical inefficiency due to the excessive computational cost. In this paper, we question the necessity of dedicating a full, generalist model to a single persona, hypothesizing that a specific character identity relies on only a fraction of the model's total capacity. We observe that naively pruning LMs often severely degrades the role-playing performance for a specific persona; it does not distinguish between redundant knowledge and essential character traits. We propose Persona-Pruner, a framework that sculpts a lightweight role-playing model by isolating persona-specific sub-networks from a single descriptio
arXiv:2605.20247v2 Announce Type: replace-cross Abstract: Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs). Although Mixture-of-Experts (MoE) architectures offer an efficient path to scaling, existing LoRA-based MoE continual learning methods still face a fundamental trade-off: they either isolate experts too aggressively, limiting knowledge transfer across tasks, or allow task-specific updates to overwrite important existing parameters, leading to severe forgetting. To address this, we propose CP-MoE, a continual learning framework built around a transient expert that captures early task-specific updates and guides their integration into stable experts. CP-MoE introduces a consistency-preserving routing bias, which uses the transient expert to estimate representation similarity with stable experts and steer routing towards more compatible expert selection, and a transient expert-guided regularisat
arXiv:2605.16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling. We propose a stage-wise preference optimization framework for hallucination reduction through targeted multimodal data construction. Rather than directly optimizing on generic instruction-following data, our approach progressively constructs hallucination-focused preference pairs near known failure boundaries. The framework emphasizes ambiguous spatial orientation, object relationships, OCR uncertainty, and adversarial false-premise training. Hallucinated negatives are generated through minimally perturbed yet visually inconsistent alternatives, enabling Direct Preference Optimization (DPO) to better separate grounded reasoning from plausible hallucination.
arXiv:2604.22191v2 Announce Type: replace-cross Abstract: In agentic workflows, LLMs frequently process retrieved contexts that are legally protected from further training. However, auditors currently lack a reliable way to verify if a provider has violated the terms of service by incorporating these data into post-training, especially through Reinforcement Learning (RL). While standard auditing relies on verbatim memorization and membership inference, these methods are ineffective for RL-trained models, as RL primarily influences a model's behavioral style rather than the retention of specific facts. To bridge this gap, we introduce Behavioral Canaries, a new auditing mechanism for RLFT pipelines. The framework instruments preference data by pairing document triggers with feedback that rewards a distinctive stylistic response, inducing a latent trigger-conditioned preference if such data are used in training. Empirical results show that these behavioral signals enable detection of una
arXiv:2604.16742v2 Announce Type: replace-cross Abstract: Scientists have long sought to accurately predict outcomes of real-world events before they happen. Can AI systems do so more reliably? We study this question through clinical trial outcome prediction, a high-stakes open challenge even for domain experts. We introduce CT Open, an open-access, live platform that will run four challenge every year. Anyone can submit predictions for each challenge. CT Open evaluates those submissions on trials whose outcomes were not yet public at the time of submission but were made public afterwards. Determining if a trial's outcome is public on the internet before a certain date is surprisingly difficult. Outcomes posted on official registries may lag behind by years, while the first mention may appear in obscure articles. To address this, we propose a novel, fully automated decontamination pipeline that uses iterative LLM-powered web search to identify the earliest mention of trial outcomes. We
arXiv:2604.03750v2 Announce Type: replace-cross Abstract: Reverse engineering (RE) is central to software security, particularly for cryptographic programs that handle sensitive data and are highly prone to vulnerabilities. It supports critical tasks such as vulnerability discovery and malware analysis. Despite its importance, RE remains labor-intensive and requires substantial expertise, making large language models (LLMs) a potential solution for automating the process. However, their capabilities for RE remain systematically underexplored. To address this gap, we study the cryptographic binary RE capabilities of LLMs and introduce CREBench, a benchmark comprising 432 challenges built from 48 standard cryptographic algorithms, 3 insecure crypto key usage scenarios, and 3 difficulty levels. Each challenge follows a Capture-the-Flag (CTF) RE challenge, requiring the model to analyze the underlying cryptographic logic and recover the correct input. We design an evaluation framework comp
arXiv:2604.01280v2 Announce Type: replace-cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires Multimodal Large Language Models (MLLMs) to identify and combine fine-grained visual cues with retrieved textual evidence. However, retrieval often introduces noisy and partially relevant content, while images contain distracting visual regions, causing pretrained MLLMs to overlook the evidence that actually supports the answer. To address this, we introduce Look Twice (LoT), a training-free inference-time framework that turns the model's own internal attention into an explicit multimodal evidence-selection mechanism. LoT first leverages the model's internal attention patterns to identify query-relevant image regions and textual sentences, filters attention sinks and distracting content, and reformulates the input to explicitly highlight the selected evidence before answer generation. The method requires no parameter updates, auxiliary models, or architectural modificat
arXiv:2604.01170v2 Announce Type: replace-cross Abstract: While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. These inefficiencies can be attributed to the miscalibration of post-trained language models, and the lack of calibration in popular sampling techniques. Here, we present Online Reasoning Calibration (ORCA), a framework for calibrating the sampling process that draws upon conformal prediction and test-time training. Specifically, we introduce a meta-learning procedure that updates the calibration module for each input. This allows us to provide valid confidence estimates under distributional shift, e.g. in thought patterns that occur across different stages of reasoning, or in prompt distributions between model development and deployment. ORCA not only provides theoretical guarantees on conformal risks, but also empirically shows higher efficiency and generalization across differen
arXiv:2601.03895v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs). However, GRPO inherits PPO's token-level clipping while replacing token-level advantages with a single sequence-level advantage. Through a four-quadrant analysis of the (likelihood-ratio, advantage) space, we show that this combination leaves one quadrant -- negative advantage combined with an increased likelihood ratio (Q4) -- structurally unbounded, so that a few high-ratio tokens can receive very large suppressive updates that collapse entropy and narrow the reasoning boundary. To address this, we propose All-Quadrant Bounded Clipping GRPO (ABC-GRPO), which applies unconditional clipping in all four quadrants through sign-dependent boundaries. ABC-GRPO clips the likelihood ratio before multiplying by the advantage, adding a trust-region floor in Q2 and a cap in Q4 -- its negative-advantage
arXiv:2509.06861v3 Announce Type: replace-cross Abstract: Test-time scaling increases inference-time computation through longer reasoning chains and has shown strong performance gains across many domains. However, frontier models still suffer from factuality hallucinations, raising the question of whether increased computation is effective on closed-book knowledge-intensive tasks. In this work, we evaluate 14 reasoning models under different test-time scaling strategies. Our results challenge its effectiveness: increasing test-time computation does not consistently improve accuracy and often leads to more hallucinations. We find that changes in hallucination rates are largely driven by the model's willingness to answer, as longer reasoning encourages more attempts, many of which are incorrect. We also observe patterns consistent with confirmation bias, where extended reasoning reinforces early incorrect beliefs with fabricated details. Finally, we provide an information-theoretic persp
arXiv:2608.04586v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from the curse of multilinguality: languages at different resource levels compete for limited representation capacity, leading to strong high-resource performance but substantial degradation on low-resource speech. To address this problem and improve multilingual consistency, we propose MSRT, a novel framework built around a resource-aware Mixture of Speech Encoders (MoSE). MoSE uses an explicit language router to assign each utterance to an appropriate expert encoder. A frozen expert preserves high-resource language capabilities, while a trainable expert adapts to and specializes in medium- and low-resource languages. We further introduce a five-stage curriculum learning strategy that substantially reduces data
arXiv:2608.02616v2 Announce Type: replace Abstract: We present what is, to our knowledge, the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter model that converts an autoregressive language model into a bidirectional PII detector, across 32 benchmarks spanning 14 languages and 5 domains. Our most practically actionable finding is a domain-dependent labeled-data crossover: fine-tuned XLM-RoBERTa surpasses OPF's zero-shot performance with only ~500 labeled examples on English synthetic PII (~100 on non-English Kiji), and ~1000 on synthetic medical PII. Crucially, per-class fine-tuning (17 PII entity types, a subset of OPF's 33) is less data-efficient than binary labels at small n -- at n=100, binary F1=0.634 vs. per-class 0.360. Zero-shot, OPF achieves F1=0.464 on the SPY medical benchmark and F1=0.855 on AI4Privacy, substantially outperforming Presidio and XLM-RoBERTa-large-NER. However, OPF degrades sharply outside its PII training distribution: F1=0.04--0
arXiv:2608.01347v3 Announce Type: replace Abstract: Two prompts can request the same code change and produce the same correct patch, yet cause a coding agent to perform radically different kinds and amounts of work. We study this effect in a preregistered benchmark spanning 4,644 valid runs, 24 deterministic coding tasks, seven reasoning models, and two real agent harnesses. The central finding is that prompt wording does not merely scale total effort; it changes where that effort is spent. Multiple approaches and deep thinking primarily inflate reasoning. Multiple approaches increases reasoning by 2.4x to 7.4x across all six open models and creates about three elaborated but discarded solution branches, while still yielding only one implemented solution and no success gain. Maximum certainty activates a different pathway: repeated verification propagates into extra test runs, tool calls, turns, latency, and context growth. Runs with high redundant verification cost 18x the clean-run m
arXiv:2606.21155v2 Announce Type: replace Abstract: Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions would deter negligent filers, we found over 1,000 filings containing fabricated citations---with this number growing year-over-year. This study evaluates whether AI-based systems can mitigate these errors by automatically detecting hallucinations. We propose a taxonomy of legal citation hallucinations grounded in actual court filings and introduce a dataset of 1,300 brief excerpts containing injected errors. Benchmarking five models in agentic and non-agentic settings as well as Claude Code reveals that while the latest iterations perform better---GPT-5 achieves 84.4% recall and a 55.0% F1 score in an agentic framework---all models struggle with subtle error categories. Agentic verification remains resource-intensive, with
arXiv:2606.13227v2 Announce Type: replace Abstract: Post-training methods such as supervised fine-tuning (SFT) and preference optimization typically align language models toward a single global assistant behavior. While effective for improving average helpfulness, this can suppress the natural variation of human responses across languages, tasks, and dialogue settings. We study this problem as conditional human-distribution alignment: models should match the human response distribution appropriate to the current interaction context, rather than a universal response style. We introduce PolyAlign, a distribution-aware alignment framework that organizes bilingual interaction data into bucket-specific human reference distributions defined by language, interaction track, response family, and length. PolyAlign combines Bucket-Aware SFT, which balances optimization across heterogeneous buckets, with Human-Distribution Preference Optimization (HDPO), which regularizes preference learning using
arXiv:2606.10921v2 Announce Type: replace Abstract: Long-document question answering (QA) requires large language models (LLMs) to reason over evidence scattered across lengthy documents, where answers often depend on event order, section-level context, and cross-part evidence connections. Although retrieval-augmented generation (RAG) reduces the input context by retrieving relevant evidence, existing structured RAG methods still face three limitations: costly query-agnostic knowledge organization, insufficient use of original document structure, and no reuse of historical reasoning experience. To address these limitations, we propose DocTrace, a multi-agent RAG framework for long-document QA that supports query-triggered knowledge organization, document-structure-aware and experience-guided reasoning. DocTrace preserves document hierarchy with a lightweight document structural tree index, constructs agent-shared hypergraph-structured working memory on demand during reasoning, and stor
arXiv:2606.02776v4 Announce Type: replace Abstract: When large language models (LLMs) are used in high-stakes scenarios, such as legal, medical and financial advice, even a single conversation history is enough to drive differences in outcomes between users. Prior work has demonstrated that this results in outcome disparities between sociodemographic groups, with some groups receiving more advantageous outcomes than others. In this work, we demonstrate that LLMs actually struggle to infer user sociodemographics from a single conversation history and that although there are disparities between sociodemographic groups, they are minimal in magnitude. To investigate what is the main driver of disparities between users, we compare user sociodemographics to a range of (psycho)linguistic features of conversations, including conversation topic, emotions, and readability. We find that conversation topics are most predictive of LLM-generated advice within a conversational context, which, to some
arXiv:2605.12519v2 Announce Type: replace Abstract: Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifiable rewards typically optimizes only final outcomes, which can improve task accuracy at the expense of reasoning quality, producing inaccurate, incomplete, or inconsistent traces. We propose verifiable process supervision (VPS), a post-training framework that jointly optimizes prediction accuracy and reasoning quality by supervising structured intermediate claims. We first apply supervised fine-tuning to induce a structured reasoning format, enabling deterministic extraction and verification of intermediate claims for process-level rewards. To address the heterogeneous difficulty of reasoning subtasks, we introduce adaptive weighting that prioritizes components with the largest remaining errors, creating an implicit curriculum. We evaluate VPS on chess as a controlled testbed where reasoning steps