Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
When districts adopt evidence-based practices like Structured Literacy, it’s often with a surge of excitement and momentum. Yet the real challenge lies not in the initial adoption, but in sustaining and scaling these practices to create lasting instructional change.
With the new school year now rolling, teachers and school leaders are likely being hit with a hard truth: Many students are not proficient in reading.
July has seen a slew of executive hires, exits and layoffs across the healthcare industry. For instance, Aledade, Providence and Whoop named new executives. There were also layoffs at organizations including Novartis, Adventist Health and Wellstar Health System. The post Healthcare Moves: A Monthly Summary of Hires, Exits and Layoffs appeared first on MedCity News .
Grove AI’s technology uses voice agents to recruit eligible patients for clinical trials. Formed in 2024, the startup’s rapid growth led to its acquisition by Hippocratic AI earlier this year. The post Startup to Acquisition in 2 Years: Grove AI’s Journey to Transform Clinical Trials appeared first on MedCity News .
Daffodil Health launched an AI-powered solution to help payers manage No Surprises Act disputes by automating claims review, negotiations and arbitration workflows. The post Daffodil Health Launches No Surprises Act Dispute Management Solution appeared first on MedCity News .
CHICAGO — A bipartisan group of state lawmakers is calling for an overhaul of the nation’s childcare system, which they say is failing children, parents and providers. The group of 13 Republican and Democratic lawmakers, who have spent the last year studying childcare access and affordability problems, offered an array of policy recommendations during the […]
A major and final piece of West Virginia lawmakers’ 2023 hallmark education legislation goes into effect this school year when schools are required to retain third graders whose reading and math abilities aren’t at grade level. Lawmakers carved out some exceptions for students with disabilities and other young learners. The robust education bill, known as […]
The state grants will cover up to 45% of the cost for large-scale projects, such as implementing energy efficiency measures.
Similar proposals have kicked around for years in Congress, but they’ve yet to become law despite sharpened public scrutiny of the practice.
Similar proposals have kicked around for years in Congress, but they’ve yet to become law despite sharpened public scrutiny of the practice.
The district offered free school meals for the past four years through the Community Eligibility Provision but no longer qualifies for the federal program.
The department will no longer include information on nonbinary students and has excluded "gender identity" from its definition of rape.
Short of dropping a dead weight copy of the state Education Code on your foot, a State Board of Education agenda item on the California School Dashboard normally wouldn’t evoke tears. But Jessica Sawko’s tears earlier this month were spontaneous and heartfelt during her one-minute testimony before the State Board of Education. They expressed joy […]
Caring for patients with ADHD, or those who potentially have ADHD, extends far beyond diagnosis and prescribing. The post The Battles We Face: ADHD From A Patient Perspective, And What Providers Should Know appeared first on MedCity News .
What healthcare needs now is a connected, deeply integrated system of action that works with the system of record to complete work across the enterprise, supported by accountable partners willing to stand behind the outcomes they create. The post Three Ways Distressed Healthcare Must Evolve appeared first on MedCity News .
A lawsuit filed by three parents argues that the displays violate a state religious freedom law and asks for them to be removed before school starts in August. The post Texans Try new tactic to remove Ten Commandments from schools appeared first on District Administration .
To these teens, the stakes for getting ahead of AI are clear. They pointed out how adult leaders were slow to protect students from the risks of other tech developments. The post Adults have struggled to set rules for AI in school. These teens figured it out appeared first on District Administration .
In the Pilsen neighborhood on Chicago’s West Side, there’s a mural of my grandfather selling ice cream. My dad worked construction and sold ice cream too, and he and my mom both spent years working in factories. The mural is a tangible reminder of where I came from and how far my family has traveled […]
Anthropic recently rolled out its Claude for Teachers tool, making a free version of Claude with advanced technical capacities available to any K-12 U.S. educator who applies to the company using a school email address and asserts that CfT will be used for educational purposes. The tool has many features that will make it appealing […]
To ensure children build foundational reading skills in the early elementary years, educators have embraced the Science of Reading with its emphasis on phonics and phonemic awareness.
Function Health raised $450 million in growth financing from General Catalyst’s Customer Value Fund, just eight months after its $298 million Series B. The lab-testing startup plans to use the capital to expand access to its testing and imaging services. The post Why General Catalyst Is Betting $450M on Function Health’s Preventive Care Push appeared first on MedCity News .
Innovative Leader Award - Matt Kuhn of Volusia County Schools shares how his district has implemented AI for use by students, staff, and leaders.
Grade inflation is all the rage in higher education. Harvard has a proposal to reduce the number of A’s it awards. A Yale faculty committee proposed that 3.0 should be the mean grade. The post Higher education’s grade inflation conundrum appeared first on eCampus News .
Keep Strunk & White in the Bathroom and Other Ways to Obsess Over Writing Well sara.custer@in… Fri, 07/31/2026 - 03:00 AM A working writer describes rituals for chasing the tension and surprise that make stories worth reading. Byline(s) Susan D’Agostino
Florida English Professor Fired for Teaching ‘Political’ Story Sues College Emma Whitford Fri, 07/31/2026 - 03:00 AM Byline(s) Emma Whitford
Dissolved New College of Florida Alumni Board Not Giving In, Chair Says Ryan Quinn Fri, 07/31/2026 - 03:00 AM Byline(s) Ryan Quinn
Personal Aphorisms Sara Brady Fri, 07/31/2026 - 03:00 AM When in doubt … Byline(s) Matt Reed
U.S. Universities Among 7 Approved to Set Up Greek Branch Campuses sara.custer@in… Fri, 07/31/2026 - 03:00 AM Institutions from the U.S., U.K. and France have been given the go-ahead following a controversial change in the law governing private institutions. Byline(s) Seher Asaf for Times Higher Education
It’s Never Too Late to Go Back to College Joshua.Bay Fri, 07/31/2026 - 03:00 AM In this week’s Voices of Student Success episode, ReUp CEO Terah Crews explores why students stop out and what colleges can do to help them return. Byline(s) Joshua Bay
Report: Struggling Florida High Schoolers Take CLT to Graduate Johanna Alonso Fri, 07/31/2026 - 03:00 AM Byline(s) Johanna Alonso
A Green Mountain College Resurrection? Josh Moody Fri, 07/31/2026 - 03:00 AM An evangelist wants to reopen a closed campus in rural Vermont. But the finances and strategy are unclear, and his prior attempt at running a college raises questions. Byline(s) Josh Moody
Critics of the AAUP and the Delusions of Neutrality Sara Brady Fri, 07/31/2026 - 03:00 AM The AAUP is not and must not be neutral in defense of academic freedom. Byline(s) John K. Wilson
GOP Rift Over Education Department’s Future jessica.blake@… Fri, 07/31/2026 - 03:00 AM In a 13-to-9 vote Thursday, two Republicans sided with the Democrats to protect four key education offices. Byline(s) Jessica Blake
Under the initiative, companies will provide funding so students can conduct dissertation research at their sites for at least a year.
Veteran educators are among those most likely to exit, leaving leaders to rely on novice teachers at a time when students need stability and experience.
From enrollment trends to student eligibility for the new federal school choice program, what did you learn from our recent stories?
As a researcher, I knew the evidence. As a father, I learned that when making decisions for your child, emotions matter as well.
arXiv:2605.26494v2 Announce Type: replace-cross Abstract: We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modify
arXiv:2605.24973v2 Announce Type: replace-cross Abstract: VLM-based OCR models have become the de facto choice for document parsing, as they can accurately extract page-level elements (e.g., paragraphs within individual pages) together with their bounding boxes and textual content. However, downstream applications such as RAG require coherent document-level information, whereas these models often break cross-page continuity and fail to recover disrupted structures, such as paragraphs and tables truncated by page boundaries. Such relationships are not confined to a single page; instead, they require joint analysis of titles, paragraphs, tables, and images spanning multiple pages. A natural solution is therefore to reuse existing OCR outputs and reconstruct document-level logical structures through post-processing. To this end, we propose MinerU-Popo, a lightweight and universal framework for POst-Processing OCR outputs, which converts page-level results from diverse parsers into coheren
arXiv:2605.15040v3 Announce Type: replace-cross Abstract: Agentic modeling aims to transform LLMs into autonomous agents capable of solving complex tasks through planning, reasoning, tool use, and multi-turn interaction with external environments. We present Orchard, an open-source framework for scalable agentic modeling. At its core is Orchard Env, a lightweight Kubernetes-native environment service that provides reusable primitives for sandbox lifecycle management across task domains, agent harnesses, and training stages. On top of Orchard Env, we build three agentic modeling recipes. Orchard-SWE targets software engineering agents. We introduce credit-assignment supervised fine-tuning and a progression of RL signals: Balanced Adaptive Rollout (BAR) for sparse-reward optimization, on-policy distillation (OPD) and rubric-based process reward (RPR) for dense supervision, and historical experience distillation, which compresses rollouts from prior experiments into a compact value model
arXiv:2604.16557v2 Announce Type: replace-cross Abstract: Current post-training methodologies for adapting Large Vision-Language Models (LVLMs) generally fall into two paradigms: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). Despite their prevalence, both approaches suffer from inefficiencies when applied in isolation. SFT forces the model's generation along a single expert trajectory, often inducing catastrophic forgetting of general multimodal capabilities due to distributional shifts. Conversely, RL explores multiple generated trajectories but frequently encounters optimization collapse - a cold-start problem where an unaligned model fails to spontaneously sample any domain-valid trajectories in sparse-reward visual tasks. In this paper, we propose Supervised Group Relative Policy Optimization (S-GRPO), a unified post-training framework that integrates the guidance of imitation learning into the multi-trajectory exploration of preference optimization. Tailored for di
arXiv:2602.08159v2 Announce Type: replace-cross Abstract: When a language model asserts that "the capital of Australia is Sydney," does it know this is wrong? Models assert misconceptions with the same fluency as facts, so the question cannot be answered from output uncertainty. Truth-related signals are known to exist in the residual stream, but not their geometry: how many dimensions carry the signal, how simple a detector can be, and whether it transfers. We characterize this geometry across 11 models (124M-14B) and test it causally with activation steering, concept erasure, and distributed alignment search. The structure is simple: two class centroids in a 2-8 dimensional subspace match a trained linear probe, and 25 labeled examples recover 90% of full-data AUC on GPT-2. Steering shifts hallucination rates by 9.1 points on six models, erasure drops detection to chance, and distributed alignment search, the only method that bounds rank, localizes at most five causal dimensions. The
arXiv:2601.17151v2 Announce Type: replace-cross Abstract: Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit major competency gaps in multimodal understanding and reasoning especially in high-value verticals such as biomedicine. Medical imaging report generation is a prominent example. Supervised fine-tuning can substantially improve performance, but they are prone to overfitting to superficial boilerplate patterns. In this paper, we introduce Universal Report Generation (UniRG) as a general framework for medical imaging report generation. By leveraging reinforcement learning as a unifying mechanism to directly optimize for evaluation metrics designed for end applications, UniRG can significantly improve upon supervised fine-tuning and attain durable generalization across diverse institutions and clinical practices. We trained UniRG-CXR on publicly available chest X-ray (CXR) data and conducted a t
arXiv:2502.20780v2 Announce Type: replace-cross Abstract: The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which the models may generate seemingly plausible results that are in fact incorrect. Such hallucinations can jeopardize clinical decision making, potentially harming the diagnosis and treatments. In this work, we propose MedHallTune, a large-scale benchmark designed specifically to evaluate and mitigate hallucinations in medical VLMs. Comprising over 100,000 images and 1,000,000 instruction pairs, MedHallTune includes both hallucination and non-hallucination samples, each with ground-truth annotations. We conduct a comprehensive evaluation of current medical and general VLMs using MedHallTune, assessing their performance across key metrics, including clinical accuracy, relevance, detail level, and risk level. The experimental results show that fine-tuning with MedHallTune successfully improves t
arXiv:2403.18591v3 Announce Type: replace-cross Abstract: Broadcast protocols are programs designed to be executed by networks of processes. Each process runs the same protocol, and communication between them occurs in synchronously in two ways: broadcast, where one process sends a message to all others, and rendez-vous, where one process sends a message to at most one other process. In both cases, communication is non-blocking, meaning the message is sent even if no process is able to receive it. We consider two coverability problems: the state coverability problem asks whether there exists a number of processes that allows reaching a given state of the protocol, and the configuration coverability problem asks whether there exists a number of processes that allows covering a given configuration. These two problems are known to be decidable and Ackermann-hard. We show that when the protocol is Wait-Only (i.e., it has no state from which a process can both send and receive messages), th
arXiv:2607.26497v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled study that varies corpus size along 28 strictly nested tiers spanning roughly 450-fold, while holding questions and a fixed bedrock of relevant and adversarial documents unchanged. Under one reader model and one judging protocol, we measure official accuracy, construction and query tokens, and latency. The results reveal a scale-dependent crossover rather than an unconditional winner. File-System Agent leads at the smallest shared tiers, but its sequential exploration costs 39 times more query tokens at the bedrock and becomes less effective as the search space grows. Around 10 million corpus tokens, BM25 overtakes it and leads at every larger shared
arXiv:2607.02383v3 Announce Type: replace Abstract: LLM-based retrieval-augmented generation (RAG) is increasingly used for automated fact-checking (AFC) and related tasks. By grounding LLM outputs in retrieved evidence, RAG-based systems provide transparent justifications while allowing external information to be updated independently of the underlying model. However, existing approaches often assume retrieved evidence is reliable, although real-world information may be conflicting, outdated, and can originate from unreliable or biased sources. Recent work on *source-critical reasoning* addresses this challenge through media background checks (MBCs) (Schlichtkrull, 2024), which assess the credibility of evidence sources to support downstream fact verification. However, generating MBCs relies on costly proprietary search APIs, limiting reproducibility. To mitigate this issue, we introduce MEDIAREF, a publicly available knowledge store of web-sourced documents that enables reproducible,
arXiv:2605.10579v2 Announce Type: replace Abstract: Evaluating whether AI agents can proactively assist humans in daily activities, ranging from routine household tasks to urgent safety-critical situations, requires diverse visual data. However, collecting such scenarios in the real world is often difficult, costly, or unsafe, and simulation environments often lack the social commonsense needed to simulate the consequences of different actions. In this work, we present VISTA, a controllable platform that uses a user-provided scenario seed, defined as a short natural-language description of the intended assistance situation, to generate editable plans, egocentric videos, and an auditable review trail. VISTA structures scenario intent around three interaction modes, including reactive, explicit proactive, and implicit proactive, and two consequence families, including safety-critical and everyday inconvenience, with no-assistance cases as controls. Its six-stage pipeline exposes the desi
arXiv:2605.00086v2 Announce Type: replace Abstract: High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTimbau and Albertina PT-BR, we introduce NorBERTo, a modern encoder based on the ModernBERT architecture, featuring long-context support and efficient attention mechanisms. NorBERTo is trained on Aurora-PT, a newly curated Brazilian Portuguese corpus comprising 331 billion GPT-2 tokens collected from diverse web sources and existing multilingual datasets. We systematically benchmark NorBERTo against Strong baselines on semantic similarity, textual entailment and classification tasks using standardized datasets such as ASSIN 2 and PLUE. On PLUE, NorBERTo-large achieves the best results among the encoder models we evaluated, notably reaching 0.9191 F1 on MRPC and 0.7689 accuracy on RTE. On ASSIN 2, NorBERTo-large attains the highest entailment F1 (~0.904) among all encoders considered, alt