Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2606.28332v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains poorly understood. We introduce \textsc{MedHarm}\footnote{Code and data will be released upon acceptance. Due to the sensitive nature of high-risk medical queries, data access will be available to qualified researchers upon request.}, a high-risk medical safety benchmark with 1,100 medically grounded queries across 10 safety-critical categories, including toxicology, pharmacology, covert poisoning, anesthesia, and fetal harm. Unlike broad medical QA benchmarks, \textsc{MedHarm} targets realistic clinical, educational, and technical prompts that require refusal, caution, or safe redirection rather than direct helpfulness. We evaluate 15 LLMs spanning general-purpose, medical-purpose, closed-source, and downstream SFT models, together with 4 representative guardrail models. Results reveal a sub
arXiv:2606.28331v1 Announce Type: new Abstract: The widespread deployment of generative artificial intelligence (AI) models has raised serious concerns about the proliferation of AI-generated content. This has led to a surge of interest in, and demand for, reliable tracking and detection mechanisms for content that is AI-generated, such as watermarking, metadata tagging, content tagging, and more. The problem has captured the attention of policymakers as well as the popular media, and a spate of recent bills in the US have sought to regulate the spread of AI content, and enforce or promote methods to track and label it. This work performs a critical analysis of the policy discourse surrounding generative AI content transparency in the US and EU. Through a broad document selection methodology, we first collect a broad corpus of documents containing legislative language and policy-relevant discourse on the topic. We then analyze these through inductive coding, and leverage our coding to
arXiv:2606.28325v1 Announce Type: new Abstract: Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a seven-axis measure of digital support, and applied it to the 300 writing systems of the Global Script Database (Fukui, 2026). Only 29 scripts (9.7%) are fully supported by contemporary digital infrastructure; among 158 living scripts, 60 (38.0%) lack complete support. Tokenizer efficiency varies by a factor of 31.7 across 45 scripts measured with parallel text. A serial mediation model -- imperial intervention to speaker population to web corpus to tokenizer efficiency -- is consistent with full mediation, with the direct effect of empire indistinguishable from zero (beta = -0.22, p = 0.39) and structural equation model fit indices indistinguishable from saturation at n = 45; the bias-corrected bootstrap CI grazes zero, and we treat the mediation as suggestive rather than confirmatory. Across
As generative AI technologies evolve, educators are moving away from fears about AI-enabled cheating and are embracing the idea that AI can open new doors for teaching and learning.
In the growing conversation around AI in education, speed and efficiency often take center stage, but that focus can tempt busy educators to use what’s fast rather than what’s best.
Article URL: https://www.codepuzzle.io/html-studio/2RC7MY56 Comments URL: https://news.ycombinator.com/item?id=45728493 Points: 1 # Comments: 0
Article URL: https://papertalk.org/papertalks/31999 Comments URL: https://news.ycombinator.com/item?id=40501840 Points: 1 # Comments: 0
Throne Science — a startup selling an AI-powered toilet camera that tracks gut health, hydration and urinary function — closed a $10 million Series A funding round. CEO Scott Hickle said the goal is to become a “smoke detector for colorectal cancer,” filling a monitoring gap he noted the wearables industry has largely ignored. The post Smart Toilet Sensor Startup Is Flush with $10M Funding appeared first on MedCity News .
Included Health announced plans to acquire Firefly Health to accelerate its alternative health plan strategy for employers. The post Included Health to Acquire Firefly Health to Expand Alternative Health Plan Offering for Employers appeared first on MedCity News .
Altimmune’s pemvidutide achieved a statistically significant reduction in heavy drinking days compared to a placebo in a Phase 2 trial. This peptide drug, which is designed to activate the GLP-1 and glucagon receptors, is also on track to begin a pivotal test in the fatty liver disease MASH. The post Altimmune’s GLP-1 Drug Reduces Heavy Drinking in Alcohol Use Disorder Trial appeared first on MedCity News .
When the leaders of a new elementary school applied for a four-digit public school code from the Colorado Department of Education last summer, they didn’t mention a key detail: The school would be religious. Instead, they said Riverstone Academy would offer traditional academics and trade-themed electives. But education department officials soon learned that the 30-student […]
The ruling marks another victory for the U.S. Department of Justice, which has sued more than a dozen states over similar policies.
A teacher has issues accessing a learning platform. Students struggle to log in to an online assessment. A classroom full of devices suddenly loses connectivity. For K–12 IT teams, solving those problems has become increasingly complicated, with today’s tech environments spanning on-premises networking equipment, cloud-hosted applications, Software as a Service (SaaS) platforms and thousands of endpoint devices spread across multiple schools. When an issue occurs, the root cause could reside almost anywhere. That complexity is pushing many districts to look beyond traditional monitoring tools…
As New Hampshire public schools grapple with funding challenges, declining enrollments, and increasing fiscal scrutiny from voters, some metrics suggest that the schools continue to perform well nationally. On Monday, Gov. Kelly Ayotte touted the latest analysis; a study by WalletHub that used a mix of proficiency scores, graduation rates, advanced placement course takeup, attendance, […]
For as long as anyone remembers, no child care center on Kauaʻi has had a single opening for an infant or toddler. Not one. On this tropical island lush with vegetation, the child care landscape is desolate. Kauaʻi is the only major Hawaiian island that has gone without center-based infant and toddler care for most […] The post The island without child care for babies appeared first on The Hechinger Report .
For as long as anyone remembers, no childcare center on Kauaʻi has had a single opening for an infant or toddler. Not one. On this tropical island lush with vegetation, the childcare landscape is desolate. Kauaʻi is the only major Hawaiian island that has gone without center-based infant and toddler care for most of the […] The post The island without childcare for babies appeared first on The Hechinger Report .
Article URL: https://www.cato-unbound.org/2009/04/13/peter-thiel/education-libertarian/ Comments URL: https://news.ycombinator.com/item?id=49085418 Points: 2 # Comments: 3
Chronic absenteeism is playing a major role in students failing to reach pre-pandemic academic levels, a new report from NWEA shows. The report, issued last week, shows that the academic progress of students in districts with the highest rates of chronic absenteeism are more likely to lag behind districts that haven’t experienced drastic rises in […]
Done well, NICU UM is so much more than an approval process. It is a clinical credibility engine that — when integrated with Case Management (CM) — drives every NICU admission toward the most appropriate, impactful care. The post Why Specialized UM Delivers on Every NICU Outcome appeared first on MedCity News .
A student being bullied at school. Struggling with their homework. Bringing a gun to school. These are familiar scenarios, but what happens when the person encountering them first isn’t a person at all but an AI-programmed teaching assistant? The post I interviewed Realbotix’s AI teaching assistant. Here’s what happened. appeared first on District Administration .
A few minutes of interaction with a digital tablet and speaking can reveal patterns that once required advanced imaging, years of observation, or were just completely undetectable. The post Spotting Alzheimer’s Years Earlier Through AI and Clinically Meaningful Insights appeared first on MedCity News .
For the second year in a row, the U.S. Department of Agriculture will let schools seek a temporary “accommodation” if they can’t meet new limits on foreign food in school meals. The post Big banana shortage avoided in school cafeterias this year appeared first on District Administration .
There’s a long-held misconception that students’ opportunities are constrained by the size of their town. In many people’s minds, rural career and technical education and career-connected learning are synonymous with agriculture and manufacturing, while college prep is limited to a path from the graduation stage to the nearest state school. Unfortunately, this mentality can be […]
Nearly 6 million children in California would be eligible for a scholarship under the new federal tax credit for education, more than in any other state, new data shows. But unlike other states with millions of eligible students, like Texas, Florida and New York, California hasn’t opted into the new Treasury Department program, and so […]
Therapy, mentoring and professional development over a meaningful period of time can help teachers better show up for themselves and their students.
Walk into almost any elementary classroom today and you will likely hear the same concern from teachers across grade levels: Many students can swipe, tap, and navigate a screen with ease, but struggle to hold a pencil comfortably, form letters fluently or sustain writing for more than a few minutes without fatigue.
Field trips have always been more than a day away from desks. They are the moments students remember, the experiences that build empathy and ignite curiosity. A new Getting Smart post examines what the research says about virtual field trips and offers a classroom-ready framework that helps educators design experiences worthy of that same legacy. For leaders looking to expand access and deepen engagement without waiting for a bus, this is a practical and visionary read. The post Beyond the Classroom Walls: What Research Says About Virtual Field Trips appeared first on Getting Smart .
Has Apple finally built a laptop that belongs in the Chromebook conversation?
Stop Telling Students Computer Science Is Dying Elizabeth Redden Tue, 07/28/2026 - 03:00 AM The data says otherwise. Byline(s) Christine Julien
3 Questions for SFBU President Nick Ladany on MBA Pro+ joshua.m.kim@d… Tue, 07/28/2026 - 03:00 AM A conversation about San Francisco Bay’s new $10,000 M.B.A. Byline(s) Joshua Kim
How Will New Student Visa Rules Impact Athletics? Johanna Alonso Tue, 07/28/2026 - 03:00 AM NCAA policy allows athletes to compete for five seasons, but new federal rules limit international students to just four years in the U.S., cutting off collegiate play for some. Byline(s) Johanna Alonso
Rider University Ends Chinese University Partnership Emma Whitford Tue, 07/28/2026 - 03:00 AM Byline(s) Emma Whitford
Compton College Plans to Tie Basic Needs to Academic Milestones. Will It Work? Sara Weissman Tue, 07/28/2026 - 03:00 AM Starting in fall 2027, students need to meet a set of academic requirements to get free meals and other supports. College leaders say the move will increase completion rates. But basic needs experts are worried. Byline(s) Sara Weissman
Board Dismisses Florida Keys President Josh Moody Tue, 07/28/2026 - 03:00 AM Byline(s) Josh Moody
DOJ Sues Colorado Over In-State Tuition for Undocumented Students Johanna Alonso Tue, 07/28/2026 - 03:00 AM Byline(s) Johanna Alonso
Why One Professor Abandoned the AI Resistance kathryn.palmer… Tue, 07/28/2026 - 03:00 AM Scott Latham has embraced the new technology and shifted his focus to measuring the AI readiness of colleges and universities. Byline(s) Kathryn Palmer
FSA Staff Slated to Move to Treasury Office Starting in August jessica.blake@… Tue, 07/28/2026 - 03:00 AM Byline(s) Jessica Blake
The Disappearing Student Safety Net Joshua.Bay Tue, 07/28/2026 - 03:00 AM New analysis from the Hope Center found that even though more colleges are offering aid for emergencies, awareness gaps and limited funding mean fewer students are receiving support. Byline(s) Joshua Bay
The nation’s first and sweeping school choice program kicks in next year, when 51.7 million students are expected to be eligible, according to a report.
Nearly half of those polled said they would rather invest in AI tools than hire and train a recent college graduate.
Institutions can immediately consider Classic Learning Test scores for certain applicants as the state higher education board looks to revise its policies.
School systems like Spokane Public Schools are boosting extracurricular activities — and hoping payouts from litigation will sustain the investments.
ASHLAND, Ore. — When Ulysses McCready was in fifth grade, a group of Southern Oregon University music students hauled instruments to Orchard Hill Elementary School to help get local kids excited about joining the middle school band. It worked on McCready, who picked up the flute and soon fell in love with music. “That day […] The post Could a public campus really close? The fight to save Southern Oregon University appeared first on The Hechinger Report .
A high school student creates an online class for middle schoolers — and gains perspective on teaching.
arXiv:2607.22083v2 Announce Type: replace-cross Abstract: We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diversity of executable environments, task assets, and agentic scaffolds through real-world deployment and large-scale synthesis. Our RL pipeline applies mixed-mode RLHF over Think and Non-Think responses to improve overall model quality and reduce failure cases, length-controlled reasoning RL to balance accuracy and reasoning efficiency, and agentic RL with outcome and process rewards to stabilize long-horizon training. Extensive evaluations
arXiv:2607.13408v2 Announce Type: replace-cross Abstract: Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limited supervision on instruction-level correctness. We propose an instruction-level framework that uses audio-aware large language models (ALLMs) as fine-grained judges to verify target event presence and temporal relations in generated audio. After validating ALLM judgments on benchmarks and through human verification, we use their feedback to construct preference pairs for direct preference optimization. We further introduce S3Bench, a narrative benchmark for evaluating multi-event temporal instruction following. Experiments show that our method improves event completeness, temporal ordering, and joint instruction-following accuracy across existing benchm
arXiv:2606.18037v2 Announce Type: replace-cross Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, databases, clinical records, and formulary tools. Standard factuality metrics usually test whether an answer is supported by pooled evidence, missing a provenance-sensitive failure mode: a claim may be supported somewhere while being attributed to the wrong source. We call this cross-source conflation. We introduce ProvenanceGuard, a source-aware verifier for MCP-grounded answers. It consumes captured MCP traces with stable tool IDs, source IDs, and raw outputs; decomposes answers into atomic claims; routes claims to source-specific evidence; checks support with NLI and a token-alignment proxy; compares stated attribution with the routed source; and returns per-claim verdicts plus an answer-level allow/block decision. Blocked answers can be repaired with retrieval-augmented answer revisio
arXiv:2605.09877v4 Announce Type: replace-cross Abstract: Recall presents a difficult choice: transformers have a linearly growing memory that slows each successive token, while linear RNNs typically have fixed costs but limited recall. We present Key-Value Means ("KVM"), a novel block-recurrence for attention that can accommodate either fixed-size or growing state. Equipping a strong transformer baseline with fixed-size KVM attention layers yields a strong $O(N)$ chunked RNN, while adding only an insignificant number of new parameters. We train a transformer with a growable KVM cache and show it performs competitively on long-context tests with only subquadratic prefill time and sublinear state growth. KVM is implementable with standard operations and without custom kernels, and supports chunk-wise parallelizable training and prefill. It provides many of the benefits of both traditional transformers (expandable context memory, chunk-wise parallelizable training and prefill) and RNNs i
arXiv:2601.09624v2 Announce Type: replace-cross Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably across individual samples: some are reliably erased, while others persist despite the same procedure. We argue that this disparity is not only a data-side phenomenon, but also reflects model-internal mechanisms that encode and protect memorized information. We study this problem from a mechanistic perspective based on model circuits--structured interaction pathways that govern how predictions are formed. We propose Circuit-guided Unlearning Difficulty (CUD), a {\em pre-unlearning} metric that assigns each sample a continuous difficulty score using circuit-level signals. Extensive experiments demonstrate that CUD reliably separates intrinsically easy and hard samples, and remains stable across unlearning methods. We identify key circuit-level patterns that reveal a mechanistic signature of di
arXiv:2507.22080v2 Announce Type: replace-cross Abstract: Acquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation. While automated synthesis has emerged as an alternative to expensive manual curation, current approaches often rely on rigid heuristics, yielding data that is ungrounded or lacks logical complexity. We propose CodeEvo, a dual-agent architecture comprising a Coder for iterative solution synthesis and a Reviewer to orchestrate the generation trajectory. To transcend the limitations of existing heuristics, the Reviewer formulates a Schema to systematically architect logic and complexity through an interleaved synthesis of instructions and code. This process is further reinforced by a hybrid verification protocol synergizing deterministic compiler feedback with semantic evaluation. Under this framework, we construct CodeEvo-100K, a large-scale dataset of instruction-code pairs with stepped difficulty levels. Extensive e