Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2607.15134v1 Announce Type: new Abstract: We study how a representative sample of United States adult AI-assistant users (n=1,999; June 2026) choose among platforms, allocate tasks across them, evaluate provider trustworthiness, and value data-handling features. Estimates are weighted to the AI-user population using external adoption benchmarks. Four patterns emerge. The market is concentrated but internally differentiated: ChatGPT is the primary assistant for 58% of users and Gemini for 25%, yet smaller platforms hold defensible task niches--Claude captures a third of coding tasks despite a 7% overall share. Task allocation is thus organized by platform far more than by user, and technical use falls steeply with age. Trust is earned through use rather than reputation: Claude is ranked most trustworthy in every head-to-head among users of both platforms, and shows by far the largest gap between how its users and non-users rate it. Finally, privacy concern is near-universal but ac
arXiv:2607.15051v1 Announce Type: new Abstract: Canadian organizations deploying artificial intelligence systems face a fragmented regulatory landscape spanning federal requirements (the Treasury Board Directive on Automated Decision-Making) and divergent provincial regulations across Ontario, Quebec, Alberta, Manitoba, and British Columbia. The death of Bill C-27 (Artificial Intelligence and Data Act) in January 2025 - and the federal government's June 2026 confirmation that it will pursue targeted instruments rather than omnibus AI legislation - leaves organizations without unified compliance guidance. Global frameworks such as NIST AI RMF 1.0, the EU AI Act, and ISO/IEC 42001 provide valuable guidance but lack systematic methodologies for adaptation to multi-jurisdictional national contexts. We present SCITUS (Systematic Canadian Integration for Trustworthy and Unified Standards), a comprehensive framework adapting NIST AI RMF 1.0 to Canadian federal and provincial AI regulations si
arXiv:2607.14575v1 Announce Type: new Abstract: Generative AI chatbots promise to transform English as a Foreign Language (EFL) writing by providing immediate, personalised feedback. However, their pedagogical value depends on how learners engage with them - a process often treated as a "black box." This study uses Transition Network Analysis to model the temporal dynamics of Japanese EFL learners using "Penny," an LLM-powered writing chatbot. Analysis of over 4,500 writing sessions and 21,000 chatbot interactions reveals two dominant behavioural loops: a "Revision Loop," where feedback leads directly to successful error correction, and a "Chat Loop," where learners engage in sustained dialogue with the chatbot following feedback. Crucially, EFL proficiency significantly shapes interaction: high-proficiency learners engage more in open dialogue and negotiation with the chatbot, while low-proficiency learners rely more heavily on repetitive corrective feedback cycles. The findings demon
arXiv:2607.14532v1 Announce Type: new Abstract: Recent work shows that it is possible to extract verbatim or near-verbatim text of some copyrighted works from some large language models (LLMs or models). That is evidence that the model weights encode the works in some form - that the model has "memorized" those works from its training data. But LLMs don't store information in the same format as familiar databases. Rather, their weights store statistical relationships between tokens that have been learned from the training data, and those relationships inform a generation process that is often probabilistic rather than deterministic. In the case of memorization, those relationships are strong enough that, in many circumstances, the model might generate a copyrighted work from its training data with some probability. Copyright law has not previously had to decide whether storing information that might or might not produce output similar to a copyrighted work is itself a copy of the work.
arXiv:2607.14479v1 Announce Type: new Abstract: As large language models become increasingly capable, concerns about their potential to assist with biological misuse continue to grow. Prioritization of safety differs across the model ecosystem, with some models freely providing high-risk information that could be misused, and others refusing benign scientific content, potentially hindering legitimate research. Both failures stem from a lack of targeted mitigation to distinguish the most dangerous information from broader scientific content. To address this, we introduce BioTIER (Biological Targeted Information for Exclusion and Refusal), a benchmark designed to enable more targeted biological risk mitigation. BioTIER organizes biological content into three risk sets: Catastrophe Avoidance (CA), Biomedical DURC (BD) and Related Biology (RB). These sets represent a spectrum from extremely narrow high-risk topics to a broad range of benign and beneficial biological knowledge. The benchmar
arXiv:2607.14447v1 Announce Type: new Abstract: AI assistants increasingly retrieve web content at inference time to provide fresh and grounded answers, yet it remains unclear whether these search-augmented capabilities respect website-owner restrictions expressed through robots.txt. We present a controlled empirical study of ten widely used AI assistants with advertised web-search capabilities. For each assistant, we first identify a configuration that actually produces observable web-browsing behavior and record the user-agent exposed during retrieval. We then evaluate compliance with controlled robots.txt rules across four complementary conditions: allowed for all user-agents, disallowed for all user-agents, allowed only for the assistant-specific user-agent, and disallowed only for that user-agent. Using server-side logs and secret codes embedded in target pages, we distinguish actual page access from user-visible answer correctness across 200 trials. Our results show substantial v
arXiv:2607.14353v1 Announce Type: new Abstract: As automated decision-making and data-driven technologies pervade society and are used to manage consequential outcomes, understanding the technology's capabilities, limitations, and attendant risks in context requires analysis of full sociotechnical systems. Sociotechnical analysis of risks in highly complex systems provides clear lessons for the design and evaluation of AI systems, transcending a technical focus on reliable or "responsibly designed" components to understand risks at a systems level. Human-made catastrophes have been studied for decades because of the severity of these events: consider Chernobyl, Three Mile Island, Fukushima-Daiichi, Bhopal, the Challenger disaster. A common misconception is that these kinds of events are freak accidents, resulting from the inherently unforeseeable interactions in complex systems. Closer examination reveals that the risks and hazards were well-known beforehand but not acted upon due to s
Recent findings on the negative impacts of AI on learning might be sparking national debate, but they are unsurprising to learning scientists.
When you walk into a math classroom in Charleston County School District, you can feel the difference. Students aren’t just memorizing steps--they’re reasoning through problems, explaining their thinking, and debating solutions with their peers.
New York is currently standing at a historic crossroads. With a rare alignment of executive leadership in Albany and NYC and a tireless advocacy community, the state is poised to transform the promise of universal early childhood education (ECE) into a reality for tens of thousands of families.
AI may make it easier to manipulate athletic performance, but students often underestimate how easily it can be exposed
Article URL: https://www.mdpi.com/2078-2489/16/8/653 Comments URL: https://news.ycombinator.com/item?id=44915580 Points: 2 # Comments: 0
AI is now at the center of almost every conversation in education technology. It is reshaping how we create content, build assessments, and support learners. The opportunities are enormous.
PatientRightsAdvocate.org sued the American Medical Association this week, arguing the trade group has no valid copyright over the CPT codes that providers must use to bill for care. It follows a public letter Senator Bill Cassidy sent to the AMA last year accusing it of “abus[ing] [its] government-backed monopoly by charging exorbitant fees to anyone using the CPT code set.” The post Patient Group Sues AMA, Says the Org Shouldn’t Be Charging for CPT Codes appeared first on MedCity News .
Former Cambridge Professor Jason Arday Found Dead Susan H. Greenberg Fri, 08/14/2026 - 04:25 PM Police say the 41-year-old was found unresponsive at an address in south London and pronounced dead at the scene. Byline(s) Tom Williams for Times Higher Education
Venrock led Aligned Marketplace’s Series A round. In total, the company has raised $31 million to date. The post Aligned Marketplace Secures $20M for Advanced Primary Care Marketplace appeared first on MedCity News .
Bristol Myers Squibb’s Zenbexus is the first FDA-approved drug in a new class of cancer drugs called CELMoDs. The pharma company is positioning this molecule as a successor to its legacy products for multiple myeloma. The post Bristol Myers Squibb Protein Degrader Wins First-in-Class FDA Nod in Multiple Myeloma appeared first on MedCity News .
COLUMBIA, S.C. — Elementary and middle school students posted their best scores yet on end-of-year standardized tests, though more than half still can’t calculate as expected for their age, according to state testing data released Monday. Third- through eighth graders continued to improve in both math and reading, though scores were still generally stronger in […]
Under the Trump administration, the agency has vocally cracked down on diversity, equity and inclusion work under the auspices of Title VII.
The federal government last week proposed sweeping changes to Head Start, describing them as a way to expand access for children. For families of children with disabilities, the question is what support will still be there once a child enters the classroom. I’ve been around long enough to know that when it comes to meeting […]
A task force called for making the university's general education more narrow and focused on topics such as Western civilization and broad U.S. history.
Three years into Texas’ takeover of the district, one of its most troubled schools has risen from seven consecutive F ratings to its first-ever A.
Backpacks can be fashionable and functional, but they can also be too heavy — weighed down by digital devices, musical instruments, sports equipment and more. Some kids carry home a laptop or tablet and textbooks, too. It’s good to be prepared, but kids who walk to school or participate in extracurricular activities may be lugging more […]
The scores grade how well specific campuses and school districts as a whole perform academically and prepare students for college and to enter the workforce. The post 15% of Texas Schools scored below a ‘C’ on the state accountability scores appeared first on District Administration .
Most younger professionals entering the workforce have little visibility into interoperability, digital health infrastructure, or healthcare data architecture as career pathways. The post Healthcare Keeps Buying AI. But Nobody’s Building the Workforce to Run It. appeared first on MedCity News .
It’s move-in week for college students across North Carolina. And once students have gotten settled into their dorms, their attention will turn to getting the classes needed to fulfill their chosen major. But which career paths are a safe bet in 2026? U.S. employers cut 23,000 jobs in July. Job creation for May and June […]
For all the noise and attention and — laudable — political energy behind it, the science of reading is not a thing. Or, rather, it is not a thing that schools can implement. The science of reading is a series of facts that, in combination, form a field consensus about how children best learn to […]
For the first time, a federal tax credit is opening a clear path for public school students to access expanded learning opportunities through Scholarship Granting Organizations (SGOs).
The outcome-based contract model can be a good way to increase efficiency and get more bang for the edtech budget buck.
Graduate education is demanding and isolating, especially for adult learners balancing coursework, research, employment, and family responsibilities. National data show half of doctoral students leave their programs before earning their degree. The post Cohort connections matter: Strategies to help graduate students persist and succeed appeared first on eCampus News .
U of Michigan Attempts to Disrupt the Transactional System johnw@mcsweeneys.net Fri, 08/14/2026 - 03:00 AM Making space for student self-regulation is good, actually. Byline(s) John Warner
The Small Section Cut Session Sara Brady Fri, 08/14/2026 - 03:00 AM Enrollment shifts aren’t evenly distributed. Byline(s) Matt Reed
Michigan Narrows Gap in Who Can Get Free Community College Tuition Ryan Quinn Fri, 08/14/2026 - 03:00 AM Byline(s) Ryan Quinn
Accreditor Puts Lane Community College on Warning kathryn.palmer… Fri, 08/14/2026 - 03:00 AM Byline(s) Kathryn Palmer
Stockton Pulls Genocide Reference After Donor Complains Josh Moody Fri, 08/14/2026 - 03:00 AM The public university in New Jersey abruptly deleted a reference to genocide in Gaza from its website. The author, a Jewish scholar of genocide, believes officials caved to donor pressure. Byline(s) Josh Moody
Cambridge to Review Hiring Processes After Arday Scandal sara.custer@in… Fri, 08/14/2026 - 03:00 AM The university says its investigation into sociology professor Jason Arday will feed into a wider inquiry on recruitment of senior academics. Byline(s) Tom Williams for Times Higher Education
Troy Punishes Professor for Writing to Lawmakers on University Letterhead Emma Whitford Fri, 08/14/2026 - 03:00 AM Byline(s) Emma Whitford
How 4 Colleges Are Supporting Student Parents Joshua.Bay Fri, 08/14/2026 - 03:00 AM From scholarships and free childcare to pre-orientation programs and workforce training, institutions are helping student parents and caregivers persist. Byline(s) Joshua Bay
DOJ Declares 3 Race-Based STEM Programs Unconstitutional Olivia.sanchez Fri, 08/14/2026 - 03:00 AM The programs received about $104 million in federal funds this past year. Byline(s) Olivia Sanchez
Institutions have until Jan. 15 to submit two rounds of data required by the Biden-era gainful employment and financial value transparency rules.
The eliminations come after a bruising budget-saving push to eliminate degrees last year at the system’s Lincoln campus.
The agency’s free guides for district leaders come as the education sector continues to face a perfect storm of cyber vulnerability and limited resources.
From the latest data on written state special education complaints to a lawsuit against the Education Department, what did you learn from our recent stories?
arXiv:2607.27539v2 Announce Type: replace-cross Abstract: Exact deletion from persistent language-model memory depends on whether a record's effect remains addressable after later computation. Native Kimi Delta Attention (KDA) gives a negative result for the tested receipt interface: the corpus-pooled raw recurrent contribution changes by 12-49% with the suffix and remains 8-49% after a decay-ledger correction. Native omission also changes later transition and write terms and other active caches. Frozen-input transport succeeds on its fixed-input control; the changed terms place native omission outside the tested receipt classes. Checkpoint replay supplies the evaluated recomputation path; zero residual on final logits and all 80 audited KDA arrays verifies restoration across the declared checkpoint surface. The complementary result is constructive. We retrofit support-vector memory into frozen Gemma 3 without attention transfer, low-rank recovery, distillation, adapters, or language-m
arXiv:2606.04032v3 Announce Type: replace-cross Abstract: Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and the impact of omitting some remain poorly understood. We systematically evaluate three projection sharing constraints: a) Q-K=V (shared key-value), b) Q=K-V (shared query-key), and c) Q=K=V (single projection). The last two variants produce symmetric attention maps; to address this, we also explore asymmetric attention via 2D positional encodings. Through experiments spanning synthetic tasks, vision (MNIST, CIFAR, TinyImageNet, anomaly), and language modeling (300M and 1.2B parameter models on 10B tokens), we discovered that our transformers perform on par or occasionally better than the QKV transformer. In language modeling, Q-K=V projection sharing achieves 50% KV cache reduction with only 3.1% perplexity degra
arXiv:2605.18852v2 Announce Type: replace-cross Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy. Small observed differences can be comparable to variability introduced by finite evaluation samples, LLM-based judges, and ambiguous multimodal evidence, while validation loss may not identify the checkpoint preferred by downstream evaluation. We formulate late-stage checkpoint selection as a stability-aware decision problem under evaluation uncertainty and propose a progressive framework combining pointwise filtering, listwise ranking, and pairwise refinement. Repeated evaluation-set subsampling is used to characterize ranking stability, while percentile-based aggregation accounts for lower- and upper-tail behavior. Experiments show that multimodal data evaluability is critical: quality-aware curation of OCR-heavy inputs reduces ranking flip rate fro
arXiv:2604.27906v3 Announce Type: replace-cross Abstract: Persistent AI memory is often reduced to a retrieval problem: store prior interactions as text, embed them, and ask the model to recover relevant context later. This design is useful for thematic recall, but it is mismatched to the kinds of memory that agents need in production: exact facts, current state, updates and deletions, aggregation, relations, negative queries, and explicit unknowns. These operations require memory to behave less like search and more like a system of record. This paper argues that reliable external AI memory must be schema-grounded. Schemas define what must be remembered, what may be ignored, and which values must never be inferred. We present an iterative, schema-aware write path that decomposes memory ingestion into object detection, field detection, and field-value extraction, with validation gates, local retries, and stateful prompt control. The result shifts interpretation from the read path to the
arXiv:2603.14501v2 Announce Type: replace-cross Abstract: Large Language Models excel in high-resource programming languages but struggle with low-resource ones. Existing research related to low-resource programming languages primarily focuses on Domain-Specific Languages (DSLs), leaving general-purpose languages that suffer from data scarcity underexplored. To address this gap, we introduce CangjieBench, a contamination-free benchmark for Cangjie, a representative low-resource general-purpose language. The benchmark comprises 248 high-quality samples manually translated from HumanEval and ClassEval, covering both Text-to-Code and Code-to-Code tasks. We conduct a systematic evaluation of diverse LLMs under four settings: Direct Generation, Syntax-Constrained Generation, Retrieval-Augmented Generation (RAG), and Agent. Experiments reveal that Direct Generation performs poorly, whereas Syntax-Constrained Generation offers the best trade-off between accuracy and computational cost. Agent
arXiv:2601.10560v2 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) coordinate multiple LLM-powered agents through structured workflows, gaining reasoning power but incurring high inference latency from multi-step execution and repeated model invocations. Existing orchestration methods primarily optimize task performance and inference cost, leaving latency largely unaddressed. In MAS, end-to-end latency is governed by the \textit{critical execution path}, so reducing total cost alone does not reliably reduce latency. Moreover, optimizing latency while preserving accuracy remains non-trivial: naive latency optimization can misassign operator-level credit and degrade task accuracy. To address this gap, we propose \textbf{L}atency-\textbf{A}ware \textbf{M}ulti-\textbf{a}gent \textbf{S}ystem (\textbf{LAMaS}), a latency-aware orchestration framework for learning-based multi-agent systems. LAMaS addresses this challenge at two levels: at \emph{training time}, it learns latenc
arXiv:2510.22282v2 Announce Type: replace-cross Abstract: Urban socio-economic sensing plays a vital role in advancing global sustainable development goals. With the advent of Large Vision-Language Models (LVLMs), new opportunities have emerged to address this challenge by framing it as a multi-modal perception and reasoning task. However, recent studies show that LVLMs still struggle to make accurate and interpretable socio-economic predictions from visual data. To overcome these limitations and fully exploit the potential of LVLMs, we propose CityRiSE, a novel framework for Reasoning urban Socio-Economic status in LVLMs via reinforcement learning (RL). With carefully curated multi-modal dataset and verifiable reward design, our approach guides the LVLM to focus on semantically meaningful visual cues, enabling structured and goal-oriented reasoning for generalist socio-economic status prediction. Experiments demonstrate that CityRiSE, equipped with emergent reasoning, significantly ou