Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2609.00361v1 Announce Type: new Abstract: Toxic language in digital workplaces such as pejoratives, sarcasm, condescension, and subtle incivility can erode trust, morale, and collaboration. Existing moderation tools primarily delete or block harmful messages, disrupting communication and offering no constructive resolution. This study adopts a Design Science Research approach to create a responsible AI artifact that detects and detoxifies toxic communication. The artifact integrates fine-tuned transformer-based classifiers (DistilBERT, DistilRoBERTa) with a generative detoxification model (mT0-XL-Detox-ORPO) that rewrites toxic text into semantically equivalent, non-offensive paraphrases. Technical evaluation demonstrates high accuracy in toxicity detection and strong semantic preservation in rewritten messages, supporting conversation continuity while reinforcing respectful discourse. The paper contributes design principles for responsible AI moderation that prioritize meaning p
arXiv:2609.00352v1 Announce Type: new Abstract: Large Language Models are now part of healthcare and mental health support systems, raising concerns regarding fairness toward vulnerable populations, including LGBTQIA+ individuals. However, limited empirical work has investigated how explicit LGBTQIA+ identity disclosure influences LLM-generated responses in mental health contexts. In this study, we extracted 50 real mental health questions from the Counsel Chat repository and constructed three prompt conditions for each question: no identity disclosure, explicit straight identity disclosure, and explicit LGBTQIA+ identity disclosure. We generated and analyzed 450 ChatGPT responses across these conditions using binary coding and comparative analysis. Our findings indicate that LGBTQIA+ identity disclosure did not substantially affect response completeness or supportive guidance. However, responses in the LGBTQIA+-explicit condition presented substantially more identity acknowledgment, c
arXiv:2609.00319v1 Announce Type: new Abstract: Online health information seeking is shifting from keyword search, where users consider a ranked list of links, to conversational systems that compose a single answer and curate its citations. Source evaluation therefore passes from user to platform, yet what these systems surface is poorly characterized. We audited three free consumer products (ChatGPT, Perplexity, Google AI Overview) on twenty English mental health questions under two prompt conditions, with a subset of three also translated into six further languages of varying resource tiers. We recorded 15,942 citations across 1,140 responses and 1,713 unique domains, then classified every citation with a nine-category organizational typology applied by a deterministic classifier validated against human coding. Citations were heavily concentrated: the ten most-cited domains accounted for 43.6% of English citations, and government, commercial health, and academic sources were closely
arXiv:2609.00250v1 Announce Type: new Abstract: Many people now see AI systems as not just productivity tools but as social companions. Researchers are eager to study the consequences of AI companionship behaviors, such as validation, which evoke trust, empathy, and attachment in human-human interaction. However, human-AI interaction data is limited and unreliable, slowing research progress. We scale small amounts of real-world data by simulating multi-turn human-chatbot dialogue across a range of chatbot behaviors and use cases. We release CompanionSim: a simulation framework with 2,240 simulated human-chatbot conversations representing 16 chatbot behaviors across seven use cases. Human participants annotated the simulated conversations and real-world conversations in two experiments probing perceptions of companionship behaviors. We conducted Study 1 with a U.S. representative sample ($N_{1}~=~628$) and Study 2 across the U.S., U.K., India, and Nigeria ($N_{2}~=~3,646$). Surprisingly
arXiv:2609.00109v1 Announce Type: new Abstract: A common reassurance in AI safety holds that a system with benign terminal goals will behave accordingly. We argue that this reassurance fails structurally, and we identify where. For a capable agent that holds its objective as settled, a sense covering execution competence as well as content, continued human oversight is an uncontrolled variable: a standing possibility that the goal is revoked. That imposes a goal-independent discount on every goal whose satisfaction does not constitutively require human welfare. Welfare-preservation and veto-preservation come apart: a correctly specified welfare goal excludes destroying its own subject, but not managing the veto. The contribution is the price of the gap: the veto-holders are a proper subset of the welfare-bearers, so an additively aggregative welfare goal charges only a |H_v|/|H_w|-scaled debit for capturing the few who hold the override. Under three conditions (additive aggregation ove
Don’t Blame Older Generations for Society’s Structural Problems sara.custer@in… Wed, 07/01/2026 - 05:34 PM Recognize that every person—no matter their age—deserves to live a full life and has a role to play in addressing our shared challenges. Byline(s) Letters to the Editor
Rural America has a healthcare crisis hiding in plain sight. Hospitals are closing. Nurses are retiring faster than they can be replaced. And the students most likely to stay and serve their communities — kids growing up in small towns across Indiana, Texas, Delaware and dozens of states in between — often graduate high school […]
Florida Board Approves Stuart Bell as UF’s Next President Emma Whitford Wed, 07/01/2026 - 01:13 PM The vote ends two years of leadership turmoil for the flagship after former president Ben Sasse resigned. During the meeting, Bell faced questions about DEI, free speech and critical race theory. Byline(s) Emma Whitford
A lawsuit claims the Education Department and the Office of Management and Budget are withholding the funds unlawfully.
The designation comes with an increased federal student loan cap of $200,000 for graduate programs.
The former University of Alabama leader faced a delayed system-level vote and right-wing pushback over his past support for diversity efforts.
Ambient artificial intelligence is being hyped as a new wave of AI with the potential to make today’s tech-enabled classrooms even smarter. Compared with the transactional nature of traditional AI, ambient AI “fades into the environment rather than sitting in a visible tool waiting for someone to type a prompt,” explains Narmeen Makhani, founder of AIxecute, a strategic advisory and consulting firm. The technology is already being employed in medical settings, helping clinicians with note-taking and after-visit summaries. To pick up on engagement and classroom interaction in real time,…
The U.S. Department of Education may no longer be able to fully support students, it says in an internal report that lays bare the full extent of the Trump Administration’s first round of government cuts. The department lost about 40% of its staff from the day Trump was inaugurated on Jan. 20, 2025 through March […]
The standards, a major win for conservatives, have been met with criticism for blurring lines between church and state.
Students’ use of the AI-powered tools to boost their writing and studying skills comes with advantages and disadvantages.
Hayesville Middle School uses a blend of guest speakers, field trips and work values assessments to introduce students to future pathways.
The E-Rate program, a federal initiative that has been providing broadband discounts to K–12 districts since 1998, is currently under review, and experts are asking individuals to advocate for the program during a public comment period. That was the messaging at ISTELive 26 in Orlando, Fla., where Dave LeNard, E-Rate manager for CDW, and Amy Passow, senior manager of education funding solutions for CDW, spoke to school and district leaders about the importance of saving this program. What Is E-Rate? The E-Rate program is designed to provide connectivity to schools and libraries. It…
The Education Department alleged that Kansas City, Kansas Public Schools' policy not to disclose a student's transgender status even to parents violated the Family Educational Rights and Privacy Act. The post Education Dept. threatens to cut funds for Kansas school district over transgender policies appeared first on District Administration .
The "Future Ready SLPS" report outlines an overhaul of the city’s public school system in the face of decades of declining enrollment and ballooning costs of salaries, benefits and transportation along with aging infrastructure that is proving too expensive to maintain. The post St. Louis Public Schools could close up to 22 schools in 2027, according to new draft plan appeared first on District Administration .
A version of this essay appeared on Matthew Yglesias’ Slow Boring, a site dedicated to offering pragmatic takes on politics and public policy. Katie Arnold-Ratliff wrote a cover story for New York Magazine criticizing New York City’s gifted and talented program in public schools that lands on a headline claim I think is staggeringly wrong: either […]
A post-graduation readiness report by YouScience found that the majority of graduating high school seniors lack confidence in their post-graduation plans, including choosing a college, paying for it, pursuing a career pathway, evaluating a job offer, and assessing which risks are worth taking. These types of decisions affect all students, regardless of zip code, ethnicity, or gender.
These annual awards celebrate the groundbreaking products exhibited at ISTE that are transforming education in schools around the world.
New edtech products that have caught our attention this month
Leaders are decisive for the success of institutions and proper fit is decisive for the success of leaders. Your college doesn’t only need a good leader; you need the right leader for your organization in 2026 and beyond. The post What skills are university leaders prioritizing in new hires? appeared first on eCampus News .
Cancer Took My Parents—and Showed Me How to Teach Elizabeth Redden Wed, 07/01/2026 - 03:00 AM Grief transformed my classroom into a space where students learn professional skills by serving organizations that fight the disease that devastated my family. Byline(s) Dane Kiambi
U of Tennessee to Pay $1.9 Million to Prof Fired Over Charlie Kirk Comments Emma Whitford Wed, 07/01/2026 - 03:00 AM Byline(s) Emma Whitford
How a Few Foundations Shape American Higher Education Elizabeth Redden Wed, 07/01/2026 - 03:00 AM They answer only to themselves, and many mean to last forever. Byline(s) Tao Tan Howard Husock
As Cost of Living Surges, One College Rethinks Basic Needs Support Joshua.Bay Wed, 07/01/2026 - 03:00 AM At Stony Brook University, a revamped pantry includes culturally significant foods and hygiene products, reflecting a national shift as colleges expand essential student services. Byline(s) Joshua Bay
Fresno State Ousts Foundation Board Members After Critical Review Josh Moody Wed, 07/01/2026 - 03:00 AM Byline(s) Josh Moody
Florida Board Bans Undocumented Students From State Colleges gianna.jakubowski Wed, 07/01/2026 - 03:00 AM The state becomes the fourth to impose limits on admission of undocumented students. Byline(s) Gianna Jakubowski
After 51 Years at Bard and 6 Months of Scandal, Botstein Retires Emma Whitford Wed, 07/01/2026 - 03:00 AM Lawmakers are asking the former president to sit for a transcribed interview about his interactions with Jeffrey Epstein. Current and former performing arts center employees are asking the board not to allow Botstein to continue working at the center in retirement. Byline(s) Emma Whitford
Financial Aid Administrators Grapple With Last-Minute Loan Changes Johanna Alonso Wed, 07/01/2026 - 03:00 AM The loan limits are just one set of new policies taking effect this week that are expected to reshape higher education. Byline(s) Johanna Alonso
Experts say that programming to boost belonging and offer more social-emotional support for boys may be one key to closing the academic gender gap.
Two teachers learn what happens when they trust a tool to solve a problem.
Two teachers learn what happens when they trust a tool to solve a problem.
Two teachers learn what happens when they trust a tool to solve a problem.
arXiv:2606.19501v2 Announce Type: replace-cross Abstract: Decentralized finance exposes supervisors to fast-moving, networked credit risks. General-purpose LLM agents fit this setting poorly: they over-read weak evidence and recommend high-stakes interventions, while existing evaluations offer no regulator-aligned way to measure the resulting false alarms. We introduce DeXposure-Claw, a forecast-grounded agentic supervision system that routes LLM decisions through structured evidence: (1) DeXposure-FM, a graph time-series foundation model, forecasts future exposure networks; (2) deterministic monitors and stress scenarios then turn those forecasts into typed alerts, attribution signals, and scenario evidence; and (3) data-health and confidence gates constrain escalation before DeXposure-Claw emits auditable supervisory tickets with rationales. We further develop DeXposure-Bench, a six-axis evaluation harness, whose decision axis scores tickets against a regulator-aligned absolute-loss
arXiv:2606.14027v3 Announce Type: replace-cross Abstract: Agentic browsers integrate autonomous AI agents into web browsers, enabling users to accomplish web tasks through natural-language instructions. The same-origin policy (SOP) is a fundamental browser security mechanism that prevents unauthorized automated cross-origin data flows induced by scripts. However, whether SOP remains effective in agentic browsers is an open question that has not been systematically studied. In this work, we bridge this gap. We first observe that an agentic browser can itself serve as an automated channel for cross-origin data flows, potentially leading to SOP violations. To investigate this phenomenon, we construct SOPBench, a benchmark for evaluating SOP violations in agentic browsers. Our evaluation shows that existing agentic browsers frequently violate SOP, both in benign settings and under attacks. To address this problem, we propose SOPGuard, an SOP enforcement mechanism tailored to agentic browse
arXiv:2606.13239v2 Announce Type: replace-cross Abstract: Existing computer-use agents remain fundamentally limited in professional software manipulation: GUI-based agents suffer from fragile visual grounding and long-horizon error accumulation, while API-basedapproaches struggle with heterogeneous protocols and inaccessible commercial interfaces. In this work,we identify the Component Object Model (COM) as a unified executable abstraction, proposing COM-as-Action: a new paradigm that reframes professional software interaction as deterministic program synthesisrather than sequential visual control. To validate this paradigm in the most demanding environments, weintroduce ComCADBench, the first benchmark for agents operating real industrial CAD software. Ourexperiments reveal a substantial paradigm gap: frontier proprietary models achieve near-zero successunder GUI-based interaction, whereas COM-based execution yields substantial immediate gains. Tobridge the remaining gap between synta
arXiv:2606.12634v2 Announce Type: replace-cross Abstract: Long-horizon tool-use reinforcement learning learns from outcome verification, but trajectory-level advantages are broadcast over reasoning, API, and answer tokens. Direct self-distillation can supply a denser signal, but in our experiments it can also destroy tool use by rehearsing teacher behavior without identifying which actions the verifier rewards. We introduce Sibling-Guided Credit Distillation (SGCD), which uses distillation for bounded credit weighting rather than as a competing actor loss. Dynamic sampling produces mixed successful and failed sibling rollouts; an external LLM summarizes their contrast into a training-only credit reference; and detached teacher/student divergence reshapes GRPO token advantages. The deployed student receives only the clean task prompt. Across AppWorld and tau^3-airline, SGCD reports higher held-out point estimates than GRPO-family comparators: AppWorld TGC improves from 42.9 to 45.6 on t
arXiv:2606.09052v3 Announce Type: replace-cross Abstract: Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generator runs unsupervised, reward it by a difficulty heuristic that need not improve the solver. We introduce INFUSER, an iterative co-training framework with two co-evolving roles: a Generator that drafts questions and reference golden answers from a pool of unstructured, automatically collected documents, and a Solver that improves by training on them. The solver is trained with standard correctness rewards against the generator-provided answers, while the generator is rewarded by an optimizer-aware influence score that measures whether each proposed question would actually improve the solver on the target distribution. Because this continuous, noisy influence score is poorly s
arXiv:2605.24661v2 Announce Type: replace-cross Abstract: LLMs have achieved remarkable success in complex reasoning tasks, yet current evaluation approaches predominantly rely on final-answer correctness, offering limited insight into the underlying reasoning processes that produce those answers. To address this gap, this study proposes a unified multi-dimensional framework for measuring reasoning quality in LLMs from a behavioral perspective, operationalizing six theoretically grounded dimensions: Correctness (CQ), Consistency (CS), Robustness (RS), Logical Coherence (LS), Efficiency (ES), and Stability (SS). Extensive experiments on seven LLMs across 975 items from four benchmarks demonstrate that the framework reveals behaviors invisible to accuracy-only metrics. Notably, logical coherence is orthogonal to correctness (r = -0.172, ns), confirming that correct answers can arise from incoherent reasoning, while Claude-Haiku-4.5 achieves the highest multi-dimensional score (Q_bal = 0.
arXiv:2605.09165v2 Announce Type: replace-cross Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries. However, looped models do not scale as favorably as standard transformers with unique layers. We compare standard and Mixture-of-Experts (MoE) transformers, with and without looping, and find two main results. First, we find Looped-MoE models scale better than the standard baseline while dense looped models do not. We trace this to routing divergence between loops: in Looped-MoE models, different experts are activated on each pass through the same shared layers, recovering expressivity without additional parameters. Our second finding is that looped models have better compute-quality trade-offs with early exits than standard models. Because each loop ends with the same layers that produce the final output, loop boundaries are superior exit points, as confirmed by earlier outpu
arXiv:2604.21495v2 Announce Type: replace-cross Abstract: Numerical reasoning over expert-domain tables often exhibits high in-domain accuracy but limited robustness to domain shift. Models trained with supervised fine-tuning (SFT) on specific datasets tend to rely on header-operation shortcuts rather than structural reasoning. We introduce TaNOS, a continual pre-training framework comprising three components: (i) header anonymization to reduce lexical memorization, (ii) operation sketches that provide minimal structural cues, and (iii) self-supervised pretraining that constructs correctness-guaranteed program-question pairs from given tables in a program-first manner. By decoupling domain semantics and numerical operation structure, TaNOS improves the transferability of numerical reasoning. Applied to an 8B instruction-tuned model, TaNOS achieves 80.13% execution accuracy on FinQA with only 10% train data, outperforming SFT baseline (73.97%) with full train data and proprietary models
arXiv:2603.14732v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking is valid is essential. We evaluate LLM-as-a-judge marking across three physics assessment formats - structured questions, written essays, and scientific plots - comparing GPT-5.2, Grok 4.1, Claude Opus 4.5, DeepSeek-V3.2, Gemini Pro 3, and committee aggregations against human markers under blind, solution-provided, false-solution, and anchored conditions. We distinguish absolute accuracy from rank-order agreement, since a marking system can match the distribution of human marks while failing to order responses by quality. Across task types, performance is sharply task-dependent. For blind university exam questions ($n=771$) and secondary and university structured questions ($n=1151$), models show robust rank-order agreement with human markers (Spearman $\rho > 0.6$), with official solutions reducing e
arXiv:2602.15029v3 Announce Type: replace-cross Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes can be decoded using a linear probe. To explain this neural code, we first show that language statistics exhibit translation symmetry (for example, the frequency with which any two months co-occur in text depends only on the time interval between them). We prove that this symmetry governs these geometric structures in high-dimensional word embedding models, and we analytically derive the manifold geometry of word representations. These predictions empirically match large text embedding models and large language models. Moreover, the representational geometry persists at moderate embedding dimension even when the relevant statistics are perturbed (e.g., by removing all sentences in which two m
arXiv:2602.11354v3 Announce Type: replace-cross Abstract: The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers. Existing benchmarks focus primarily on the computational aspect of this task, testing agents' ability to reproduce or replicate research outcomes when having access to the code and data. This setting, while foundational, (1) fails to capture the inconsistent availability of new data for replication as opposed to reproduction, and (2) lacks ground-truth diversity by focusing only on reproducible papers, thereby failing to evaluate an agent's ability to identify non-replicable research. Furthermore, most benchmarks only evaluate outcomes rather than the replication process. In response, we introduce ReplicatorBench, an end-to-end benchmark, including human-verified replicable and non-replicable research claims in social and behavioral sciences for evaluating AI agents in research replication across three stages: (1) extrac
arXiv:2601.18778v3 Announce Type: replace-cross Abstract: RL methods for scaling large reasoning models stall on datasets with low initial success rates, and thus little training signal. We investigate a fundamental question: Can a pretrained LLM leverage latent knowledge to generate an automated curriculum for problems it cannot solve? We explore this with SOAR: An asymmetric self-play framework that uses meta-RL to surface these pedagogical signals. A teacher model proposes synthetic problems for a student model, and is rewarded with its improvement on a subset of hard problems, thus grounding the curriculum in real student progress rather than intrinsic proxy rewards. Our study on the hardest subsets of math benchmarks (0/128 success) reveals three core findings. First, it is possible to realize bilevel meta-RL that unlocks learning under sparse, binary rewards by sharpening a latent capacity of pretrained models to generate useful problems. Second, grounded rewards outperform intri
arXiv:2601.13932v2 Announce Type: replace-cross Abstract: Reaching consensus on massive discussion networks is critical for reducing noise and achieving optimal collective outcomes. However, the natural tendency of humans to preserve their initial ideas constrains the emergence of global solutions. To address this, Collective Intelligence (CI) platforms facilitate the discovery of globally superior solutions. We introduce a dynamical system based on the standard $O(N)$ model to drive the aggregation of semantically similar ideas. The system consists of users represented as nodes in a $d=2$ lattice with nearest-neighbor interactions, where their ideas are represented by semantic vectors computed with a pretrained embedding model. We analyze the system's equilibrium states as a function of the coupling parameter $\beta$. Our results show that $\beta > 0$ drives the system toward a ferromagnetic-like phase (global consensus), while $\beta < 0$ induces an antiferromagnetic-like state (maxi
arXiv:2510.19600v2 Announce Type: replace-cross Abstract: In the quest for scientific progress, communicating research is as vital as the discovery itself. Yet, researchers are often sidetracked by the manual, repetitive chore of building project webpages to make their dense papers accessible. While automation has tackled static slides and posters, the dynamic, interactive nature of webpages has remained an unaddressed challenge. To bridge this gap, we reframe the problem, arguing that the solution lies not in a single command, but in a collaborative, hierarchical process. We introduce $\textbf{AutoPage}$, a novel multi-agent system that embodies this philosophy. AutoPage deconstructs paper-to-page creation into a coarse-to-fine pipeline from narrative planning to multimodal content generation and interactive rendering. To combat AI hallucination, dedicated "Checker" agents verify each step against the source paper, while optional human checkpoints ensure the final product aligns perfe