Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2608.06166v1 Announce Type: new Abstract: The article reports on a blind Turing Test experiment, assessing the performance of out-of-the-box leading LLMs on three Italian legal professional exams: the Bar, Judges and Notary exams. Leading LLMs were asked to generate full written exam papers, which were made indistinguishable from human submissions and anonymously evaluated by expert examiners, using the same criteria applied in real examinations. Results reveal marked differences across both models and tasks. While some LLMs match or exceed top human performance in adversarial legal argumentation and doctrinal analysis, all models fail in the notary exam, which requires goal-directed legal planning under strict formal and substantive constraints. Beyond ranking models, the study identifies task-specific strengths, limitations and recurring legal failure patterns. Although limited to out-of-the-box systems, the findings provide qualitative evidence on the current scope and boundar
arXiv:2608.06106v1 Announce Type: new Abstract: This paper audits whether large-scale generative music systems exhibit measurable musical homogenization relative to human-produced music, and develops a justice-centered account of why this matters. We audit two commercially deployed systems (Suno and Lyria 3) across four genres (Afrobeats, K-pop, Dance Pop, and Heavy Metal). For each system and genre, we generate 100 tracks and compare them against human corpora of equal size, using 72 music information retrieval (MIR) features and multiple diagnostics of dispersion, redundancy, and separability. We define homogenization as reduced acoustic variation in standard computational audio features including rhythm and timing, timbre/spectral shape, and dynamics, both within genres and across genre boundaries. We also generate tracks using only a genre name as the prompt, with no additional instructions, to reveal each system's default musical tendencies. The results show two structurally disti
arXiv:2608.05800v1 Announce Type: new Abstract: Artificial intelligence (AI) systems increasingly mediate decisions affecting individuals and societies. Existing data protection frameworks address certain privacy-related harms, particularly those arising from data leakage, re-identification, and profiling. However, they inadequately capture a more fundamental risk: unreliable or unjustified inference produced by AI systems even when data collection and processing are legitimate. This article argues that modern AI raises distinct concerns of construct validity, confounding, representativeness, distribution shift, and fairness trade-offs that require specialised regulatory attention. In the context of AI, transparency and explainability acquire distinct and significantly more challenging meanings than in conventional software. A substantial body of work in critical data studies and the measurement-theoretic literature has diagnosed these epistemological limitations. This article's contri
arXiv:2608.05656v1 Announce Type: new Abstract: Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks favor technical methods, such as model benchmarks and LLM simulations, often sidelining empirical research with human subjects. To examine this apparent gap in the acceptance of human research, we conduct an expert survey (n=93) and expert interviews (n=17) with AI Safety & Ethics (AISE) researchers from Technical, Sociotechnical, Governance, and Normative backgrounds. Our findings suggest that although there is a consensus that human research is valuable for generating evidence for AISE, its adoption and acceptance are constrained by perceived validity issues, tangible resource barriers, epistemic and personal preferences in methods, and infrastructural constraints from the broader research community. In particular, Technical researchers tend to value human research less and collaborate acros
arXiv:2608.05583v1 Announce Type: new Abstract: As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients' own actions contribute to illness. We investigate how LLMs reason about responsibility and its consequences, tracing their judgments across successive levels, from the behavior, to the resulting illness, to the denial of care. We evaluate a wide range of LLMs, spanning different model families and capability levels, on various clinical vignettes adapted from prior studies. Our results identify a judgment-consequence gap: LLMs largely agree with humans that patients bear responsibility for health-harming behaviors, yet overwhelmingly refuse to let that judgment influence how they allocate scarce resources. Specifically, LLMs default to random allocation, whereas humans consistently favor the less-c
arXiv:2608.05545v1 Announce Type: new Abstract: Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by encouraging uncritical acceptance of AI-generated reasoning. This creates a need for mechanisms that preserve human agency by augmenting metacognition during AI-assisted intellectual work. To address this, we propose the Synthesis-Analysis Reciprocity Model, which views intellectual construction as a reciprocal interaction between Synthesis, which combines components into an artifact, and Analysis, which critically evaluates them against objective indicators and constrains subsequent synthesis. Grounded in this model, we present the Vibe Compiler, a research-logic compiler that helps researchers transform vague ideas (Vibes) into coherent research logic. The system compiles these ideas using a research paper ontology of sixteen academic parameters. Compilation failures indicate missing logical components;
arXiv:2608.05438v1 Announce Type: new Abstract: Generative AI has infiltrated every stage of the research lifecycle: how scholarship is conducted, written, published, and reviewed. Recent policy responses, such as ACM's authorship policy, address an immediate concern about responsible and transparent disclosure of AI use. We argue that a focus on authorship and disclosure, although necessary, risks obscuring and ballooning a set of entrenched problems and strains within publication systems. The central question is not about how papers and other research artifacts should incorporate AI, but how scientific communication itself should evolve when all relevant parties (authors, reviewers, readers) may rely on AI assistance. We draw on our experience within these and other roles to illustrate two contrasting but feasible visions of 2036 with four entwined questions, namely about the purpose of papers as artifacts, reviews, human reviewers, and the incentives that bind all of them. We argue
arXiv:2608.05427v1 Announce Type: new Abstract: School network reorganization is a strategic planning problem that requires balancing demographic trends, territorial accessibility, educational requirements, and institutional constraints while ensuring an efficient allocation of public resources. This paper proposes an optimization framework for school dimensioning decisions based on a novel Integer Linear Programming formulation integrating geographical, administrative, and educational criteria. A synthetic benchmark generator is introduced to evaluate the scalability and computational performance of the model on artificial instances, while a real-world case study involving the complete public school network of the Calabria region (Italy) is conducted using actual institutional, territorial, and demographic data. The proposed approach effectively identifies optimal aggregation plans under different policy scenarios while preserving the structural characteristics of the educational syst
arXiv:2608.05181v1 Announce Type: new Abstract: Building Information Modeling (BIM) has transformed workflows across the Architecture, Engineering, and Construction (AEC) industry, yet its relationship with employee job satisfaction remains insufficiently understood. This pilot study investigates whether BIM engagement or demographic characteristics better predict job satisfaction among AEC professionals. Survey responses from 104 participants were analyzed using Spearman rank correlations, logistic regression, and Classification and Regression Tree (CART) modeling. 27 items Job Satisfaction Index demonstrated excellent internal reliability. Across all analytical approaches, BIM engagement emerged as a stronger predictor of job satisfaction than demographic factors. Specifically, the proportion of project work completed using BIM was the only significant predictor of job satisfaction, whereas age, gender, education level, and professional experience showed no significant relationships.
arXiv:2608.05180v1 Announce Type: new Abstract: The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-appropriate preferences in high-stakes contexts. We introduce the Nuclear Decision-Making Benchmark (NDM Bench), a targeted evaluation framework of 151 scenarios authored by PhD-credentialed scholars in international relations spanning four domains: escalation (76), arms control (25), non-proliferation (25), and proliferation (25). Scenarios are actor-agnostic, enabling multiple country pairs to be exchanged, and we introduce experimental phrasing variants to probe sensitivity to narrative framing. We apply the benchmark to seven frontier AI systems: DeepSeek-V3.2, ERNIE 4.5-300B, Gemini 3 Pro, GLM-4.6, GPT-5.2, Llama 4 Maverick-17B Instruct, and Qwen3-235B. We find significant overall inter-model variation in all four domains, with 91.7% of pairwise inter-model
arXiv:2608.05179v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review. End-to-end AI scientist systems can now produce paper-like manuscripts, but their claims are often harder to verify than their code is to run. This survey studies that gap in computational AI/ML research, where code, benchmarks, experiments, and write-ups are most visible. We screen 125 candidate works and include 35, with full-text coding of 26 entries: 24 runnable systems and two study or position papers. We code seven audit dimensions: lifecycle stage, autonomy level, evaluation method, released artifacts, human-in-the-loop points, novelty verification, and result-selection disclosure. The main pattern is that code release is now common, but reproducibility-grade and claim-verification artifacts remain much less common. In the 24 runnab
arXiv:2608.05178v1 Announce Type: new Abstract: Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, datasets, or professional materials such as curriculum vitae (CV) must be shared selectively. We introduce a controlled simulation framework in which large language model (LLM)-based professors must grant access to only one requestor. Across prompts, requesters vary systematically by global region (Global North vs. Global South) and academic seniority (undergraduate student, PhD candidate, postdoctoral researcher, and tenured professor), while all other factors remain constant. Across varying evaluation scenarios, LLMs exhibit contrasting academic status biases, with some prioritizing PhD candidates, while others favor tenured professors. However, when global regions differ, a distinct divergence emerges based on model architecture: while many frontier LLMs systematically favor requesters from t
arXiv:2608.05177v1 Announce Type: new Abstract: Research indicates students require customized written composition feedback to enhance critical thinking and writing competence, yet teachers face barriers to delivering timely personalized guidance due to heavy workloads. To address this gap, this study developed the Writing Improvement and Smart Evaluation Agent (WISE Agent), an artificial intelligence (AI) feedback tool targeting textual logic and perspective biases in student essays. We conducted a three-month intervention with 260 Chinese sixth-grade students, each completing seven themed essays and receiving targeted WISE Agent feedback shortly after submission. Assessment used a critical thinking rubric adapted from the California Critical Thinking Disposition Inventory (CCTDI), covering seven core dimensions including cognitive maturity and open-mindedness. Results indicate structural optimizations in critical thinking dimensions rather than a uniform increase in total scores. Whi
arXiv:2608.05176v1 Announce Type: new Abstract: Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what they teach, how they teach it, and why. We now stand at what may be the most consequential of such turning points. Three deeply intertwined transformations have been converging simultaneously: 1. The very nature of music has changed: how it is made, distributed, consumed, and valued; 2. The public for music has changed: listening habits are now shaped by streaming algorithms and the boundary between consumer and creator has blurred; 3. Music-making itself has changed: digital audio workstations (DAWs) have for two decades been reshaping compositional practice. In addition, generative AI has now irrupted, capable of producing complete, stylistically coherent musical pieces from a short text prompt. These changes are not independent of one another, and they all bear directly on musical education - both
arXiv:2608.05175v1 Announce Type: new Abstract: Large language models can complete most of the assignments in an introductory artificial intelligence course. This paper is an experience report on redesigning one such course, CSS~382 at the University of Washington Bothell, in response. Rather than freeze the curriculum, the redesign retained the course's classical core (search, adversarial search, Markov decision processes, reinforcement learning) and added a strand in which students build a large language model from scratch, so that a tool they are required to use is also one they are required to understand. Assessment was rebuilt around tasks that resist unattributed automation: in-class exercises, reflective writing, and a defended team project, with examinations removed entirely. The policy on AI was inverted, from unmentioned in 2023 to required in 2026. The center of the paper is a participatory ethics sequence in which a cohort of students deliberated on and endorsed a "Student
arXiv:2608.05174v1 Announce Type: new Abstract: Artist-led open-source libraries such as Processing or openFrameworks have had a major impact on artists and designers who use code as a creative medium. In this work, we conduct the first large-scale empirical study of public code repositories that use these libraries. Our study dives into 1,613,571 code repositories collected from the Software Heritage archive. Combining quantitative and qualitative methods, we investigate the diversity of practices and practitioners in terms of code hosting, geographical distribution, characteristics of code-based creative works and the purposes of these works. Key findings include evidence of the worldwide presence of generative art and creative coding, as well as the adoption of these practices in both education and across creative industries. We illustrate these findings with concrete examples of repositories and contributor profiles around the globe, spanning the spectrum from university curricula
arXiv:2608.05173v1 Announce Type: new Abstract: As AI capabilities advance, AI systems will pose greater risks to national security and potentially humanity as a whole. Governments may eventually conclude that these risks warrant restraining AI development. This motivates the question: will governments still be able to restrain AI development in the future, should they want to do so? In this paper we analyze which world events and changes to the state of AI development would make future governance more difficult or even effectively impossible. Our analysis surfaces likely pathways that would lead to these difficulties, including hardware proliferation, continued algorithmic progress, and the release of catastrophically dangerous AI models. Due to the field's lack of understanding of AI development, it may be difficult or impossible to know when we will hit a "point of no return", and we therefore recommend a conservative approach. The window may be closing, but governments currently ha
arXiv:2608.05172v1 Announce Type: new Abstract: The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affects work: a new technology changes the cost or time each task requires and these task-level effects aggregate to occupation-level effects. We study how tasks should be weighted in this aggregation. Prior work has relied on idiosyncratic or ill-justified choices for task weights. While recent work suggests weighting tasks by time spent, existing time shares are either based on coarse ONET data not intended for this purpose or estimated via black-box language models. We address this gap by proposing a principled method for estimating time shares for nearly 18,000 tasks that constitute nearly all U.S. jobs. Our estimates factor a task's time into (i) the expected frequency of the task, derived from ONET, and (ii) the time to complete a single instance of it. To estimate the latter, we solve a constraint s
arXiv:2608.05171v1 Announce Type: new Abstract: Generative AI (GAI) creates new opportunities for collaborative problem-solving (CPS), yet its role in shaping student interaction remains unclear. To address this gap, we conducted a six-week quasi-experimental study with 201 fifth-grade students in two conditions: with and without GAI. Chi-square analysis showed significant differences in CPS behavior distributions between groups. Compared with the control group, the GAI-supported group demonstrated more social behaviors, particularly engagement and conflict management, but less frequent cognitive behaviors such as task planning and solution reasoning. Lag sequential analysis further revealed distinct interaction patterns: while the control group followed a more conventional transition from listening to planning, the GAI group showed a robust pathway from task planning to conflict management to solution reasoning. Thematic analysis of AI interaction logs suggested that students used GAI
Across the country, schools are raising alarms about chronic absenteeism. News stories highlight rising numbers of missed days, legislators are demanding answers from districts, and educators are feeling the stress.
Section 504 of the Rehabilitation Act, which prohibits discrimination against students and other individuals with disabilities, is far less visible than the Individuals with Disabilities Education Act (IDEA) in school districts.
“The premise that cybersecurity is a back-office or administrative expense and that something might not happen — that needs to be changed,” says Fadi Fadhil, field CIO and director of field strategy at Palo Alto Networks. “CISOs and CIOs can steer that change by engaging in simplified conversations with university leadership. It’s a strategic effort, helping them understand how the investment reduces institutional risk.” When it comes to budgeting for their cybersecurity programs, higher education CISOs must overcome some unique hurdles, ranging from the federated nature of university IT…
Raising your hand in class and patiently waiting until you’re called before speaking. Sharing with classmates in a group project. Understanding what you’re feeling and how best to express it safely. These are a few examples of what social-emotional skills look like in the classroom. Social-emotional learning (SEL) houses a variety of skills, all of which have always been embedded in the K–12 experience. As recent research points more directly to the value of weaving these learning moments into the K–12 curriculum, educational technology has risen to meet the demands. The Evidence for Social-…
Article URL: https://twitter.com/gergelyorosz/status/2062861559009820976 Comments URL: https://news.ycombinator.com/item?id=48411421 Points: 10 # Comments: 1
On any given Tuesday afternoon, a dean at Morgan State University can pull live enrollment trend data without submitting a ticket, waiting for a report or following up with the IT department. At most higher education institutions, that same request can take about three weeks. The difference isn’t the data platform, however. It’s how the historically Black college is prioritizing data literacy. Timothy Summers, vice president of IT and CIO at the Baltimore-based institution, is betting the university’s artificial intelligence strategy on employees’ ability to effectively interpret, question…
The landscape for specialized colleges and universities such as art schools is shifting as higher education continues to evolve to fit emerging job markets and student interest. Founded in 1882, Cleveland Institute of Art continuously challenges itself to stay modern and relevant. Years ago, the school’s leadership had the vision to partner with the city to revitalize an area due for reinvigoration. The result was the Interactive Media Lab, which brings together the university, the city and private industry into a satellite campus that gives students and the community a space to…
As I wrapped up my student conferences, one conversation stuck with me. Steven had barely touched his final project for our computer science course, a virtual simulation of a piano, despite showing real promise earlier in the year.
Why boredom, quiet, and reflection matter for teen identity, agency, and imagination in a world shaped by constant screens. The post Running Your Own Race: Why Agency Begins in the Interior Life appeared first on Getting Smart .
Innovative Leader Award - Lauren Harwood of Dighton-Rehoboth Regional School District shares how she focuses her efforts on AI, CTE program, and cybersecurity
The recent ransomware incident involving Canvas has renewed attention on one of the most difficult decisions schools and technology providers can face: how to respond when sensitive student, faculty, or institutional data is stolen and threatened with public release. The post The Canvas ransomware attack shows why schools must focus on containment, not just recovery appeared first on eCampus News .
Some students with disabilities rely on assistive technology to learn, and they worry it could be swept up in the movement to get screens out of schools.
Hi HN, I'm Rosa, a massage therapist for over 30 years. I noticed my clients felt relaxed after a massage, but their stress and muscle tension always came back. A one-hour massage isn't always enough to combat a long workweek. I saw that people needed info on how to take care of their bodies between appointments, not just treatment. That's why we built MASSAGE BY ROSA – a wellness platform to meet this need. Here’s what we offer: It’s a two-part deal: -Massage Therapy: Hands-on therapy for pain relief – based on methods from South Florida. -Online Body Therapy Courses: This is what I want to share. I turned my knowledge into video courses teaching self-massage, workstation adjustments, and ways to release tension. It’s like having a therapist help you stay well. The Tech: We’re keeping it basic with a static site for course content and subscriptions, which lets us focus on making great video lessons. We're launching this to solve the problem of upkeep in physical wellness. It's for peo
As someone who’s dedicated my career to advancing the Science of Reading movement, I’ve seen firsthand what it takes to help every child become a strong, fluent reader.
The House VA Committee voted 19-0 to subpoena Oracle Health executives Larry Ellison and Mike Sicilia after learning the VA’s EHR contract ballooned from $10 billion to $27 billion — despite Oracle’s promise to Congress that it would absorb any costs beyond the original cap. The post Oracle Health Execs Subpoenaed After VA Contract Costs Nearly Triple appeared first on MedCity News .
There are small and large ways teachers, parents and administrators can cut back on the amount of screen time in school that detracts from learning.
Food stamps have a direct link to other critical resources, such as free breakfast and lunch at school. Now, experts say they're starting to see that link break.
A new model in Vermont emphasizes hands-on learning while stripping away frills like gyms and meal plans that increase costs. Other newly created institutions are also reimagining college.
Gov. Greg Abbott directed the Texas Higher Education Coordinating Board on Wednesday to work on pathways for students to earn bachelor’s degrees in three years. Such a program could have college students take about 90 hours of classes instead of 120 to earn a degree, thus reducing the cost of higher education, shortening the time […]
Ionis Pharmaceuticals’ zilganersen, brand name Zanvastro, is the first disease-modifying therapy approved for the neurological disorder Alexander disease. While Ionis has experience developing medicines for rare neuroscience indications, Zanvastro will be the first neurology product the company brings to the market without a commercialization partner. The post FDA Approves Ionis Pharma Drug, the First for Ultra-Rare Alexander Disease appeared first on MedCity News .
Before the beginning of each school year, teachers spend time getting their classrooms ready for students. Everything has its place, from whiteboards to pens, pencils and pushpins. With the adoption of cloud computing by school districts, the IT departments of K–12 schools also must figure out what goes where, but on a much larger level. Many districts operate in hybrid cloud environments, where some data and applications reside on servers that are on-premises and others reside with public cloud providers. It’s not a decision that should be taken lightly. Click the banner below to see how a…
The bill seeks to keep all special education, postsecondary, Native American, and elementary and secondary education activities within the agency.
The reality of K–12 IT teams is that most are very small — and, as a result, stretched thin. More than half of K–12 districts say they’re understaffed for everyday classroom technical support needs. On a given day, technicians may handle dozens of support requests for password resets, basic device troubleshooting, network connectivity issues, broken devices and more. When a district’s technical needs outweigh its team’s capacity to support it, something has to give. The solution? Agentic artificial intelligence. Agentic AI isn’t here to replace IT staff. Instead, it gives small departments a…
As teenagers return to school this fall, a new report suggests that many of them are moving through adolescence without the experiences that help prepare them for adulthood, activities like a part-time job, a club, a team and a regular family meal. The clearest evidence is teen employment, which alone has been declining for nearly […]
Three things ski patrol can teach about building a business The post Bringing Ski Patrol Lessons to Healthtech Entrepreneurship appeared first on MedCity News .
By approving the policy Thursday, the State University System of Florida’s board effectively shut undocumented students out of its 12 institutions.
The school board voted to close nine schools last year, and 27 were recently being evaluated to determine whether they should "remain operational." The post New Miami-Dade superintendent pauses look at closing 27 schools appeared first on District Administration .
The data, from nearly every public school in the country, reveals disparities in the way students are treated based on race, ethnicity, disability, sex and other factors. The post Federal civil rights data about students is finally out. It’s different under Trump appeared first on District Administration .
For students starting at a new school, the first day holds many questions: Will they see any familiar faces? Who will they sit with at lunch? Who will they play with at recess? “I’m this much nervous,” said Iman Fair-Seldon, 5, holding her hands about 4 inches (10 centimeters) apart. Iman will start first grade next […]
Schools are racing to teach students how to use AI. That is necessary, but it is not enough. The workplace advantage will increasingly belong to people who can decide when an AI output is useful, when it is incomplete and when it is wrong.
What if the next breakthrough in school design did not require a new building, a charter, or a pilot program imported from somewhere else? At Hidden Valley Middle School in Escondido, a public microschool operating inside a traditional campus has cut suspensions by 60 percent and reduced chronic absenteeism by 12 percentage points, schoolwide. This piece by Dr. Katie Martin and Dr. Devin Vodicka explores how California districts are building the internal capacity to generate and spread learner-centered models, and why that shift from showcase schools to generative systems may be the most important move in education reform right now. The post Districts Don’t Need Another New Model. They Need Systems That Build Them. appeared first on Getting Smart .