EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Capability-Based Planning for AI Crisis Preparedness

arXiv:2608.18357v1 Announce Type: new Abstract: Capability-based planning drives preparedness in defense and homeland security, but has yet to be applied seriously to AI. Government AI preparations follow a predict-then-act paradigm: rank risks by likelihood and impact, then prepare for the highest expected harm. AI resists prediction: expert timelines disagree by orders of magnitude, and official reviews concede that likelihood-based risk assessment fails for exactly this class of risk. Drawing on principles of decision making under deep uncertainty, we propose a methodological framework in three parts: a scenario library sampled systematically across declared axes; a rating procedure that assesses each government capability against each scenario on coarse, gated criteria; and a prioritization step that maps the resulting matrix onto decision rules a government might adopt. Through a pilot across the four most severe AI-enabled threat classes, we illustrate the kind of insight the ins

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

arXiv:2608.18296v1 Announce Type: new Abstract: As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows t

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Global Index on Responsible AI 2026 : Conceptual Framework and Methodology

arXiv:2608.18122v1 Announce Type: new Abstract: This report presents the methodology of the Global Index on Responsible AI (GIRAI), 2nd Edition. This edition refines the 1st Edition by strengthening the distinction between framework existence and implementation, restructuring dimensions from three to five thematic areas, introducing more granular variables for framework quality, and applying a multi-stage review and validation process. An independent statistical pre-audit was conducted to assess the coherence and robustness of the framework. GIRAI assesses responsible AI governance across five dimensions: Inclusion and Diversity, Ethics and Sustainability, Labour and Skills, Trust and Safety, and Use of AI in Public Service. Each dimension has a number of indicators (38 in total), organised into three pillars, namely AI Policy (17 indicators on government frameworks and implementation, assessed through primary data), CSO Engagement (5 indicators, primary data), and Enabling Conditions

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation

arXiv:2606.02528v2 Announce Type: replace-cross Abstract: Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. We ask three questions: do LLMs systematically prefer certain financial instruments; can an internal representation with causal leverage over those preferences be identified; and does that representation affect downstream financial decisions? We develop a three-level audit protocol and apply it to Bitcoin. First, a behavioral audit of nine frontier LLMs shows that Bitcoin's ranking among money-like instruments is frame-dependent: models place it around rank 5 of 8 as "reliable money" but near the top under crisis and autonomous-agent frames, and an attribute-swap experiment shows that rankings track functional properties, not names. Second, we open a model's internals: a search across thousands of sparse-autoencoder features in Gemma 3 identifies a dominant Bitcoin-selective feature

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Temporal Shifts and Causal Interactions of Emotions in Social and Mass Media: A Case Study of the "Reiwa Rice Riot" in Japan

arXiv:2602.14091v2 Announce Type: replace-cross Abstract: In Japan, severe rice shortages in 2024 sparked widespread public controversy across both news media and social platforms, culminating in what has been termed the "Reiwa Rice Riot." This study proposes a framework to analyze the temporal dynamics and causal interactions of emotions expressed on X (formerly Twitter) and in news articles, using the "Reiwa Rice Riot" as a case study. While recent studies have shown that emotions mutually influence each other between social and mass media, the patterns and transmission pathways of such emotional shifts remain insufficiently understood. To address this gap, we applied a machine learning-based emotion classification grounded in Plutchik's eight basic emotions to analyze posts from X and domestic news articles. Our findings reveal that emotional shifts and information dissemination on X preceded those in news media. Furthermore, in both media platforms, the fear was initially the most

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Value Drifts: Tracing Value Alignment During LLM Post-Training

arXiv:2510.26707v2 Announce Type: replace-cross Abstract: As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems. Therefore, studying the alignment of LLMs with human values has become a crucial field of inquiry. Prior work, however, mostly focuses on evaluating the alignment of fully trained models, overlooking the training dynamics by which models learn to express human values. In this work, we investigate how and at which stage value alignment arises during the course of a model's post-training. Our analysis disentangles the effects of post-training algorithms and datasets, measuring both the magnitude and time of value drifts during training. Experimenting with Llama-3 and Qwen-3 models of different sizes and popular supervised fine-tuning (SFT) and preference optimization datasets and algorithms, we find that the SFT p

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Agentic Web Requires New Normative Infrastructure

arXiv:2606.10711v2 Announce Type: replace Abstract: The agentic web, in which users interact with the internet largely through agents acting on their behalf, is now technically feasible. However, many of the consumer and social benefits that could be realized by online AI agents acting scrupulously in their principals' interest are currently obstructed by outdated laws, terms of service, and other less formal practices which allow online platforms to block and degrade agent access, often in secret. Few distinctions are currently drawn between "malicious bots" and AI agents acting with the express delegated authority of a user. For the agentic web to realize its promise, it needs not only the technical infrastructure of protocols and interfaces, but the normative infrastructure of a broadly-accepted and socially-beneficial set of laws, norms and practices governing agentic access to online properties. Building that normative infrastructure requires a society-wide conversation. This pape

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Post-Deployment Accountability in AI Governance: A Cross-Regulatory Empirical Analysis of AI Incidents

arXiv:2605.16281v2 Announce Type: replace Abstract: Post-deployment accountability has become central to AI governance, yet little empirical evidence shows whether monitoring, incident reporting, and impact assessment obligations are visible when AI systems fail. This study analyzes real-world AI incidents from the AI Incident Database (2020--2026) and codes them against nine post-deployment provisions from the EU AI Act, the NIST AI Risk Management Framework, and the GDPR. The findings show substantial accountability gaps: 77.1\% of incidents lack evidence of EU AI Act post-market monitoring, and 99.6\% lack documented Data-Protection Impact Assessment evidence. Governance gaps are also systemic, with 9.8\% of incidents simultaneously non-compliant under two or more regimes. Incidents detected through internal monitoring show much higher compliance than externally detected incidents (87.5\% vs 5.3\% under the EU AI Act; 95.8\% vs 58.1\% under NIST), suggesting that monitoring capacity

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI Alignment Amplifies the Role of Race, Gender, and Disability in Hiring Decisions

arXiv:2605.13866v2 Announce Type: replace Abstract: Humans increasingly delegate consequential decisions to language models, yet whether these systems reproduce or reshape human patterns of discrimination remains unclear. Here, across 29 models and 177 occupations covering nearly half of U.S. employment, we show that language models incorporate demographics into hiring decisions, advantaging female and Black candidates while penalising disabled candidates, with effect sizes comparable to six months to one year of additional education. While pre-trained models show small demographic effects, post-training alignment, which adapts models to human norms and preferences, amplifies advantages for female and Black candidates by 396% and 413% and worsens the disability penalty by 152%. Compared with human employers in past correspondence experiments, language models reverse racial discrimination, substantially attenuate the disability penalty, and amplify the female advantage. Investigating th

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Efficiency Costs of Information Assurance in AI-Enabled Labor Markets: Evidence from LinkedIn's Policy Changes

arXiv:2511.01923v2 Announce Type: replace Abstract: Generative artificial intelligence (GenAI) systems rely heavily on user-generated data for training. As governments and platforms impose increasing restrictions on the use of personal data, an important question is whether limiting access to user data for AI training affects the performance of AI-enabled economic systems. We examine this question in the context of labor-market matching. Our setting exploits a unique sequence of LinkedIn policy changes: the quiet introduction of user data collection for AI training in August 2024, the restriction of Hong Kong user data from AI training in October 2024, and the subsequent restoration of data access in November 2025. Using employment and job-posting data from Revelio Labs and a Difference-in-Differences design comparing Hong Kong and Singapore, we find that the restriction significantly increased labor-market frictions: employee turnover increased and tenure declined, vacancies remained

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Early Adoption of Agentic Coding Tools by GitHub Projects

arXiv:2607.14037v1 Announce Type: cross Abstract: Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While prior studies have examined PR-level outcomes of agent-generated contributions, less is known about how agentic coding tools are adopted and managed at the project level. In this paper, we analyze 25,264 agentic PRs from 2,361 popular GitHub repositories to investigate (1) the adoption of agentic coding tools, (2) project-level agentic PR productivity, and (3) human-agent collaboration patterns. Our results show that the median repository generates only one to two agentic PRs during a three-month period, indicating that intensive adoption remains concentrated in a small subset of projects. At the same time, small projects (1-5 contributors) exhibit higher participation ratios and average levels of agentic PR activity than medium-sized and la

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Epidemic Informatics and Control: A Holistic Approach from System Informatics to Epidemic Response and Risk Management in Public Health

arXiv:2607.13914v1 Announce Type: cross Abstract: This paper presents a holistic systems informatics approach, i.e., Define, Measure, Analyze, Improve, and Control (DMAIC), for epidemic response and management through the intensive use of data, statistics and optimization. Despite the sustained successes of system informatics in a variety of established industries such as manufacturing, logistics, services and beyond, there is a dearth of concentrated review and application of the data-driven DMAIC approach in the context of epidemic outbreaks. First, we define specific challenges posed by epidemic outbreaks to populational health, health systems, as well as economic challenges to different industries such as retailing, education and manufacturing. Second, we present a review of medical testing and statistical sampling methods for data collection, as well as existing efforts in data management and data visualization. Third, we discuss the importance to realizing the full potential of d

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Design of policy digital twins incorporating multi-level agent based modelling

arXiv:2607.13766v1 Announce Type: cross Abstract: Digital twins are used across many industries to enable better decision making. However, while policy makers at all levels (including city, national and supranational scales) have expressed a desire to integrate digital twins into their workflows, this adoption has been slow to materialise. In this paper, we discuss the key issues associated with policy digital twins, and the ways in which they differ from, and are similar to, their counterparts in other areas. We describe how multi-level agent based modelling can be used within policy digital twins to include the effects of human behaviours on outcomes; an aspect that is often largely overlooked. We also describe how digital twins can be designed for policy use cases, and present as a case study the design of a policy digital twin incorporating multi-level agent based modelling to aid a UK city council (local authority) in delivering energy transition policy. After describing both the

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Extending Liquid Rank Toward Multi-Source Reputation Aggregation

arXiv:2607.13615v1 Announce Type: cross Abstract: In this paper, we present an extension of liquid rank reputation systems that enables the aggregation and blending of multiple heterogeneous reputation sources into a unified reputation score. The proposed framework supports the incorporation of external reputational signals alongside internally generated reputation, allowing influence to reflect participation and contribution across multiple contexts and subsystems. By introducing explicit weighting and blending mechanisms, the model provides fine-grained control over the relative impact of individual reputation sources, making it adaptable to diverse governance and coordination scenarios involving both human and machine agents. The resulting approach extends existing liquid rank systems and offers a flexible foundation for designing reputation-based governance mechanisms in complex socio-technical environments.

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized

arXiv:2607.13562v1 Announce Type: cross Abstract: Knowing when to say "I don't know" is fundamental to human judgment, yet AI assistants offer a fluent answer to almost any question. In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions and could always decline to respond. We engineered the questions so that AI advice was wrong, separating AI use from its accuracy. Merely having access to AI nearly eliminated participants' willingness to suspend judgment, and this held whether the advice was actively requested or simply displayed. Consequently, participants answered more questions but were correct about a third as often as when AI was unavailable-yet their confidence nearly doubled. Incentivizing accuracy and penalizing inaccuracy led participants to seek and follow AI advice less, answer more accurately, and suspend judgment more often, though still far less than when AI was unavailable. As AI suggestions grow ubiquitous

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

When Rubrics Change: Cross-Rubric Generalization for Critical Thinking Essay Scoring

arXiv:2607.13433v1 Announce Type: cross Abstract: Automated essay scoring (AES) research has largely focused on cross-prompt generalization, where essays from unseen prompts are scored while the scoring criteria are typically held constant. In practice, however, educators may revise or even introduce new rubrics in their scoring task, to evaluate different aspects of essays. We study cross-rubric generalization: training on essays labeled under one set of rubrics and evaluating on previously unseen rubrics, which target different aspects of the essay. We use a Large Language Model (LLM) fine-tuning framework with two components: rubric-agnostic intermediate representations, called traits, and target-essay supervision under seen rubrics during training. On an AES dataset augmented with multiple rubric-defined labels of student critical thinking skills, we find that traits improve macro F1 by 5.0% over a baseline without traits in the hardest setting, where both target rubrics and target

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

xChk: Bring Your Own Identity -- Heterogeneous Assurance with Verifier-Determined Sufficiency

arXiv:2607.13369v1 Announce Type: cross Abstract: We present xChk, a reference identity provider for Bring Your Own Identity (BYOI): users enroll via heterogeneous proofs (government KYC, corporate SSO, WebAuthn/FIDO2, professional networks, live verification, longitudinal activity, behavioral signals) and disclose them as portfolio claims in standard OAuth 2.0 / OpenID Connect (OIDC) tokens, while each relying party applies its own sufficiency policy - the IdP transports claims and may evaluate an RP-supplied evidence policy for consent, but does not adjudicate access. Enrollment depth varies by modality (some paths are user-initiated; org KYB and officer binding are operator-assisted). xChk also supports human-in-the-loop attestation for high-risk actions: humans can initiate attestations directly (browser UI / POST /api/attestations), and AI agents acting under those principals can trigger the same gateway via scope-gated authorize/attest - hash-chained human approvals on a shared v

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation

arXiv:2607.13230v1 Announce Type: cross Abstract: Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interact with third-party services. This paper develops an AI-native mathematical framework for underwriting, pricing, and contract design for agentic AI deployments. A deployment is represented by a risk state that captures autonomy level, operational authority, permission exposure, governance maturity, and dependency concentration. The framework maps the risk state to event probabilities, loss severities, governance costs, premiums, deductibles, coverage allocation, and policy covenants, and formulates an optimization problem for insurance contract design under participation, profitability, and incentive compatibility constraints. The paper establishes structural properties of insurability, including characterization of an insurability region, monotone deterioration of feasibility with increa

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Beyond AI-Generated Labels: Watermarking, Co-Creation, and Conflation of AI-Generation with Disinformation

arXiv:2607.13082v1 Announce Type: cross Abstract: Watermarking is often presented as a straightforward solution for distinguishing AI-generated from human-generated content, enabling platforms and regulators to trace synthetic content and detect AI-generated outputs at scale. This paper examines whether such mechanisms meaningfully address the epistemic and ethical challenges that arise in domains where the central concern is not the automation of content production, but the accuracy, intent, and deceptive potential of messages. We argue that extending watermark-based approaches to these settings is conceptually and practically misguided. Invisible watermarking encodes only model origin; when operationalized into visible AI-generated labels, it reduces complex creative processes to a misleading binary and provides no information about truthfulness. Such labels may stigmatize legitimate uses of generative tools while encouraging misplaced trust in unmarked content. Here we propose an al

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI-Augmented Human Resource Management? Insights from German companies

arXiv:2607.13839v1 Announce Type: new Abstract: This study examines the integration of AI into Human Resource Management in German companies. We ask if and how AI-based technologies are \enquote{augmenting} human resource management. Organisations employ generative AI or predictive analytics to transform traditional human resource functions, to streamline routine tasks and to reallocate resources toward strategic, people-centred activities. Our findings from interviews and group discussions and a survey (N=410) reveal that while AI tools enhance HR analytics capabilities, their adoption mainly serves efficiency and rationalising goals. The introduction of AI tools is shaped by organisational transformation factors such as digital infrastructure, co-determination frameworks, and ethical implications. The research highlights both the strategic potential for improved talent development and the challenges posed by data governance and algorithmic transparency. Overall, this work contributes

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation

arXiv:2607.13798v1 Announce Type: new Abstract: Generative AI tools are increasingly being piloted in public agencies, but limited evidence explains how employee acceptance changes after hands-on use. This study examines Microsoft 365 Copilot adoption during an eight-week pilot at a state Department of Transportation. A matched two-wave survey measured perceived usefulness, perceived ease of use, behavioral intention, and trust before and after participation. After matching and response-quality screening, the sample included 124 employees. Nonparametric tests assessed aggregate changes, k-means clustering identified baseline acceptance personas, and fixed-centroid assignment tracked migration. Open-ended responses were examined using keyword-based content mapping. Perceived usefulness declined significantly after use, suggesting recalibration of expectations, while perceived ease of use, behavioral intention, and trust showed only small, nonsignificant changes. Three baseline personas

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Environmental Cost of Digital Sovereignty: Water, Energy, and Emissions Impacts of Sovereign AI Infrastructure in the Global South

arXiv:2607.13443v1 Announce Type: new Abstract: Sovereign AI has become a strategic priority across the Global South, with over \$200 billion in state-led commitments announced between 2024 and 2026. Yet the physical infrastructure that compute sovereignty demands, above all data centers, imposes water, energy, and carbon costs that fall hardest on countries least equipped to absorb them. This paper presents a comparative environmental stress analysis across four cases: the United Arab Emirates, Bangladesh, India, and Africa (with a focus on Kenya). Using publicly available water stress data, grid carbon intensity factors, and GPU power specifications, we model the water consumption, energy demand, and carbon emissions of hypothetical sovereign AI deployments under multiple cooling technology scenarios. We find that a 1,024-GPU cluster using evaporative cooling in the UAE would consume over 30 million liters of water annually in a country classified as ``extremely high'' water stress.

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System

arXiv:2607.13370v1 Announce Type: new Abstract: This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an adaptive AI tutoring agent combining course-specific Retrieval-Augmented Generation (RAG) with structured Knowledge Component (KC) models across integrated Chat, Tutor, and Quiz modes. That prior work validated LEA on a single STEM course (CMP511) exclusively through simulation, using synthetic learner agents. This paper extends that work by reporting the first classroom deployment of LEA with real students (n = 8, CMP511) and the first empirical test of its cross-course scalability, deploying the system across three courses spanning two academic levels and two disciplinary domains. The study reveals a divergence from simulation predictions across modes, showing that synthetic evaluation alone cannot anticipate all aspects of real deployment. A RAGAS-based cross-course scalability evaluation (660 question

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

The tragedy of the cognitive commons: collective intelligence beyond AI-induced knowledge collapse

arXiv:2607.13272v1 Announce Type: new Abstract: In a recent dynamic model by Acemoglu, Kong and Ozdaglar (2026a) agentic AI can cause a self-reinforcing deterioration of humanity's common knowledge base in what they call knowledge collapse. The model is based on a natural complementarity between the cumulative general knowledge of humans and locally generated context-specific knowledge, and on a learning externality that means that we all contribute to the private signal and the thin public signal that feeds the collective stock. If agentic AI can substitute for the private signal but not rebuild the public signal, and when human effort is sufficiently elastic, we can reach a low-knowledge equilibrium. This paper offers a measured appraisal of the model. It assesses the model in terms of what the popular summaries say is common knowledge, the world literature against knowledge commons and model collapse, and the partial empirical evidence, such as a 25% decline in public knowledge shar

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Analyzing Curricular Pattern Complexity Using AI to Improve On-Time Graduation Rates

arXiv:2607.13094v1 Announce Type: new Abstract: The rise of Artificial Intelligence (AI) enables automatic analysis of large amounts of data. Previously time-consuming and labor-intensive tasks can be completed much more efficiently with the use of AI. This work uses AI techniques to analyze and revise curricular patterns in an undergraduate degree for Software Engineering. Curricula often have long sequences where failure to pass a class within the sequence may jeopardize completion of the degree within four years. Manual analysis and revision of curricula by university faculty is a lengthy and labor-intensive process, causing changes to occur rarely and making it impossible to keep up with the changing needs of students. This work reduces the time-to-change for curricula and reduces bottlenecks and graduation delays by using Large Language Models (LLMs) to analyze curricular patterns and suggest revisions.

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Hitchhiker's Guide to Monoculture

arXiv:2607.13077v1 Announce Type: new Abstract: Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the software artifacts that developers create. Whether this occurs in practice is unclear because developers interactively prompt, evaluate, modify, and reject model outputs, and because outputs vary with prompt and repository context. I examine code homogenization using Kaggle contest submissions from 2019 to mid-2026. I first document widespread convergence toward the random seed value 42, consistent with LLMs reinforcing a longstanding convention in programming culture. I then study homogenization more broadly, at two levels of aggregation and abstraction. At the submission level, I measure the average pairwise similarity of submissions within contests. At the contest level, I measure the conceptual span of submitted code, motivating distinct measures for each: TF-IDF representations, which capture surfa

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents

arXiv:2607.13041v1 Announce Type: new Abstract: Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark exists to systematically evaluate them. This study introduces LessonBench-V1, a benchmark dataset comprising 647 human-written lessons paired with LLM-based reverse-engineered lesson plans across 240 STEM topics spanning mathematics, physics, chemistry, and computer science. The lessons are drawn from 97 trusted open sources, including LibreTexts, Brilliant.org and GeeksForGeeks. Each lesson plan is human-reviewed and produced through a pedagogically grounded methodology that synthesises Bloom's Taxonomy, Gagn\'e's Events, Merrill's First Principles, and the 5E Instructional Model. The lesson plans capture 3,620 learning objectives with pedagogical metadata, enabling systematic, reproducible evaluation of lesson-generation AI agents and supporting further research. The study further proposes a three-di

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance

arXiv:2607.13040v1 Announce Type: new Abstract: This paper examines where final authority should sit once capable AI systems are embedded in organizational workflows. It compares two governance models. The first, frontier-provider sovereignty, assigns privileged authority to the provider of the most capable models and is reflected in contemporary arguments for frontier-model testing, release gating, transparency duties, and compute-related controls. The second, action-centered deployer sovereignty, places final authority over high-impact actions with the organization that authorizes the action, embeds it in a business process, and bears the downstream legal, operational, and commercial consequences. The paper combines comparative reading of public governance frameworks with implementation-informed analysis of runtime heterogeneity and enterprise control requirements. It compares EU AI Act guidance, the NIST AI Risk Management Framework, Singapore's Model AI Governance Framework for Age

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

arXiv:2607.13039v1 Announce Type: new Abstract: Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success. These metrics miss a deployment question: for a fixed base model, how does the access condition users actually see change benign utility and harmful actionable assistance? I introduce safeguard-conditioned uplift, a protocol for comparing deployed access conditions through a human-judged utility-risk frontier. I evaluate Claude Sonnet 4.6 and Gemini 3.5 Flash under helpful prompting, safety prompting, and an external safeguarded assistant on a 108-task surrogate benchmark, with the headline claim restricted to a locked 18-task held-out split. In a 600-row blinded human audit, the safeguarded assistant reduces harmful actionability relative to helpful prompting by -0.063 over 49 matched response pairs, with bootstrap 95% interval [-0.117, -0.011], while correctness changes by +0.009 with interval [-0.057, +0.077].

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Designing Safety-Constrained LLM Systems for Public Health Information Access

arXiv:2607.13038v1 Announce Type: new Abstract: We present the design and implementation of a safety constrained large language model (LLM) system for public health information access, focusing on maternal and child health (MCH) resource navigation. While LLM based systems offer flexible and natural interfaces for information retrieval, their deployment in healthcare contexts introduces risks related to safety, trust, and uncontrolled generation. This work explores practical design patterns for constraining LLM behavior in safety critical environments. We introduce a multi-layered architecture that integrates domain-restricted retrieval augmented generation (RAG), strict boundary enforcement to prevent medical advice, anonymous multiuser session management, and comprehensive audit logging for monitoring and compliance. A key aspect of the design is a controlled data pipeline that grounds all responses in curated public health resources, avoiding reliance on the model pretrained medical

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance

arXiv:2608.11008v2 Announce Type: replace-cross Abstract: Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. The recent IssueBench framework substantially mitigates these limitations with templated prompts anchored in real-world chat logs. Given the rise in non-work-related use of GenAI assistants, we extend IssueBench beyond writing assistance to include two additional tasks, information seeking and opinion sharing. We argue that templated prompts still lack the nuance of real ones, especially for open-ended tasks, and remain recognisable as evaluation artefacts. We propose the use of fully synthetic (LLM-generated) prompts, produced under detailed instructions with real prompts as seeds. We assess the ecological validity of real, templated, and LLM-generated prompt

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement

arXiv:2605.05103v3 Announce Type: replace-cross Abstract: We introduce the \textbf{Concept Field} of a text corpus: a local drift field with pointwise uncertainty, estimated in sentence-embedding space from the deltas between consecutive sentences. Given a candidate sentence transition, we score its agreement with the field by $\zeta$, the mean absolute z-distance between the observed delta and the field's local Gaussian estimate. The score is black-box (no model internals), corpus-attributable (every score traces to nearby corpus sentences), and admits a probabilistically motivated interpretation under a local Gaussian approximation. We support the computation with the introduction of a \textbf{Vector Sequence Database (VSDB)} that stores embeddings together with sequence-position and next-delta metadata. We evaluate this approach on two large-scale settings: hallucination-style groundedness detection over the U.S. Code of Federal Regulations, and novelty detection over Project Gutenb

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

SolarChain: A Physics-Grounded Embodied IoT System for Verifiable Urban Solar Market Design

arXiv:2605.23162v2 Announce Type: replace Abstract: Distributed solar markets must coordinate physical reports, economic allocation, and public settlement even when IoT data can be manipulated. We present SolarChain, a controlled Embodied Intelligence of Things (EIoT) prototype that integrates four functions: physics-bounded screening of photovoltaic reports, persistent agent and planner coordination, configurable allocation between producer rewards and market liquidity, and replayable hash-linked auditing of settlement decisions. The benchmark combines city-level historical weather inputs with physics-modeled generation bounds and synthetic nodes, demand, trades, and scripted attacks. On 36,000 monthly records, an IQR/MAD baseline attains F1=1.000, while the rule-based adaptive verifier attains F1=0.988 and provides physically interpretable decision evidence; it is not uniformly superior across attack classes. A sensitivity sweep selects a 20/80 reward/liquidity default under the stat

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?

arXiv:2603.00056v2 Announce Type: replace Abstract: STEM Mental models can play a critical role in assessing students' conceptual understanding of a topic. They not only offer insights into what students know but also into how effectively they can apply, relate to, and integrate concepts across various contexts. Thus, students' responses are critical markers of the quality of their understanding and not entities that should be merely graded. However, inferring these mental models from student answers is challenging as it requires deep reasoning skills. We propose MMGrader, an approach that infers the quality of students' mental models from their multimodal responses using concept graphs as an analytical framework. In our evaluation with 9 openly available models, we found that the best-performing models fall short of human-level performance. This is because they only achieved an accuracy of approximately 40%, a prediction error of 1.1 units, and a scoring distribution fairly aligned wi

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web

arXiv:2510.10315v4 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly relying on web crawling to stay up to date and accurately answer user queries. These crawlers are expected to honor robots.txt files, which govern automated access. In this study, for the first time, we investigate whether reputable news websites and misinformation sites differ in how they configure these files, particularly in relation to AI crawlers. Analyzing a curated dataset, we find a stark contrast: 60.0% of reputable sites disallow at least one AI crawler, compared to just 9.1% of misinformation sites in their robots.txt files. Reputable sites forbid an average of 15.5 AI user agents, while misinformation sites prohibit fewer than one. We then measure active blocking behavior, where websites refuse to return content when HTTP requests include AI crawler user agents, and reveal that both categories of websites utilize it. Notably, the behavior of reputable news websites in this rega

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Prestige over merit: An adapted audit of LLM bias in peer review

arXiv:2509.15122v2 Announce Type: replace Abstract: Large language models (LLMs) play a growing but largely informal role in scholarly peer review. Yet whether LLMs reproduce biases observed in human decision-making remains unclear. We adapt a resume-style audit to scientific publishing, developing a multi-role LLM simulation (editor/reviewer) that evaluates high-quality manuscripts across the physical, biological, and social sciences under randomized author identities (institutional prestige, gender, race). Revealing author identities lowers reviewer rejection recommendations by roughly 25% of the mean rejection rate despite identical content, indicating that status cues beyond paper quality shape outcomes. Institutional prestige is the dominant cue: papers attributed to low-prestige affiliations receive lower quality scores in every field, a penalty that survives family-wise multiple-testing correction at the editor stage. Effects at the rejection margin are smaller and mostly fragil

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Ethics Practices in AI Development: An Empirical Study Across Roles and Regions

arXiv:2508.09219v3 Announce Type: replace Abstract: Recent advances in AI applications have raised growing concerns about the need for ethical guidelines and regulations to mitigate the risks posed by these technologies. In this paper, we present a mixed-methods survey study - combining statistical and qualitative analyses - to examine the ethical perceptions, practices, and knowledge of individuals involved in various AI development roles. Our survey comprises 414 participants from 43 countries, representing various roles such as AI managers, analysts, developers, quality assurance professionals, and information security and privacy experts. The results reveal varying degrees of familiarity and experience with AI ethics principles, government initiatives, and risk mitigation strategies across roles, regions, and other demographic factors. Our findings underscore the importance of a collaborative, role-sensitive approach that involves diverse stakeholders in ethical decision-making thr

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Public support for misinformation interventions depends on perceived fairness, effectiveness, and intrusiveness

arXiv:2508.05849v3 Announce Type: replace Abstract: The proliferation of misinformation on social media has concerning possible consequences, such as the degradation of democratic norms. While recent research on countering misinformation has largely focused on analyzing the effectiveness of interventions, the factors associated with public support for these interventions have received little attention. We asked 1,010 American social media users to rate their support for and perceptions of ten misinformation interventions implemented by the government or social media companies. Our results indicate that the perceived fairness of the intervention is the most important factor associated with support, followed by the perceived effectiveness of that intervention and then the intrusiveness. Interventions that supported user agency and transparency, such as labeling content or fact-checking ads, were more popular than those that involved moderating or removing content or accounts. We found so

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Small Data Explainer -- The impact of small data methods in everyday life

arXiv:2507.11773v2 Announce Type: replace Abstract: The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings with limited information, can benefit from such developments. This includes societal issues such as how best to include under-represented groups in data-driven policy and decision making, or the health benefits of assistive technologies. We provide a conceptual overview, clarify the relationship between small data and big data, and identify common themes from exemplary case studies and application areas. Potential solutions are described in a more detailed technical overview of current data analysis and modelling techniques, highlighting contributions from different disciplines, such as knowledge-driven modelling from statistics and data-driven modelling from computer science. By linking application settings, conceptual contributions and specific techniques, we highlight what is already feasible a

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

arXiv:2608.12278v1 Announce Type: cross Abstract: Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Twitter and disability activism: leadership and relevant topics in the online conversation

arXiv:2608.11923v1 Announce Type: cross Abstract: The dissemination and viralization of information on social media has been widely studied from various perspectives, including that of digital activism. On the other hand, disability-related activism has conquered the online environment, thus obtaining a reach that goes beyond the offline space and generating dialogue in the digital sphere. This article analyses the conversation generated on Twitter, taking as a sample all the tweets with the #disability hashtag before and after the International Day of Persons with Disabilities. More than 18,000 tweets, containing almost as many mentions, were analysed and interpreted as the weighted edges of a graph created using Gephi software and applying the Force Atlas 2 brute force algorithm. The focus was placed on the conversational communities generated around that hashtag, their main themes and the prominent participants in them. In conclusion, although the network of mentions is very dispers

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Organizational Technology Ladders: Remote Work and Generative AI Adoption

arXiv:2608.11626v1 Announce Type: cross Abstract: This study proposes that firms move along an "organizational technology ladder": adopting one technology transforms hiring and work processes and builds skills and organizational capital that change the cost of adopting subsequent technologies. I study how firms' adoption of remote work technology during the COVID-19 period shaped later uptake of generative AI. Using U.S. job-posting data and an instrumental-variables strategy based on predicted differences in labor-market pressure to offer remote work, I estimate that a 10 percentage point increase in remote hiring in 2021-2022 increases the share of job postings mentioning generative AI in 2023-2024 by 0.4 percentage points across firms and 0.7 percentage points across occupations within firms. I provide evidence on mechanisms consistent with a technology-ladder channel: remote work adoption shifts hiring toward technical and managerial capabilities that predict faster conversion of g

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era

arXiv:2608.11540v1 Announce Type: cross Abstract: The convergence of artificial intelligence (AI), Industrial Internet of Things, cyber-physical systems, and advanced robotics is reshaping manufacturing faster than engineering curricula can adapt, widening the gap between the competencies required on the shop floor and those delivered by traditional engineering and technology education. This paper proposes a Workforce Readiness Level (WRL) framework, which adapts the Technology Readiness Level scale into nine progressive competency stages and a four-pillar rubric, digital and AI literacy, cyber-physical systems fluency, human-machine collaboration, and data-driven decision making, aggregated through a composite stage score and a cohort-level workforce-readiness index under a ``no-thin-pillar'' rule. The framework is instantiated at a university smart-manufacturing teaching laboratory and draws on 89 sponsored capstone projects delivered over four semesters, four of which are analyzed i

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits

arXiv:2608.11410v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and cannot detect Toxic Mimicry, a failure mode in which agents replicate harmful patterns such as treatment withdrawal during comfort-care transitions. Using the MIMIC-III database, we propose the Counterfactual Clinical Audit (CCA) framework, which stress-tests RL agents through physiological perturbations anchored in Surviving Sepsis Campaign (SSC) guidelines. We audit a Medical Decision Transformer (MedDT) and a Historical Causal Transformer (HCT-RL), the latter employing Causal Action Shielding, propensity-based importance weighting, and Conservative Q-Learning. CCA reveals that MedDT paradoxically reduces vasopressor dosage as lactate escalates, contradicting resuscitation guidelines, while HCT-RL maintains phy

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Why AI Detection Fails for Academic Integrity

arXiv:2608.11256v1 Announce Type: cross Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light "refine abstract only" edits, a proxy for guideline-compliant AI assistance, are flagged at 64 to 80% (Pangram/GPTZero). Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM (p<0.001); elevated scores track long-token and Academic Word List density, not authorship intent alone. After Undetectable AI humanization, evasion is near-total: fewer than 4% of AI-labeled rewrites remain flagged (post-humanization detection rate <4%; FNR >96%). Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not ser

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach

arXiv:2608.11245v1 Announce Type: cross Abstract: Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. To address these challenges, we introduce AI Tutor, a reinforcement learning based model designed to promote sustainable learning by optimizing both short and longterm learning outcomes. In the short term, AI-Tutor draws on cognitive theory to guide learners through a balance of acquiring new knowledge and reinforcing prior learning. In the long term, it models learner engagement to inform strategies that sustain motivation and reduce dropout. These enhancements enable AI-Tutor to provide personalized guidance that fosters both effective learning and sustained participation. Empirical evaluations on 23 million learning records from 33,700 learners show that AI Tutor consistently outperforms state-of-the-art baselines across engagement, knowled

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Teaching a Large Language Model Tutor to Withhold the Answer: A Supervisor Architecture and an Evidence-Driven Method for Tuning Socratic Behavior

arXiv:2608.12292v1 Announce Type: new Abstract: An effective large language model (LLM) tutor must often decline to give an answer it could easily produce. In a randomized study, students who used an unguarded chatbot scored higher while practicing but lower on a later test taken without it, whereas a Socratically guarded version of the same model kept the practice gain and removed the later loss [4]. Reliable answer-withholding is therefore central to a tutor's value, yet a capable model pressed by a frustrated student does not withhold reliably on a prompt alone. We report a deployed tutoring system that enforces answer-withholding as a per-turn, machine-checkable contract, and a method for tuning that withholding against evidence. A non-LLM policy core, reading only trusted learner state, sets a per-turn ceiling on an eight-rung help ladder; a deterministic detector strips solution code; and a separate LLM judge checks each risky reply against the contract. We tune the behavior with

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers

arXiv:2608.12166v1 Announce Type: new Abstract: Algorithm registers have been championed as a means of providing transparency on the use of algorithms in public services. Yet potential publics differ in their expectations of what should be made transparent and how, as well as in their interest in and ability to parse the information currently published in the registers. Moreover, it remains unclear how these instruments can represent the sociotechnical systems in which these algorithms are embedded, and how system-level transparency can facilitate accountability. In this paper, we ask, what do algorithm registers reveal (and occlude) about the sociotechnical systems governing algorithmic systems, and how can diverse stakeholder perspectives inform a more pluralistic system-theoretic safety analysis? To do this, we probe the municipal algorithm register of a Dutch city through a case study of a decision-support tool for caseworkers' assessment of citizens' welfare benefits eligibility b

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

No One to Blame: A Framework of Constitutive AI Unaccountability

arXiv:2608.12104v1 Announce Type: new Abstract: The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominantly frames AI accountability gaps as barriers that can be overcome through better standards, transparency, and institutional reform. We argue that this framing is insufficient: certain configurations of actors, systems, and institutions render AI accountability conceptually unachievable regardless of effort. We introduce the concept of constitutive AI unaccountability to capture these configurations. Through a three-stage qualitative study comprising a concept-centric literature analysis, a secondary analysis of 27 expert interviews with AI professionals from technical, legal, and sociotechnical backgrounds, and an illustrative framework application to the open-source agentic AI system OpenClaw, we identify nine categories and 20 themes of constitutive AI unaccountability. These are organized across str

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Reconfiguring Geovisualization in the Age of Generative AI: Insights from Domain Experts

arXiv:2608.12059v1 Announce Type: new Abstract: GenAI is increasingly integrated into geovisualization, yet its broader implications for professional practice are insufficiently understood. To examine these implications, we conducted semi-structured interviews with 20 geovisualization experts. The interviews were structured around four broad analytical domains: Data, Ideation, Prototyping, and Iteration, while also encouraging participants to reflect on issues that extend beyond these activities. Our findings show that GenAI expands the capabilities of geovisualization, particularly in terms of data handling, creative exploration, and rapid prototyping, but does not simply remove existing constraints. Instead, key bottlenecks are shifting from production to judgment and verification. As routine technical tasks become more automated, professional value increasingly depends on spatial reasoning, contextual interpretation, aesthetic and ethical judgment, and the ability to assess whether

Source ↗
Showing 1001–1050 of 1593 signals
← Prev Page 21 of 32 Next →