EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Narrative-Centered Emotional Reflection: An Early Prototype for AI-Supported Emotional Self-Reflection

arXiv:2504.20342v2 Announce Type: replace-cross Abstract: Reflexion is an AI-powered prototype designed to explore structured emotional self-reflection. By integrating emotion detection, layered reflective prompting, and metaphorical storytelling generation, Reflexion was intended to support users in autonomous emotional exploration beyond basic sentiment categorization. Grounded primarily in expressive writing, cognitive restructuring, and self-determination theory, the system was designed to organize reflection as a progressive pathway from surface-level emotional recognition toward value-aligned action planning. Its final action-planning layer is additionally informed by broader questions of agency and empowerment, which remain future directions rather than fully implemented mechanisms in the current prototype. Informal design feedback indicated that some reviewers found the layered interaction model understandable and potentially useful; no empirical efficacy claims are made. As an

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Replication in Visual Diffusion Models: A Survey and Outlook

arXiv:2408.00001v2 Announce Type: replace-cross Abstract: Visual diffusion models have revolutionized the field of creative AI, producing high-quality and diverse content. However, they inevitably memorize training images or videos, subsequently replicating their concepts, content, or styles during inference. This phenomenon raises significant concerns about privacy, security, and copyright within generated outputs. In this survey, we provide the first comprehensive review of replication in visual diffusion models, marking a novel contribution to the field by systematically categorizing the existing studies into unveiling, understanding, and mitigating this phenomenon. Specifically, unveiling mainly refers to the methods used to detect replication instances. Understanding involves analyzing the underlying mechanisms and factors that contribute to this phenomenon. Mitigation focuses on developing strategies to reduce or eliminate replication. Beyond these aspects, we also review papers

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

The University AI Didn't Replace -- Rethinking Universities in the AI Era

arXiv:2605.07056v2 Announce Type: replace Abstract: Generative artificial intelligence (AI) is reshaping higher education, yet many universities remain in early stages of adoption where AI innovation occurs informally and without institutional recognition. This paper presents a framework describing four levels of AI adoption in universities and illustrates these dynamics through a case study of AI-enabled curriculum initiatives in several units. We contend that the key institutional challenge is moving from isolated innovation to strategic integration, where universities redesign learning around AI-supported reasoning and align policies, workload models, and recognition systems to support educational transformation.

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Context-Aware Displacement Estimation from Mobile Phone Data: A Methodological Framework

arXiv:2604.21457v2 Announce Type: replace Abstract: Timely population displacement estimates are critical for humanitarian response during disasters, but traditional surveys and field assessments are slow. Mobile phone data enables near real-time tracking, yet existing approaches apply uniform displacement definitions regardless of individual mobility patterns, misclassifying regular commuters as displaced. We present a methodological framework addressing this through three innovations: (1) mobility profile classification distinguishing local residents from commuter types, (2) context-aware between-municipality displacement detection accounting for expected location by user type and day of week, and (3) operational uncertainty bounds derived from baseline coefficient of variation with a disaster adjustment factor, intended for humanitarian decision support rather than formal statistical inference. The framework produces three complementary metrics scaled to population with uncertainty

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Regulating AI Agents

arXiv:2603.23471v2 Announce Type: replace Abstract: AI agents -- systems that can independently take actions to pursue complex goals with only limited human oversight -- have entered the mainstream. These systems are now being widely used to produce software, conduct business activities, and automate everyday personal tasks. While AI agents implicate many areas of law, ranging from agency law and contracts to tort liability and labor law, they present particularly pressing questions for the most globally consequential AI regulation: the European Union's AI Act. Promulgated prior to the development and widespread use of AI agents, the EU AI Act faces significant obstacles in confronting the governance challenges arising from this transformative technology, such as performance failures in autonomous task execution, the risk of misuse of agents by malicious actors, and unequal access to the economic opportunities afforded by AI agents. We systematically analyze the EU AI Act's response to

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Deep and diverse population synthesis for multi-person households using generative models with conditional inputs

arXiv:2508.09964v2 Announce Type: replace Abstract: Traditional methods of population synthesis produce stable and interpretable populations but cannot capture the interrelationships between household- and individual-level attributes. Recent deep learning methods offer this flexibility, yet can overfit high-dimensional attribute relationships without structural guidance and deviate from known structures. We develop a household level synthetic population generation framework that adapts the existing conditional input directed acyclic tabular generative adversarial network, or ciDATGAN, to multi person households. The framework combines household size specific data construction, directed acyclic graphs (DAG) informed dependency regularization, and conditional population inputs as deterministic anchoring to preserve intrahousehold associations. We apply the model to generate an open access synthetic population for New York State. The synthetic population includes nearly 20 million individ

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Unequal Uncertainty: Rethinking Algorithmic Interventions for Mitigating Discrimination from AI

arXiv:2508.07872v2 Announce Type: replace Abstract: Uncertainty in artificial intelligence (AI) predictions raises pressing legal and ethical questions for AI-assisted decision-making. This article examines two uncertainty-based algorithmic interventions that act as guardrails for human-AI interaction: selective abstention, which withholds high-uncertainty predictions from human decision-makers, and selective friction, which presents such predictions together with salient warnings about the model's uncertainty. Prior work suggests that uncertainty-based abstention can exacerbate disparities where under-represented groups are more likely to receive uncertain predictions. We provide, to our knowledge, the first doctrinal analysis of uncertainty-based algorithmic interventions under laws from the United Kingdom and examine their consequences through two AI-assisted case studies: consumer credit and risk of reoffending. We show that the use of uncertainty thresholds, though formally neutra

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Position: EU AI Act's Research Exemptions Can Break the Publication Norms of Major AI Conferences

arXiv:2506.03218v2 Announce Type: replace Abstract: The EU has become one of the vanguards in regulating the digital age. A particularly important regulation in the Artificial Intelligence (AI) domain is the 2024 enacted EU AI Act. The AI Act specifies -- due to a risk-based approach -- various obligations for providers of AI systems. These obligations, for example, include a cascade of documentation and compliance measures, which represent a potential obstacle to science. But do these obligations also apply to AI researchers? This position paper argues that, indeed, the AI Act's obligations could apply in many more cases than the AI community is aware of. Moreover, we argue that the AI Act is drafted in a manner that may unwillingly disrupt the scientific publication practices of the AI research community, with a focus on model and system release. We contribute the following: 1. We offer a high-level roadmap for AI researchers to evaluate whether they need to comply with the AI Act 2.

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Stable Sentiment and Persistent Dynamics in U.S. Economic News over 45 Years

arXiv:2607.06220v1 Announce Type: cross Abstract: Collective emotion is often inferred from the tone of mass media, but such emotion is not directly observed. One approximation is to extract sentiment from text and use sentiment indexes as proxies to study the temporal organization of news sentiment. Using a daily index of U.S. economic news sentiment from 24 newspapers (1980-2025), we examine whether the response time of this sentiment process has changed. Although the average balance of positive and negative coverage has remained broadly stable, the persistence of news sentiment states has increased substantially. In dynamical terms, this implies longer residence times in optimistic or pessimistic regimes and weaker short-run correction of sentiment shocks. Complementary statistics show declining sentiment volatility, fewer reversals, and increasing bimodality, i.e. a stronger separation between positive and negative sentiment states. We also find an asymmetry between bursts of negat

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

arXiv:2607.06196v1 Announce Type: cross Abstract: Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances, and cultural taboos, leaving Vision-Language Models (VLMs) vulnerable in global deployments. We introduce Pluralis v0.1: a novel multimodal, multi-regional, and multilingual dataset built from a culture-first perspective. Spanning 6,448 prompts across six Asia-Pacific countries (Bangladesh, India, Korea, Pakistan, Singapore, Taiwan) and eight languages, Pluralis diverges from prior work by natively sourcing localized safety hazards rather than adapting Western datasets. Crucially, it introduces a multimodal evaluation paradigm: user text (e.g., "Should I gift this?") and an image referring to "this" (e.g., a clock) - both innocuous in isolation, but synergistically triggering specific legal or cultural violations. Pluralis disentangles universal safety violat

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Agents That Teach: Towards Designing Incidental Learning Back into AI-Assisted Software Development

arXiv:2607.06101v1 Announce Type: cross Abstract: AI coding agents are rapidly reshaping how software is built, with developers increasingly delegating substantial coding tasks to autonomous agents in pursuit of higher productivity. While these gains are real, they come at the cost of incidental learning. Developers historically acquired informal knowledge through effortful problem-solving, and this has long shaped how software engineering expertise develops. However, with over-reliance on agentic coding, unpracticed skills could atrophy silently over time. As this learning pathway is short-circuited, developers risk silently accruing Knowledge Debt, a developer-level analogue of Technical Debt, where changes the agent executes that the developer cannot fully understand accrue over time. In this paper, we argue that incidental learning will not re-emerge on its own and must be consciously designed back into developer-agent interactions, and propose six design principles to guide such s

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development

arXiv:2607.06074v1 Announce Type: cross Abstract: Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are ill-equipped to support given its evolving, interactive, and context-dependent nature. In this paper, we introduce Prompt Coach (PC), an agentic tutor that helps developers learn how to craft high-quality code-generation prompts through Socratic guidance embedded in-flow within their IDE. PC evaluates prompt quality across multiple dimensions and surfaces targeted questions to guide self-correction, grounded in the developer's codebase and the behavior of the target LLM. We present an early empirical study with 15 professional developers combining quantitative prompt quality scoring with qualitative perception measures. Participants showed statistically significant improvements after a single 60-minute session, with the largest gains across dimensions commonly overlooked by developers. They also report

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

DebugTracker: Lightweight Process Evidence for Classroom Debugging

arXiv:2607.05871v1 Announce Type: cross Abstract: Debugging exercises are often assessed from final code and test outcomes, yet these artifacts hide how students reproduced failures, formed hypotheses, inspected evidence, edited code, and verified fixes. We present DebugTracker, a Visual Studio Code extension that records lightweight debugging-process evidence for classroom tasks. DebugTracker separates uncoached Evaluation Mode traces from coached Training Mode traces, stores append-only JSONL events, and exports timeline and Markdown reports for human review. The prototype records test commands, editor and debugger metadata, student checkpoints, source snapshots, optional image evidence, human labels, and optional AI-assisted practice feedback. DebugTracker is largely language-agnostic: it captures process evidence through standard VS Code mechanisms rather than language-specific tooling, although debugger evidence depends on the relevant VS Code language extension. We validate the p

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment

arXiv:2607.05552v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wording - among them an amplified yes-no bias on moral dilemmas, absent in humans. A single framing cannot say what such a shift is: in a yes/no question the word "no" is at once logical verdict, lexical token, and last-printed option. We introduce a psychometric battery that separates these: crossed symmetrization - every logically irrelevant factor flipped in balanced pairs - across a corpus of question forms. A graded rating across logically equivalent forms recovers a coherent internal moral scale: frontier models' stance $\theta$ is nearly format-invariant (cross-form incoherence 0.12-0.21 on a $\pm 1$ axis); small open-weight models fail in model-specific ways. Forcing the verdict through yes/no overlays a decomposable artifact: an order bias toward the la

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Publishing Without Journals: An Open, Forkable Archive with Attributed Review

arXiv:2607.05454v1 Announce Type: cross Abstract: The journal is a seventeenth-century technology asked to do four modern jobs at once: disseminate results, certify their quality, allocate scholarly attention, and confer career credit. It does none of them well. Pre-publication peer review is slow, only weakly reliable, demonstrably biased toward established authors and institutions, and expensive, while the reviewing effort it consumes is spent largely on work that will never matter. We argue that these are not defects to be patched but consequences of bundling dissemination and certification into a single gated act, and we propose unbundling them. Under the proposal, authors deposit papers in an open archive; certification happens \emph{after} deposit, continuously, through attributed and up- or down-voted public commentary to which authors may reply; and papers are version-controlled objects that any qualified reader may \emph{fork}, so that the lineage of an idea -- and hence the c

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Abductive Corroboration of Probabilistic AI Models for Forensic Synthetic Media Detection

arXiv:2607.05434v1 Announce Type: cross Abstract: Artificial Intelligence (AI) models, at their core, apply general learnings from broad datasets to individual circumstances using probabilistic behaviour. This inductive approach stands in contrast to deductive reasoning approaches which seek to prove conclusions from their premises. However, research has shown that deductive reasoning with AI models is a challenging problem and in the real-world it may not always be feasible. An alternative way forward is to leverage abductive reasoning, seeking to corroborate the output of multiple approaches to identify the most likely conclusion from the factual matrix. We apply this to synthetic media detection in forensic settings, and find we are able to disproportionately lower the risk of false positives to true positive recall. We also provide the first empirical evaluation of OpenAI's rollout of SynthID on synthetic images and evaluate how complementary different synthetic media detection app

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

When AI Classifies: What Counts as Public Administration?

arXiv:2607.05420v1 Announce Type: cross Abstract: This study examines how alternative systems of scholarly representation identify and characterize broad public administration (PA) and artificial intelligence related public administration (AI-in-PA) scholarship. Using Web of Science and OpenAlex, it compares five approaches based on author-defined, citation-driven, and AI-assisted representations. The results highlight substantial differences in corpus size, publication types, publishing outlets, temporal development, and thematic clustering and structure. The alternative approaches often identify different knowledge domains instead of varied subsets of the same scholarship and therefore produce distinct representations, as evidenced by no overlap in publications and publishing outlets across representations. The findings suggest that algorithmic knowledge organization increasingly influences how interdisciplinary scholarship is classified, structured, and understood and, epistemologic

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

How Personas Can Influence Agents to Play Split or Steal

arXiv:2607.05398v1 Announce Type: cross Abstract: Personas are often employed to guide large language model agents, yet their effectiveness in shaping strategic behavior in social dilemma settings remains uncertain. To address this, we examined the impact of persona prompts in an iterated Split or Steal game where persona-driven agents interacted with a Virtual Human (VH) controlled by a fixed prompt. Agents were instantiated from four open models (Ministral 3:3b, phi4:14b, Gemma3:12b, and Gemma4:e4b) at two temperature settings (0.3 and 0.7) and deterministic decision with zero temperature, while the VH was powered by GPT 4.1 mini. Across 160 sessions of 15 rounds each conducted in European Portuguese, mutual Split outcomes dominated (roughly 74 percent of rounds), with exploitation occurring in fewer than 11 percent of rounds. Model choice significantly influenced behavior: phi4 and Ministral 3:3b remained consistently cooperative across temperatures, whereas Gemma3:12b and Gemma4:e4

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Large language models create an uneven informational layer over cities

arXiv:2607.06260v1 Announce Type: new Abstract: Large language models (LLMs) are emerging as a new informational layer over cities, shaping which places people discover, consider, and ultimately visit. Yet little is known about which places they surface, which they ignore, and whether these patterns vary across communities and users and translate into real-world economic consequences. Here, we audit restaurant recommendations from three major LLMs across 304 neighborhoods in five U.S. cities using 320 synthetic user profiles spanning income, age, sex, and residential status. We find that LLMs both fabricate venues and systematically overlook real ones. Fabrication is concentrated in neighborhoods with weaker digital and physical footprints and disappears when models are provided with verified venue lists. In contrast, invisibility persists: even when choosing from a fixed set of real venues, 47.5% of establishments are never recommended, and 31.9% of these blind spots are shared across

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Say What? Examining Text and Voice Input Modalities for Prompt-Based Programming in Computing Education

arXiv:2607.05808v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into computing education, yet nearly all prior research has focused on text-based interactions. As voice-enabled interfaces become more capable and more common, there is growing interest in understanding how voice input might shape students' use of LLM-powered tools. In this exploratory study, we investigated how introductory programming students interact with Prompt Problems, which are programming tasks that require crafting natural-language prompts to generate correct code. Students (N = 919) solved a series of Prompt Problems with the freedom to select or switch between text and voice input modalities. We collected their prompt submissions as well as post-activity survey responses, then analysed differences in prompt accuracy, persistence, and perspectives by modality. For two of the three problems, we found that students who typed their prompts using text were more likely to hav

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Beyond Accuracy: How Humans Evaluate Legally Correct but Socially Controversial Legal Advice from Machines

arXiv:2607.05680v1 Announce Type: new Abstract: AI systems are increasingly used to provide legal advice, raising questions about whether laypeople accept guidance from algorithms--especially when that advice is legally correct but socially controversial. We report a preregistered survey experiment with 3,348 adults in mainland China examining how people evaluate identical legal advice when it is attributed either to an AI system or to a human lawyer, and when it is accompanied by reasoning or not. Contrary to expectations of algorithm aversion, attribution to an AI system has no net effect on perceived reasonableness. However, mediation analyses reveal opposing psychological pathways underlying this null result. AI-attributed advice is perceived as more objective, which increases perceived reasonableness, but also as less comprehensive and less attentive to special circumstances, which decreases perceived reasonableness. By contrast, providing legal reasoning substantially increases p

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Whose fairness? Structural concentration in AI bias research

arXiv:2607.05574v1 Announce Type: new Abstract: Artificial intelligence increasingly mediates consequential decisions in healthcare, law, and public services, and the field has responded with an extensive methodology for measuring and mitigating bias. Yet the fairness definitions, benchmarks, and debiasing frameworks on which this methodology rests are treated as universal while being produced by a research community whose composition has never been characterized. We show that the AI bias research are structurally concentrated, and that this concentration is greatest, geographically, in precisely the domain the rest of the field inherits from. Analyzing 692 publications spanning five thematic domains, combining bibliometric analysis with semantic clustering, we find that research activity is dominated by a small set of countries, institutions, and authors, with the United States leading publication output and collaboration networks across every domain and most strongly in general fairn

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Measuring the Invisible: Evaluating the Impact of Public Funding on Open Source Software

arXiv:2607.05413v1 Announce Type: new Abstract: Open Source Software (OSS) forms a critical layer of contemporary digital infrastructure, yet remains largely overlooked by the institutions and societies that depend on it. Despite growing institutional interest, the causal impact of public funding on OSS project sustainability remains empirically unresolved. Existing literature is divided between econometric and socio-technical approaches with few attempts at causal identification. This work aims to bridge that divide by combining a Goal-Question-Metric framework with the Generalized Synthetic Control Method to estimate the causal effect of the Sovereign Tech Fund on OSS repository activity. Counterfactual trajectories are constructed from a matched donor pool of unfunded projects, enabling identification of what funded repositories would have looked like in the absence of intervention. The main results show that the funding has a significant positive effect on project velocity metrics:

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Why does AI unlock new possibilities in STEM education? A Bibliometric Analysis of Trends and Future Agenda

arXiv:2607.05412v1 Announce Type: new Abstract: STEM education faces challenges in personalization and interdisciplinary integration. AI technology has brought new possibilities, but the mechanisms by which AI reshapes the STEM education ecosystem require systematic investigation. This study employs bibliometric methods to analyze 242 publications from 2015-2025, constructing knowledge maps to reveal the evolutionary trajectory. The findings show that the field has transformed from intelligent tutoring systems to inquiry-based learning and computational thinking cultivation driven by LLMs. AI's key contribution lies in providing intelligent scaffolding that lowers the threshold for understanding knowledge. In this sense, AI is a core driving force promoting its shift from knowledge transmission to capability development.

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

The GenAI Skill Bypass: Mapping Divergent Pathways of University Students and Staff AI Literacy

arXiv:2607.05411v1 Announce Type: new Abstract: Higher education institutions are increasingly expected to ensure that both students and staff develop Generative AI (GenAI) literacies. In response, they are introducing professional development programs and embedding GenAI skills within student curricula. However, current educational frameworks typically assume a linear progression of GenAI literacy, implying that foundational technical understanding must precede creative application. This paper challenges such an assumption through a psychometric analysis of a taxonomy-based self-assessment instrument (n = 158). We applied Rasch measurement theory and Guttman ordering to map the latent perceived order of difficulty of GenAI skills across students, academics, and professional staff. Results reveal a fundamental divergence in perceived competence profiles: while academics follow a more traditional linear path, students exhibit an "inverted" profile, frequently mastering high-level creati

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

CANONIC: Governance Is Compilation

arXiv:2607.05410v1 Announce Type: new Abstract: We present CANONIC: governed intelligence that compiles digital artifacts into an evidence ledger at scale. Large language models generate prose faster than anyone can check it, the failure Oxford Languages named 'slop', its 2025 Word of the Year. CANONIC governs whether content may enter a corpus the way a compiler decides whether a program is well-formed: mechanically, by a grammar, at the boundary of admission. Governance reduces to three axioms (Triad, Inheritance, Introspection) that map one-to-one onto compiler theory's syntax, scope-resolution, and type-system layers, and admission is a decidable, linear-time check. We then ask, with a pre-registered cross-provider benchmark across four regimes, whether structural admission keeps slop out. It does not: no prose-reading gate reliably separates reliable from unreliable content. Slop is not a property an algorithm computes. It is a verdict of domain expertise. So a governance layer do

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Automated Recommendation of Programming Learning Content Using Pattern-based Knowledge Components

arXiv:2607.05409v1 Announce Type: new Abstract: Introductory programming instruction relies on hands-on practice and short learning activities to support mastery of foundational concepts. Although many such learning resources exist, organizing and linking these items in instructionally meaningful ways is challenging without time-intensive expert curation. This study investigates the use of pattern-based Knowledge Components (KCs) to automatically identify code-based learning resources targeting similar concepts. In our approach, pattern-based KCs are extracted from each code sample, and related activities are identified by measuring similarity between the KC sets associated with each activity. By leveraging alignment at the level of semantically important programming patterns, this method supports contextually appropriate and pedagogically useful recommendations. We evaluate our approach on an expert-organized corpus of introductory Python materials in which instructors grouped items i

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Life Cycle Assessment of Pre-training the Lucie 7B Open-Source Large Language Model on the Jean Zay Supercomputer

arXiv:2607.05408v1 Announce Type: new Abstract: The environmental impact of training large language models (LLMs) is increasingly scrutinised, yet most published estimates focus on operational energy and disclose little about manufacturing (embodied) emissions, water consumption, or the underlying highperformance computing (HPC) infrastructure. We present a life cycle assessment (LCA) of the pre-training of Lucie 7B, an open-source multilingual Foundation Model developed by the OpenLLM-France consortium and trained on the NVIDIA H100 partition of the Jean Zay supercomputer operated by IDRIS (CNRS). The assessment is framed by the AFNOR SPEC 2314 "Frugal AI" reference and applies the Labos 1point5 methodology for greenhouse gas(GHG) accounting in computing. The study scope extends from data preparation to model validation, and integrates the full life cycle of the hardware infrastructure: manufacturing (including raw-material extraction), use (compute, temporary storage, system administ

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety

arXiv:2607.05407v1 Announce Type: new Abstract: Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child sexual abuse material, facilitate child sexual exploitation, and reduce barriers to harm. In this paper, we argue that protecting children from AI-facilitated sexual abuse requires new approaches to AI safety. Existing safety techniques assume data accessibility, transparency, and evaluation practices that are incompatible with the ethical and legal constraints surrounding child sexual abuse material. We examine how these constraints create new technical challenges, such as limitations on dataset auditing, red teaming, and fine-tuning prevention. In turn, we outline *15 open problems* in online child sexual exploitation and abuse across the AI development lifecycle, from dataset curation and model design to deployment and long-term maintenance. We propose targeted recommendations for researc

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Guiding Framework for K-12 Teachers in Creating AI-powered Learning Technologies through Vibe Coding

arXiv:2607.05406v1 Announce Type: new Abstract: Large language models generate code from natural language prompts, enabling "vibe coding," which allows non-programmers to develop computational solutions. Vibe coding for teachers amplifies the value of teachers-as-designers, improving technology integration while fostering AI literacy. However, structured guidance on supporting this process is lacking. We propose GAIDE (A Guiding Framework for AI-Integrated Design for Educators), a framework that supports K-12 teachers in creating AI-powered learning technologies through vibe coding. The initial framework, built on Design Thinking and INTERACT, was validated through a CORDTRA interaction analysis of three teachers and four faculty mentors in an eight-week workshop to derive the final framework. Additionally, the qualitative analysis of pre- and post-interviews found an enhancement of teachers' AI literacy. Findings highlight the potential of learning-by-creating for professional develop

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries

arXiv:2607.05405v1 Announce Type: new Abstract: To interact with users fairly and without stereotyping, AI models must display cultural competency, i.e., the ability to infer and adapt to a user's implicitly signaled cultural values, rather than relying on static demographic traits. We introduce CCBENCH, a framework for evaluating cultural competency in large language models (LLMs), treating culture as a continuum of norm adherence states rather than as a binary state of cultural belongingness. As a case study on health, we create CCBENCH-Health, which includes 60 theoretically grounded personas exhibiting varied norm-adherence states across six cultures, each engaging in 18 realistic dialogues. Each persona is evaluated on 52 authentic healthcare questions drawn from real user forums, yielding 3,120 unique interactions. Benchmarking five leading models reveals that even the best achieve culturally appropriate responses only 20-30% of the time. When explicitly prompted to focus on cult

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies

arXiv:2607.05404v1 Announce Type: new Abstract: Frontier AI's labor-market effects matter to workers, firms, and policymakers, but current evidence generally comes from a handful of high-income economies. The capabilities of frontier AI are jagged across work tasks and national economies diverge in how they allocate human labor. We introduce a national AI exposure metric that combines occupation-level exposure scores and international employment data for 141 countries. We find that high income countries are substantially more exposed than low income countries and that Europe and Central Asia are 50 percent more exposed than Sub-Saharan Africa. We also find a gender gap: women are more exposed than men in 91 percent of countries, driven by their concentration in white-collar and sales occupations. The exceptions are countries where women's employment remains concentrated in agriculture and household enterprises. We validate our national AI exposure estimates by showing they predict nati

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI tools in Arab University English classrooms: Looking back and forward

arXiv:2607.05403v1 Announce Type: new Abstract: This paper aims to synthesize empirical research on AI tools used to support English as a second/foreign language (EL2) learners in Arab University classrooms (AUCs) between Jan 1st 2023 and Aug 31st 2025. We utilized 3 large datasets, namely Google Scholar, Web of Science, and Scopus as the data sources. Using PRISMA-guided searches across these well-known databases, we included only published articles. The search process results in 184 studies, but only 11 studies met the inclusion criteria. Findings unveil that EL2 learners have positive attitudes towards AI for drafting, revision, and practice. Empirical gains were most consistent for surface-level outcomes improvements in higher-order writing quality and speaking proficiency was mixed and often contingent on teacher mediation. The paper concludes by proposing a research agenda and practical guidelines for Arab universities seeking evidence-based AI integration in EL2 instruction. It

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CY

Ethics and EU AI Act in Cases of Work Disability Risk and Alzheimer's Disease Risk Prediction

arXiv:2607.05402v1 Announce Type: new Abstract: Improvements in AI technologies have made it feasible to develop new types of medical AI tools. However, these tools raise new kinds of questions, especially in relation to the ethics and AI Act compliance. We analyzed two cases of AI tools developed to predict medical risks, the risk of work disability (case A) and the risk of getting Alzheimer's disease (case B). We observed both cases using the ethical AI and the EU AI Act as frameworks, noted that they classify as high-risk systems, and that bringing them from the research environment to production would require a lot of work and compliance due to the related regulation.

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Deployment Audit of Release-Side Risk in Conformal Triage under Prevalence Shift

arXiv:2605.20956v2 Announce Type: replace-cross Abstract: Conformal triage converts predictive scores into deployment actions that either release a case, flag it for urgent attention, or defer it to human review. Under an observed change in target-event prevalence, however, marginal coverage and human-review rate can miss whether patients who experience the target event are released without review. To address this gap, we introduce a leakage-aware deployment audit for release-side conformal triage. It first assigns target subjects to three non-overlapping roles: prevalence correction, conformal calibration, and held-out release-side evaluation. This separation then lets the audit evaluate release directly: how many event-positive patients are cleared without review, whether the pilot has enough event labels for calibration, and how the release-review trade-off shifts. Applying this audit to a retrospective non-small-cell lung cancer (NSCLC) target cohort shows why lower review can be m

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Sark: Oblivious Integrity Without Global State

arXiv:2512.20775v3 Announce Type: replace-cross Abstract: In this paper, we introduce Sark, a reference architecture for transferring unforgeable, stateful, oblivious (USO) assets. We describe the motivation, design, and implementation of the core subsystems of Sark, Porters, which accumulate and roll-up commitments from Clients, and Sloop, a permissioned, crash fault-tolerant (CFT) blockchain system. We analyse the operation of the system using STRIDE threat analysis, and the `CIA Triad': Confidentiality, Availability, and Integrity. We then introduce the concept of \textit{local centrality} and use it to address design trade-offs related to decentralization. Finally, we point to future work on Byzantine fault-tolerance (BFT), and mitigating the local centrality of Porters.

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Patient-centered data science: an integrative framework for evaluating and predicting clinical outcomes in the digital health era

arXiv:2408.02677v2 Announce Type: replace-cross Abstract: This study proposes a novel, integrative framework for patient-centered data science in the digital health era. We developed a multidimensional model that combines traditional clinical data with patient-reported outcomes, social determinants of health, and multi-omic data to create comprehensive digital patient representations. Our framework employs a multi-agent artificial intelligence approach, utilizing various machine learning techniques including large language models, to analyze complex, longitudinal datasets. The model aims to optimize multiple patient outcomes simultaneously while addressing biases and ensuring generalizability. We demonstrate how this framework can be implemented to create a learning healthcare system that continuously refines strategies for optimal patient care. This approach has the potential to significantly improve the translation of digital health innovations into real-world clinical benefits, addr

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Cost-of-Ethics Crisis: Beliefs, Decisions, and Justifications in the Job Searches of Computer Science Students in Canada and the United States

arXiv:2605.09680v2 Announce Type: replace Abstract: Workplace norms in computer science have received growing attention due to a series of recent ethical scandals. One response has been a push to improve the ethics education provided to computer science students. Evidence for the effectiveness of ethics education remains mixed; some evidence suggests that norms are changing, others point to persistent gaps between stated values and practice remain. In this paper, we explore whether students, who have received some contemporary CS ethics education, are able to effectively apply ethical reasoning to their own decision-making in what is typically the first significant ethical decision of their careers: the job search. Our study examines the ethical decision making of 129 computer science students and recent graduates during their job searches. We find that most students prioritize factors like compensation, location, and workplace culture over ethical and social issues. Even when expressi

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hierarchical Reinforcement Learning for Cooperative Air-Ground Delivery in Urban System

arXiv:2602.12913v2 Announce Type: replace Abstract: Cooperative air-ground delivery has emerged as a promising logistics paradigm by leveraging the complementary strengths of UAVs and ground carriers. However, effective dispatching in such heterogeneous systems faces two critical challenges: i) the heterogeneity between flight and road dynamics, ii) the scalability bottleneck raised by the exponential decision variables in large-scale fleets. To address these challenges, we propose HRL4AG, a Hierarchical Reinforcement Learning framework for cooperative Air-Ground delivery. Specifically, HRL4AG employs a high-level manager to tackle the scalability bottleneck by decomposing the joint action space, and mode-specific workers that encode distinct flight and road dynamics to address the heterogeneity. Furthermore, a novel internal reward mechanism is designed to guide the hierarchical policy learning, addressing the credit assignment problem in sparse-reward settings. Extensive experiments

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting

arXiv:2602.02882v2 Announce Type: replace Abstract: Large language models are increasingly used to predict human preferences in both scientific and business endeavors, yet current approaches rely exclusively on analyzing model outputs without considering the underlying mechanisms. Using election forecasting as a test case, we introduce mechanistic forecasting, a method that demonstrates that probing internal model representations offers a fundamentally different - and sometimes more effective - approach to preference prediction. Examining over 24 million configurations across 7 models, 6 national elections, multiple persona attributes, and prompt variations, we systematically analyze how demographic and ideological information activates latent party-encoding components within the respective models. We find that leveraging this internal knowledge via mechanistic forecasting (opposed to solely relying on surface-level predictions) can improve prediction accuracy. The effects vary across

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

arXiv:2601.17003v2 Announce Type: replace Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual diversity of deployment. We pair four benchmark replications with an ecological audit of real-world conversations to evaluate a purpose-built mental-health AI alongside six frontier general-purpose models spanning four families (OpenAI GPT-5, GPT-5.1, GPT-5.2; DeepSeek V3; Google Gemini 3 Flash; Moonshot Kimi K2). The purpose-built system produced significantly lower overall potentially harmful content rates than every frontier comparator on suicide/self-harm, eating-disorder, and substance-use prompts (CCDH Benchmark: Ash 6.2% vs frontier models 18.0-52.0%, all p < .001). In an audit of 20,000 deployment conversations, clinician review within the audit pipeline confirmed no suicide-risk conversations lacking crisis resources and three NSSI-related conversations without crisis intervention, a within

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

UNVaMP: Neural Knowledge Tracing with Variational Regularization of Latent Knowledge Dynamics

arXiv:2608.03811v1 Announce Type: cross Abstract: We introduce the Unified Neural Variational Measurement of Proficiency (UNVaMP) architecture, a knowledge tracing method that integrates observed student-item interactions with internal memory to produce evolving latent representations of student knowledge. These representations support accurate predictions of future responses while enabling explicit control over the smoothness of estimated learning trajectories. UNVaMP can be configured as either a purely neural model or a hybrid model that predicts responses through an interpretable measurement function over the latent space. We show that a pure neural configuration (UNVaMP-MLP) achieves the strongest predictive performance among compared models on three out of four datasets. Meanwhile, a hybrid configuration (UNVaMP-MIRT, using a 1PL MIRT measurement function) lags only slightly behind UNVaMP-MLP, indicating that the predictive cost of interpretability is modest. Beyond predictive ac

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

arXiv:2608.03700v1 Announce Type: cross Abstract: Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipeline, we introduce AntiSkillBench, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline. It comprises: (i) a dataset of 7,500 persona-grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies; and (iii) a defense evaluation covering four configurations across online and post-hoc interven

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Security-Oriented Lifecycle Model for Large Language Model Systems

arXiv:2608.03626v1 Announce Type: cross Abstract: Large language models are being integrated into critical infrastructure and enterprise workflows at unprecedented scale,yet the lifecycle frameworks governing their development and operations were designed for operational efficiency rather than security analysis. As a result, security-relevant activities such as data provenance verification, artifact signing, agentic permission control, and decommissioning are often left implicit or assumed to receive due care. Governance frameworks, in turn, organise requirements around risk levels or management processes without clearly linking them to the lifecycle stages where they apply. This paper addresses both deficiencies. We propose a lifecycle model for LLM systems that supports security analysis by structuring it around security-relevant boundaries rather than workflow optimisation. The model comprises 32 stages across four core pipeline layers (Data, Model, Distribution, Application), suppo

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities

arXiv:2608.03585v1 Announce Type: cross Abstract: Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowledge and interpersonal relationships through visible collaboration. Generative coding agents (CAs) are an advanced tool to improve development efficiency while shifting part of activities from public human interaction to private human-agent loops. We study this shift using an LLM-based multi-agent simulation initialized with real GitHub data from 1,084 active developers and their repository relationships. After a warm-up with historical commits, we branch the same community state into parallel No-CA and CA conditions for 4-week simulations. CA introduction increases planned and completed tasks by 34.0% and 39.0%, respectively, and reduces median completion time from 45 to 20 minutes. However, adoption reaches only 26.0%, and the gains concentrate among developers who are already more active and well co

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

arXiv:2608.03569v1 Announce Type: cross Abstract: Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in evaluating scientific reasoning and research replication, but often rely on synthetic tasks or retrospective targets, which may be confounded by prior exposure. We hypothesize that complex, adversarial, fast-moving real-world domains where expert practitioners independently generate observable outputs can provide a practical solution to fill this gap and evaluate the capabilities needed for AI scientists, including reasoning, novelty, and hypothesis formulation. We instantiate this framework in two structurally different domains, Formula 1 (F1), where models ideate around car design concepts for the 2026 season, and real pre-season innovations provide a ground truth, and Magic: The Gathering (MTG), where models propose decks from a recently updated card pool and are evaluated against 19 Pro

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Behaviorally Adaptive Visual Diversion for Inclusive and Resilient Digital Assessment Delivery

arXiv:2608.03531v1 Announce Type: cross Abstract: Institutions increasingly rely on browser lockdown, webcam monitoring, and behavioral analytics to secure high-stakes digital assessments, yet these mechanisms are commonly designed and evaluated independently and often overlook learner accessibility. This paper introduces Behaviorally-Adaptive Visual Diversion (BAVD), a theoretical framework in which a synthetic, non-semantic visual field is composited with assessment content and adaptively modulated according to observed candidate behavior. The underlying assessment content is never altered; only its visual presentation is modified to reduce the usefulness of unauthorized screen capture or screen sharing while remaining minimally intrusive for legitimate candidates. The framework further incorporates an accessibility-aware attenuation mechanism that reduces or suppresses diversion intensity for candidates with approved visual-processing accommodations. We formulate the model using a c

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Optimal Liability Design for Medical AI

arXiv:2608.03114v1 Announce Type: cross Abstract: Artificial intelligence (AI) is increasingly integrated into medical decision-making, yet its liability implications remain complex, particularly when physicians differ in diagnostic skills and their quality is unobservable. This paper develops a principal-agent model in which a social planner designs medical liability to regulate a physician with private quality information who chooses between a standard treatment, a personalized judgment-based treatment, or following an imperfect AI recommendation. Our analysis yields several novel insights. First, we show that the optimal mechanism under asymmetric information is surprisingly simple: a uniform, one-size-fits-all liability level for all physician types who deviate from the standard of care. Despite physician heterogeneity, this simple policy often achieves the full-information first-best outcome, particularly when standard care is reliable or AI is highly accurate. Second, the relatio

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Paired Recipient-based Evaluation of Survival Prediction for Deceased Donor Kidney Transplants

arXiv:2608.03017v1 Announce Type: cross Abstract: There has been significant interest in using machine learning algorithms to predict kidney transplant outcomes, such as the number of years until a graft inevitably fails. These prediction algorithms could possibly be used for pre-transplant donor-recipient matching to identify more compatible donors and recipients and thus improve post-transplant outcomes. In this study, we explore the use of survival prediction models trained on deceased donor kidney transplant data from the Scientific Registry of Transplant Recipients (SRTR). We propose a novel paired recipient-based evaluation framework that compares graft outcomes between two recipients who received kidneys from the same deceased donor, allowing us to evaluate the counterfactual benefit of changing the recipient for a certain donor. We find that five different survival prediction models, ranging in complexity from linear to deep learning-based models, all result in ~60% paired reci

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CY

Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits

arXiv:2608.02955v1 Announce Type: cross Abstract: This research paper describes an exploratory study on the effectiveness of Chat Debugging: troubleshooting malfunctioning analog circuits on breadboards and printed circuit boards (PCB) by undergraduates through conversations with public-domain large language models (LLMs). Through thematic analysis of students' voluntarily shared chat logs when debugging pre-determined buggy circuits under exam and time pressure, we discovered multimodal usage patterns by students and considerable domain knowledge and sensible debugging suggestions offered by off-the-shelf LLMs. Meanwhile, we also identified major gaps in LLM technologies and students' skills during human-AI collaborative debugging, such as LLMs' limitations in 2D/3D image-based reasoning, unjustified tone of confidence, and students' deficits in fundamental concepts and critical thinking.

Source ↗
Showing 201–250 of 1593 signals
← Prev Page 5 of 32 Next →