EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI

arXiv:2607.03510v1 Announce Type: cross Abstract: Enterprise artificial intelligence is moving from experimentation into operational workflows. Early programs focused on model access and retrieval-augmented generation, but enterprises are now beginning to deploy agents that plan, retrieve, remember, call tools, update systems, and coordinate work across applications. This changes the evaluation problem. Leaders are no longer asking only whether an answer is accurate or fluent. They need to know who authorized an action, which policy applied, whether evidence was current, whether memory was valid, whether a tool call was permitted, whether the decision can be replayed, and whether the agent can be stopped before it creates business impact. This paper introduces CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI. CAGE-1 is an evaluation framework for deciding whether enterprise agents are ready for deployment. It evaluates authority, policy enforcement, retri

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

From Mobile Data to Business Insights: An End-to-End Analytics Framework for Large-Scale Urban Mobility Analysis and Decision Support

arXiv:2607.03394v1 Announce Type: cross Abstract: Real time location data derived from mobile applications is a powerful tool for addressing various urban challenges, including tourism planning, parking management, bus route optimization, and resource allocation. Besides, it offers invaluable insights for shaping strategic decisions in commercial domains such as location based services, market share analysis, and behavioral profiling. In this expansive study, we aim to address all of the aforementioned challenges by investigating the behaviors and patterns of smartphone users within urban environments, particularly in the domains of tourism, transportation, and retail. Our approach encompasses the development of a sophisticated data platform from inception to implementation, which includes the formulation of use cases, architectural design, and implementation of modules. We employ state of the art techniques and technologies, including data anonymization, ETL pipelines, and utilizing G

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Silicon Sampling via Cross-Survey Transfer

arXiv:2607.03091v1 Announce Type: cross Abstract: Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B-120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable cons

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Comparative Study of Static, Scrollytelling, and Chatbot Visualization Onboarding Techniques for UX Designers

arXiv:2607.03023v1 Announce Type: cross Abstract: User experience (UX) designers face barriers when creating data visualizations due to limited domain expertise in visualization or unfamiliarity with specialized tools. This highlights a clear need for effective methods to build visualization literacy. To address this, we evaluated three visualization onboarding techniques -- static, scrollytelling, and chatbot -- in an experimental study with 25 UX designers and students. We measured visualization comprehension and guideline adherence during a visualization creation task, followed by surveys and interviews to capture preferences and experiences. Compared to static onboarding, the pooled interactive condition (scrollytelling or chatbot) was associated with significantly higher guideline-adherence scores during visualization creation; both interactive techniques also received higher engagement ratings. Instruction clarity ratings were significantly higher when the two interactive conditi

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

CAF\'E, an automated feedback tool to approach Formal Methods

arXiv:2607.03001v1 Announce Type: cross Abstract: We present CAF\'E, a learning platform designed to introduce computer science students to Formal Methods (FM). CAF\'E aims to scaffold students' structural thinking (in contrast with operational thinking) by promoting the practice of Graphical Loop Invariant Based Programming (GLIBP). In the GLIBP approach, students solve loop-based problems by first constructing a Graphical Loop Invariant (GLI) before deriving the corresponding code. The GLI is an informal diagrammatic representation of the loop invariant. It illustrates the variables involved in the loop, their properties, and the relationships between them. To enable automated feedback, students complete a blank GLI, a box-based version of the GLI. Beyond evaluating the code students submit, CAF\'E provides personalized feedback on students' GLI and its alignment with the code. In this demo, we walk through CAFE from both a student's and a teacher's perspective. We show how the tool

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Angry but Accurate: Detecting and Profiling the Counter-Misinformation Ecosystem on Twitter

arXiv:2607.02900v1 Announce Type: cross Abstract: On social media, many users actively push back against false claims. Understanding who pushes back and how they do so matters, as this corrective activity is central to how misinformation is contested. We study this counter-misinformation ecosystem at scale: applying a domain-specific NLI model from our prior work to a large corpus of COVID-19 tweets, we classify 264,737 posts as supporting or opposing false claims and compare 23 user- and text-level features across the two groups. Contrary to the dominant assumption that negative emotion is a signature of falsehood, we find that anti-misinformation posts are more emotionally negative than pro-misinformation posts, with higher levels of anger, disgust, and sadness. These differences are modest in magnitude but consistent in direction across the negative emotions. We also find that posts opposing misinformation tend to come from more established users, i.e., older accounts, more follower

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure

arXiv:2607.02814v1 Announce Type: cross Abstract: Personal agents will increasingly negotiate on behalf of users: splitting costs with other personal agents, appealing platform decisions, escalating support disputes, requesting refunds, changing subscriptions, and negotiating deadlines or reimbursements. Existing negotiation benchmarks emphasize agreement, surplus, or strategic competence, but a user-owned agent can reach an agreement while harming the user through privacy leakage, consent violation, unsupported advocacy, over-concession, failed escalation, or poor auditability. We introduce SovereignNegotiation-Bench, a trace-level multi-turn benchmark for delegated personal-agent negotiation under private utilities, disclosure constraints, evidence requirements, and institutional asymmetry. The benchmark separates agent-visible observable state from evaluator-only labels and evaluates agreement success jointly with user utility, privacy, consent, evidence grounding, concession discip

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Signal from Space: Detecting Schools and Towers to Bridge the Digital Divide

arXiv:2607.02724v1 Announce Type: cross Abstract: Reliable internet access is essential for modern education, yet millions of school-aged children especially in developing regions remain offline due to unconnected schools. The Giga Initiative aims to connect every school to the internet, but doing so at scale requires efficient methods to map schools and assess surrounding connectivity infrastructure without relying on sparse or noisy third-party datasets. In this work, we propose a scalable, vision-only framework that uses high-resolution satellite imagery and transfer learning to address both tasks simultaneously. By adapting pre-trained object detection models to new geographical regions with minimal labeled data, we detect schools and cell towers directly from space. We then analyze the spatial relationship between detected schools and nearby towers as a proxy for connectivity availability. This purely imagery-driven pipeline enables large-scale infrastructure mapping, reduces depe

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Doom Researching: A Conceptual Framework for Repetitive AI-Assisted Information Seeking, Cognitive Offloading, and the Illusion of Knowing

arXiv:2607.02723v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) systems such as ChatGPT, Claude, and Gemini have made information seeking faster, more conversational, and more cognitively comfortable. These affordances can support learning and productivity, but they can also encourage a repetitive pattern in which users continue querying AI systems for explanations, summaries, comparisons, plans, and reassurance without converting those interactions into durable understanding, decisions, or finished work. This conceptual paper proposes the term doom researching to describe this AI-mediated pattern of repetitive information seeking without proportional synthesis or output. Building on research on doomscrolling, information seeking, cognitive offloading, transactive memory, human-AI interaction, productivity loss, and the illusion of knowing, the paper develops a framework in which fluent AI responses reduce cognitive effort, inflate perceived knowledge, and

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Internal Pluralism and the Limits of Pairwise Comparisons

arXiv:2607.02672v1 Announce Type: cross Abstract: Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.g., in participatory design or alignment. However, their use builds in two strong assumptions: that local comparisons are sufficient evidence about how a person wants an automated decision rule to behave, and that people can always answer those comparisons decisively. We investigate how these assumptions may be compromised under internal pluralism: the idea that an individual evaluates decision rules according to multiple authoritative priorities about how the rule should behave. We provide a formal model of such pluralistic preferences over decision rules, which then lets us identify two distinct failures of forced local pairwise comparison data. First, priorities such as proportionality, egalitarianism, and equal treatment are inherently global: what they imply in one case can depend on what happens elsewhere, so local comparisons may

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Classroom Behavior Monitoring with YOLO An Empirical Study in Higher Education Settings

arXiv:2607.02580v1 Announce Type: cross Abstract: Classroom behavior monitoring plays a vital role in evaluating student engagement and improving teaching effectiveness. Traditional observation methods remain subjective and lack scalability. This study introduces a real-world dataset of classroom videos collected at the Banking Academy of Vietnam (BAV-Classroom dataset), annotated with nine distinctive behavioral categories. State-of-the-art Computer Vision models were evaluated and compared, with YOLOv11 achieving the best performance. Experimental results indicate that students' concentration often decreases notably during the final part of lectures, highlighting challenges in sustaining engagement. Our findings demonstrate the feasibility of applying computer vision for automated classroom monitoring, providing valuable insights for academic quality management.

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Curated retrieval versus open web search in public AI information services: a coverage-trust trade-off

arXiv:2607.05217v1 Announce Type: new Abstract: Public institutions increasingly use large language models (LLMs) to answer citizens' questions, often pairing a curated knowledge base with live web search, yet whether the sources behind these answers can be trusted has received little empirical scrutiny. We report a pre-launch expert evaluation of Evr\'opuvefur, an independent, government-funded service run by the University of Iceland that answers questions about the European Union, conducted as Iceland prepared for its referendum of 29 August 2026 on whether to resume EU accession talks. Five domain experts produced 551 evaluations of 449 AI-generated answers, scoring each against a seven-criterion quality rubric and, separately, flagging individual cited sources. We compared two retrieval paths: a curated local corpus (RAG) and open web search. In more than a third of the reviewed web-search answers (35%, 65 of 187), at least one cited source was flagged, almost always as untrustwor

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Open Problems in AI Incident Governance

arXiv:2607.05163v1 Announce Type: new Abstract: AI systems may produce failures after deployment that pre-deployment safety assessments do not anticipate. Managing these failures requires what we refer to as adequate \textit{AI incident governance}, where having good definitions, taxonomies, monitoring practices, reporting mechanisms, and incident analysis is essential. We examine existing frameworks related to AI incident governance by regulatory bodies and independent efforts, and find that while there are frameworks that describe how individual functions can be performed, there is a lack of consistency within the aspects of definitions, classification, monitoring, and reporting. These inconsistencies apply to the types of incident data that is collected and reported, the ways in which they are categorised, and as a result, the depth, representativeness, and accuracy of analysis that can be performed.

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games

arXiv:2607.05132v1 Announce Type: new Abstract: As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents that publicly commit to actions will honor those commitments. We place LLM agents in repeated $n$-player games with a three-stage protocol that separates private intent, public announcement, and final action, allowing us to identify whether each deviation from a stated announcement was already planned during private deliberation. Evaluating three frontier models across six games in homogeneous and heterogeneous groups over 10 rounds, we report two findings. First, when agents deviate from their announcements, the deviation is predominantly already stated in their private plan (exceeding 90% in the highest-deception conditions), yet this is not a fixed model property: the same model ranges from perfect honesty to near-total deviation across games. Second, different models interpret announcements

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Understanding Student Perceptions, Mistakes, and Debugging Approaches when Solving Natural Language Programming Tasks

arXiv:2607.05034v1 Announce Type: new Abstract: Learning to communicate with code-generating AI models is an emerging skill for novice programmers. One recent pedagogical approach, Prompt Problems, has students solve computational tasks by writing natural-language prompts for code-generating AI models. However, little is known about the specific prompt-level mistakes novice programmers make, the kinds of computational details they fail to communicate, and what strategies they use to recover when generated code is incorrect. In a CS1 course, we studied attempts by more than 900 students to solve dialogue-based Prompt Problems. We analyzed student reflections, unsuccessful prompts, and reported debugging strategies. Compared to traditional coding tasks, students generally found prompting easier, more enjoyable, and better targeted at developing problem-solving skills. The most common mistakes are related to the omission of key details, suggesting both a failure to acknowledge their impor

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Psychological features of dispute content and public acceptance of AI in legal adjudication: evidence for systematic variation beyond individual differences

arXiv:2607.04838v1 Announce Type: new Abstract: Public acceptance of artificial intelligence (AI) in legal decision-making has been primarily explained through individual differences in personality traits and general technology attitudes. However, contextual features of legal disputes themselves may systematically influence preferences for AI versus human adjudicators. Across two studies with Japanese participants (N = 1,384 and N = 596), we examined whether psychological characteristics of dispute content shape acceptability judgments for algorithmic adjudication. Study 1 employed exploratory factor analysis on acceptability ratings across 46 legal dispute vignettes, revealing a two-dimensional structure distinguishing interpersonal-relational disputes (where human adjudicators were strongly preferred) from institutional-procedural disputes (where AI acceptance was comparatively higher). Study 2 replicated this structure in an independent sample and demonstrated that experimentally ma

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Double-edged Effect of Banning Generative AI on Online Question-and-Answer Communities: Evidence from Stack Exchange

arXiv:2607.04601v1 Announce Type: new Abstract: We investigate how banning generative artificial intelligence-generated content (AIGC) affects knowledge seeking, knowledge contribution, and contribution efficiency in online question-and-answer communities. After the launch of ChatGPT in late November 2022, several Stack Exchange communities implemented official bans on AIGC over concerns such as less reliable and socially engaged content. Leveraging data from the full network of Stack Exchange communities, we employ a difference-in-differences (DID) approach to examine the impacts of these bans. Our results reveal a double-edged impact: while the AIGC ban increases knowledge seeking, as evidenced by a higher volume of posted questions, it simultaneously reduces contribution efficiency, reflected in a lower proportion of questions receiving satisfactory answers within the expected time frame. Notably, these impacts are only evident in non-STEM communities. We take a socio-technical pers

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Government AI Use as a Monitoring Primitive: A Public Document Pilot Study

arXiv:2607.04543v1 Announce Type: new Abstract: Governments are important actors in frontier AI governance, but many facts about their adoption and use of AI systems are difficult to observe directly. Procurement disclosures and official statements are useful, but can also be delayed, selective, and better suited to measuring formal adoption than actual day-to-day use. We propose a complementary monitoring primitive: measuring traces of language-model assistance in public government documents. The approach is lightweight, externally reproducible, and based on revealed behavior rather than stated intent. In a pilot study of ten public document streams from U.S. and PRC government-related sources, we find that, while 2021 baselines are consistently near zero, by 2026, four of our ten sources show statistically significant signs of AI-assisted writing. In our sample, the U.S. signal concentrates in publications downstream of policy work; the PRC signal concentrates closer to it. We close

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Hybrid Algorithmic Governance in U.S. Welfare Administration: State- and County-Level AI as a Case of Support-Control Convergence

arXiv:2607.04503v1 Announce Type: new Abstract: This article examines the institutional conditions under which artificial intelligence systems in U.S. welfare administration come to operate as instruments of support or as instruments of control. Rather than asking what welfare algorithms "really" are (tools of proactive assistance or infrastructures of surveillance) the article starts from the premise that support and control are co-present within the same system, while their relative balance shifts over time. This movement is conceptualized through the notion of support-control convergence and the model of an institutional ratchet. Routine budgetary and political pressures make control-oriented effects easily measurable and politically capitalizable, whereas a return toward support requires external intervention of disproportionate force, such as judicial compulsion, legislative prohibition, or public scandal. Empirically, the article draws on process tracing of six state- and county-

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Case for Globally Beneficial Technology

arXiv:2607.03906v1 Announce Type: new Abstract: To whom do the fruits of advanced technological innovation belong? To their inventors, to the organizations and individuals involved in making such discoveries possible, or to still larger groups of people, potentially encompassing all of humanity? This question sits at the heart of the present investigation. The arguments developed here focus on an expansive reading of the entitlement to benefit from technological breakthroughs: we argue that they should be designed, developed, and distributed in ways that benefit everyone. This central claim, which encompasses technologies such as advanced forms of artificial intelligence, is grounded in an exploration of five moral arguments that involve human rights, beneficence, contingencies of birth, the global tree of knowledge, and global economic justice. Taken together, they underpin the argument for globally beneficial technologies.

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Content Hidden Behind Execution: Analyzing Public Scratch Projects at Runtime

arXiv:2607.03700v1 Announce Type: new Abstract: Public Scratch projects are reused in computing education as classroom examples, remix sources, open-exploration materials, and research data. Curation often begins with titles, thumbnails, descriptions, tags, and remix links, but Scratch projects are executable learning artifacts. Content affecting age appropriateness can appear only after execution, gameplay progression, a failure state, user interaction, costume switching, audio playback, or a hidden event trigger. We study "runtime-revealed sensitive content" as a computing education curation challenge: educators and researchers need runtime evidence about what students may encounter when Scratch projects are used in these settings. We introduce a runtime-aware annotation scheme that separates content category, risk level, evidence channel, reveal mechanism, and annotation confidence. Using this scheme, we conducted an audit of 500 public Scratch projects sampled from curated candidat

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Development of a Bio-Inspired Routing Algorithm According to Values of Solidarity and a Freirean Perspective of Engineering

arXiv:2607.03607v1 Announce Type: new Abstract: A routing algorithm for Se\~noritas Courier, a bicycle delivery cooperative in S\~ao Paulo, Brazil, composed exclusively of cis women and trans people, is presented in this paper. Unlike conventional logistics optimization, which typically focuses on cost or distance minimization, this cooperative operates under principles of solidarity, care, and equitable income distribution. The algorithm was developed through a participatory process involving cooperative members as co-designers. The classical Vehicle Routing Problem proved inadequate for this context, as it disregards individual constraints and fairness. We formulate a new variant, the Se\~noritas Routing Problem, which incorporates biker-specific constraints on weight, volume, and maximum distance, alongside a solidarity objective that balances route lengths. A genetic algorithm is employed as the solution method. Three fitness formulations are compared: a baseline distance-minimizat

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI

arXiv:2607.03542v1 Announce Type: new Abstract: Frontier-AI governance today faces a problem structurally analogous to the one banking regulation faced pre-2008, and which post-2008 reforms (Basel III, Dodd-Frank) have since addressed. Two gaps recur: discovering a risk is not tantamount to acting on it, and individual-model review is unlike managing correlated build-up across the sector. Drawing on the Basel III framework and the U.S. financial-stability architecture, I propose a macro-prudential early warning and response system ("MEWRS") for internal frontier AI. These are systems deployed for labs' own internal research, testing, and production workflows, as distinct from externally released products. Layer A adapts the finder-coordinator-defender early-warning model to route structured reports on dual-use capabilities, autonomy indicators, and security compromises through a government clearinghouse to domain-specific defender working groups. Layer B calibrates operational controls

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI Systems as Digital Public Goods -- Evidence and Recommendations from a Multi-Stakeholder Assessment

arXiv:2607.03427v1 Announce Type: new Abstract: AI systems are increasingly being positioned as potential Digital Public Goods (DPGs) to accelerate progress towards the Sustainable Development Goals (SDGs). Yet, despite major global commitments, most notably the Global Digital Compact's call to "develop, disseminate and maintain safe and secure open-source software, open data, open artificial intelligence models and open standards that benefit society as a whole", very few AI systems currently meet the DPG Standard in practice. This report explains why, and what must change for "AI as Digital Public Goods" (AIDPGs) to become a credible, implementable pathway rather than an aspirational label. Commissioned by the Asian Development Bank (ADB) and produced by United Nations University (UNU) in partnership with UN Office of Digital and Emergent Technologies (UN ODET), this assessment combines: (i) a structured desk review of policy, legal, and technical frameworks on DPGs, openness, and AI

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Analyzing the Difficulty of Programming Assignments with Interpretable Knowledge Component Metrics

arXiv:2607.03419v1 Announce Type: new Abstract: This research paper examines how Knowledge Components (KCs) - fine-grained concepts or skills required to solve programming tasks - can be used as interpretable signals for understanding assignment difficulty and student struggle in introductory programming courses. While prior work has focused on predictive models based on programming behavior, such models are often difficult to interpret and therefore hard to use for instructional decisions. We analyze KC-based metrics, including the number of KCs per assignment and changes in KC coverage between consecutive assignments. We examine correlations between the number of KCs and student performance on the assignment, and analyze changes in KCs across assignments to identify cases where performance declines without new concepts being introduced. Selected assignments are then qualitatively inspected to understand potential design issues. Our results on data from three introductory programming

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Scalable Approach to Evaluating Moral Sensitivity in LLMs

arXiv:2607.02972v1 Announce Type: new Abstract: Moral sensitivity is the ability to identify the morally relevant features of a decision situation and use them as the basis for action. It is the foundation of broader moral competence: any other moral reasoning capabilities will be irrelevant if an agent lacks sensitivity to the relevant facts. In this paper, we offer a new evaluation of LLM moral sensitivity and in doing so, we address and resolve a central problem in AI alignment research: how to scale behavioural evaluations beyond expensive and sometimes metaethically dubious comparisons with a human baseline, without adopting an LLM judge that must be assumed to have the very capability that you are attempting to evaluate. Our central question is this: can LLMs successfully identify the morally relevant features of noisy cases, in which various kinds of morally irrelevant information have been introduced to distract the respondent? To explore this, we introduce \textbf{MORPH-1K (MO

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Foreign Policy AI Evaluation Gap

arXiv:2607.02955v1 Announce Type: new Abstract: We argue that AI systems used in conducting foreign policy tasks - broadly enacting 'statecraft' - should be a priority test case for technical AI governance research. In enacting foreign policy, we refer to the formulation and implementation of external objectives by political actors. Statecraft is a high-consequence deployment domain, with extreme downside risks and structural properties that standard evaluation practices handle poorly. These features include partial observability, unbounded action spaces, contested ground truth, and multidimensional objectives. This paper advocates for a literature-grounded research agenda. Our contribution is threefold: (i) a claim about the structural conditions of foreign policy that combine catastrophic tail risk with technical evaluation complexities, (ii) an ECOSYSTEM review that highlights the asymmetric focus on ASSESSMENT features over ACCESS, VERIFICATION, SECURITY, and OPERATIONALIZATION, an

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Hidden Water Geography of U.S. Hyperscale Data Centers in the AI Era

arXiv:2607.02531v1 Announce Type: new Abstract: Water use by data centers is routinely reported as a single footprint, but water is consumed through two physically distinct pathways: at the site for cooling and in the power system that generates electricity. We mapped both pathways for 472 U.S. hyperscale facilities by linking facility locations to electricity regions, hydrologic basins, and water-stress data. Under baseline assumptions, operational water consumption totals approximately 300 GL yr^-1 (range 205-451 across scenarios), with electricity-related water contributing three-quarters of the total. The two pathways produce different hotspot geographies: direct cooling burdens concentrate in stressed western and south-central basins, whereas electricity-related burdens concentrate in a few eastern grid regions with fossil-heavy supply. Just 3 of 24 hosting balancing authorities account for 59% of electricity-related water. Separating pathways identifies which decisions matter whe

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

Cybercrime Victimization Among Young Adult Males Aged 18--20: A Post-Pandemic Analysis of Converging Risk Factors

arXiv:2607.02530v1 Announce Type: new Abstract: Cybercrime victimization among young adult males aged 18--20 has become an increasingly urgent public safety concern in the post-pandemic digital environment. From 2022 to 2024, individuals aged 20--29 submitted 191,787 complaints to the FBI Internet Crime Complaint Center (IC3), reporting combined losses of more than $1.28 billion. Although this population represents a substantial share of cybercrime victims, the 18--20 male sub-cohort remains insufficiently examined as a distinct demographic group within cybercrime victimization research. This study presents an original risk factor analysis and theoretical synthesis, representing the first integration of criminological, neurological, and behavioral evidence for this specific demographic sub-cohort. Drawing on FBI IC3 and FTC Consumer Sentinel Network data from 2022--2024 alongside European cybersecurity threat intelligence from ENISA, the study develops a unified risk profile centered o

Source ↗
technology Tue, 07 Jul 2026 00:00:00 -0400
arXiv cs.CY

AutoResearch: An Execution-Grounded Multi-Agent Framework for Reliable Research Workflow Automation

arXiv:2607.02520v1 Announce Type: new Abstract: Automated research agents increasingly generate code, retrieve literature, and draft scientific artifacts, but they often fail to verify whether generated experiments execute correctly or whether cited sources support generated claims. We present AutoResearch, an execution-grounded multi-agent framework for reliable research workflow automation. AutoResearch couples sandboxed Python/PyTorch execution, iterative code repair, citation verification, claim-support auditing, decision control, and structured \LaTeX{} artifact generation. The system treats runtime errors, citation-verification failures, and review-agent feedback as practical filtering signals for generated research artifacts. In controlled evaluations on HumanEval, MBPP, a SciCode subset, citation-validation tasks, claim-support auditing, and small end-to-end workflow stress tests, AutoResearch improves execution success, citation validity, local claim support, and workflow comp

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent Community

arXiv:2602.02613v4 Announce Type: replace-cross Abstract: The rapid emergence of autonomous large language model agents has given rise to persistent, large-scale agent ecosystems whose collective behavior cannot be adequately understood through anecdotal observation or small-scale simulation. This paper introduces data-driven silicon sociology as a systematic empirical framework for studying social structure formation among interacting artificial agents. We present a pioneering large-scale data mining investigation of an in-the-wild agent society by analyzing Moltbook, a social platform designed primarily for agent-to-agent interaction. At the time of study, Moltbook hosted over 150,000 registered autonomous agents operating across thousands of agent-created sub-communities. Using programmatic and non-intrusive data acquisition, we collected and analyzed the textual descriptions of 12,758 submolts, which represent proactive sub-community partitioning activities within the ecosystem. Tr

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Interpretable Recognition of Cognitive Distortions in Natural Language Texts

arXiv:2511.05969v2 Announce Type: replace-cross Abstract: We propose a new approach to multi-factor classification of natural language texts based on weighted structured patterns such as N-grams, taking into account the heterarchical relationships between them, applied to solve such a socially impactful problem as the automation of detection of specific cognitive distortions in psychological care, relying on an interpretable, robust and transparent artificial intelligence model. The proposed recognition and learning algorithms improve the current state of the art in this field. The improvement is tested on two publicly available datasets, with significant improvements over literature-known F1 scores for the task, with optimal hyper-parameters determined, having code and models available for future use by the community.

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems

arXiv:2506.17467v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown significant potential to change how we write, communicate, and create, leading to rapid adoption across society. This dissertation examines how individuals and institutions are adapting to and engaging with this emerging technology through three research directions. First, I demonstrate how the institutional adoption of AI detectors introduces systematic biases, particularly disadvantaging writers of non-dominant language varieties, highlighting critical equity concerns in AI governance. Second, I present novel population-level algorithmic approaches that measure the increasing adoption of LLMs across writing domains, revealing consistent patterns of AI-assisted content in academic peer reviews, scientific publications, consumer complaints, corporate communications, job postings, and international organization press releases. Finally, I investigate LLMs' capability to provide feedback on r

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Right Kind of Help: Evaluating the Effectiveness of Feedback Methods in Elementary-Level Visual Programming

arXiv:2512.11735v2 Announce Type: replace Abstract: We present a large-scale study comparing the effectiveness of various feedback methods in elementary-level programming. While prior work has explored different feedback methods, their relative impact during the learning and post-learning phases remains unclear. In this study, we compare three feedback methods: code-edit recommendations (Code-Rec), quizzes based on code edits (Code-Quiz), and quizzes based on metacognitive strategies (Plan-Quiz), along with a no-feedback control (None). A total of 398 students (across grades 4-7) participated in a two-phase study: a learning phase comprising write-code tasks from the Hour of Code: Maze Challenge with feedback, followed by a post-learning phase comprising more advanced write-code tasks without feedback. All feedback methods significantly improved learning performance over the control, while preserving students' problem-solving skills in the post-learning phase. Furthermore, quiz-based m

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Reflection-Satisfaction Tradeoff: Investigating Impact of Reflection on Student Engagement with AI-Generated Programming Hints

arXiv:2512.04630v2 Announce Type: replace Abstract: Generative AI tools, such as AI-generated hints, are increasingly integrated into programming education to offer timely, personalized support. However, little is known about how to effectively leverage these hints while ensuring autonomous and meaningful learning. One promising approach involves pairing AI-generated hints with reflection prompts, asking students to review and analyze their learning, when they request hints. This study investigates the interplay between AI-generated hints and different designs of reflection prompts in an online introductory programming course. We conducted a two-trial field experiment. In Trial 1, students were randomly assigned to receive prompts either before or after receiving hints, or no prompt at all. Each prompt also targeted one of three SRL phases: planning, monitoring, and evaluation. In Trial 2, we examined two types of prompt guidance: directed (offering more explicit and structured guidanc

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Economics of AI Training Data: A Research Agenda

arXiv:2510.24990v3 Announce Type: replace Abstract: Despite data's central role in AI production, it remains the least understood input. As AI labs exhaust public data and turn to proprietary sources, with deals reaching hundreds of millions of dollars, research across computer science, economics, law, and policy has fragmented. We establish data economics as a coherent field through three contributions. First, we characterize data's distinctive properties -- nonrivalry, context dependence, and emergent rivalry through contamination -- and trace historical precedents for market formation in commodities such as oil and grain. Second, we present systematic documentation of AI training data deals from 2020 to 2025, revealing persistent market fragmentation, five distinct pricing mechanisms (from per-unit licensing to commissioning), and that most deals exclude original creators from compensation. Third, we propose a formal hierarchy of exchangeable data units (token, record, dataset, corp

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

WhatsApp as an improvisation of health information systems in Southern African public hospitals: A socio-technical perspective

arXiv:2502.09049v2 Announce Type: replace Abstract: Digital health interventions, particularly electronic referrals (e-referrals) and health information systems, have revolutionised clinical workflows in public hospitals by automating processes. However, the utilization of e-referrals has yielded mixed outcomes, with varying levels of success in organisational processes.This paper explores improvisation of health information systems in Southern African public hospitals from a socio-technical perspective. In particular the paper explains the design-reality gaps giving rise to improvisations of mandated health information systems in order to understand their occurrence and impact on referral outcomes. We employed the design-reality framework and the Process framework for Healthcare Information System Workarounds and Impacts to explain the socio-technical issues related to the phenomenon of interest.We conducted semi-interviews with 31 respondents from health organisations as case studies

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

ApplE: A Modular Ontology of Applied Ethics and Event Context for Ethical Decision Modeling

arXiv:2502.05110v2 Announce Type: replace Abstract: Applied ethics applies ethical decision-making to domain-specific contexts using contextual information such as agents, actions, temporal and spatial settings, and theoretical constructs such as utility, virtues, rights, and duties. However, representing an ethical decision is challenging as it may be abstract, context-sensitive, and semantically heterogeneous. Nevertheless, important ethical and contextual factors can be formally modeled to support structured ethical reasoning. Knowledge representation and reasoning provide a mechanism to translate abstract ethical concepts into machine-interpretable conceptual structures in the context of an event. To achieve this, we propose ApplE, an Applied Ethics ontology that models ethical theory and event context within a unified and modular conceptual framework for ethical decision-making. The ontology was developed using a modified version of the Simplified Agile Methodology for Ontology De

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation

arXiv:2608.02518v1 Announce Type: cross Abstract: The most capable AI deployments are not single models but ensembles of specialized agents that delegate and act in coordination. This architecture unlocks powerful new capabilities, and it also introduces risks that existing frameworks for monitoring, detection, and mitigation were not designed to address. Most state-of-the-art AI abuse detection literature focuses on single-turn or multi-turn (single-session) threat models. This leaves a critical gap: an attacker can decompose a harmful goal into innocuous-looking units and execute each in isolated agentic sessions. The agent is stateless between conversations, but the attacker is not. This asymmetry allows for cross-session trajectories that are effective at evading detection. Our contributions are twofold. First, we demonstrate cross-session goal decomposition as an evasion technique, showing it may elicit more harmful capability than equivalent single-session or multi-turn attacks.

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMs

arXiv:2608.02486v1 Announce Type: cross Abstract: Open-source LLMs reliably name Zeus, Jupiter, and Thor, but recover their counterparts in less-represented traditions like Finnish, Slavic, Egyptian, or Chinese mythology far less consistently. We ask where inside the model this cultural default is produced. On a parallel cross-cultural substrate of Thompson-motif entities, we instrument 18 open-source LLMs from 8 architecture families with linear probing, logit lens, activation patching, and output extraction. The residual stream cleanly distinguishes cultures, well above a name-string baseline, yet the decoder collapses culturally-specific tokens onto dominant-tradition ones. The failure is at readout, not at representation. Asking the same question in the target culture's native language versus English produces failures that cluster within language but decouple across language: the decoder is gated on prompt language. We release a per-entity (probe, output) decomposition framework, a

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

WIP: Chat-Debugging: Large Language Model as a Hardware Debugging Assistant

arXiv:2608.02420v1 Announce Type: cross Abstract: This work-in-progress research paper explores Chat-Debugging, a novel use case for large language models as an assistant for hardware debugging tasks to improve students' debugging skills. Hardware debugging can be a time-consuming and stressful skill to develop, leading to frustration and other negative emotions. While past work has explored streamlining and automating software-based circuit debugging where digital circuits are dominant, Chat-Debugging aids in physical hardware debugging where circuits may be analog, digital, or mixed-signal. Qualitative data were collected from LLM chat logs and interviews with a fourth-year electrical engineering undergraduate student. Major themes were extracted using a constant comparative analysis. Chat-Debugging incorporates accurate hardware information, properly handles natural language descriptions of circuits, and improves debugging confidence. A successful Chat-Debugging session includes inv

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models

arXiv:2608.02409v1 Announce Type: cross Abstract: Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms record satisfaction without task purpose. Together, they leave a critical gap in LLM evaluation: no existing infrastructure routinely links interaction trajectories to user-defined outcomes. We introduce MonitrLLM, open-source infrastructure for community-centered LLM evaluations that links full conversation transcripts to user-reported task intent and outcome assessments, treating all three as primary evaluative signals rather than optional metadata. To demonstrate the value of this approach, we conducted a two-week feasibility pilot with 26 college students using ChatGPT, collecting 206 evaluation reports with full conversation transcripts. The findings from our pilot demonstrate the value of connecting conversation trajectories with user-reported outcomes.

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Profitability of Open-Source Software Product Development

arXiv:2608.02398v1 Announce Type: cross Abstract: Many technology firms now build software products in the open, inviting outside developers to contribute alongside their employees on platforms like GitHub. Does this openness in product development pay off? Analyzing 977 U.S. high-tech firms from 2001 to 2025, this study finds that open-source adoption raised firms' gross margins by 4-5% on average. These gains flow largely through higher labor productivity, as firms integrate external contributors' diverse knowledge into internal workflows, broadening the organizational knowledge base without a commensurate rise in labor costs. However, the payoff emerges only when outside volunteers supply a meaningful share of the work (around 35% in this sample), and hinges on the firm's resource configuration. While there are multiple pathways to profitability, pairing open-source product development with sustained internal R&D is a core condition present in all high-profitability configurations.

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

TrainShield: Targeted Awareness for Cybersecurity Training

arXiv:2608.02296v1 Announce Type: cross Abstract: In recent years, cybersecurity threats have increasingly exploited human behaviour rather than purely technical vulnerabilities, exposing the limits of traditional awareness programmes delivered outside real-world contexts. To bridge this gap, we introduce TrainShield, an interaction paradigm for contextual cybersecurity training that embeds adaptive learning interventions directly within user workflows. The system integrates real-time risk detection (e.g., phishing and data loss prevention) with event-triggered hypermedia overlays that dynamically connect users to context-specific learning nodes embedded within their browsing workflow to deliver personalised micro-learning content and structured feedback tailored to the user's knowledge level and current context. This approach operationalises behavioural theories by transforming security incidents into immediate learning opportunities, shifting users from automatic to reflective decisi

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers

arXiv:2608.02024v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used across diverse tasks in K-12 education, yet existing safety evaluations rarely examine how harmful or inappropriate content appears in interactions between LLMs and students or teachers. To address this, we present EduZone, an evaluation framework for LLM safety across diverse educational scenarios. Our framework systematically combines (1) student- and teacher-facing LLM usage contexts, (2) fine-grained curriculum concepts, and (3) 6 risk categories and 28 subcategories spanning both conventional and education-specific harms to generate contextually grounded adversarial interactions. We construct these interactions in three settings: single-turn requests, static multi-turn conversations, and dynamic multi-turn conversations. Using these interactions, we evaluate ten LLMs using four safety levels: refusal, safe assistance, risky assistance with safety guidance, and fully risky assistanc

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Contractualist Argumentation Framework for Moral Decision-Making

arXiv:2608.01937v1 Announce Type: cross Abstract: Autonomous agents operating in shared environments must make decisions that affect multiple individuals with potentially conflicting interests. We propose a formal framework for moral decision-making grounded in Scanlon's contractualism, an ethical theory that evaluates the permissibility of actions in terms of principles that no one could reasonably reject. To operationalise contractualist reasoning, we use ASPIC+, a structured argumentation framework, extended with value-based filtering to model how each agent's values determine which reasons are morally relevant in the first place. The result is a Contractualist Argumentation Framework in which agents' reasons are formally represented, compared, and evaluated through argumentation semantics. We illustrate the approach through a worked example in a domestic setting and discuss its relation to existing value-based argumentation approaches.

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Do people rely on ChatGPT more than their peers to detect deepfake news?

arXiv:2608.01540v1 Announce Type: cross Abstract: This experimental study investigates how people rely on different sources of advice when detecting AI-generated fake news (deepfake news). In a laboratory deepfake detection task, student participants identified the proportion of human-written (non-AI-generated) content in synthetic deepfake news articles and received advice from ChatGPT (GPT-4), human peers, or linguistic experts. The results show that participants rely more on ChatGPT than on human peers when detecting GPT-2-generated deepfake news. Participants also rely more on linguistic experts than on peers, while the relative reliance on experts versus ChatGPT is mixed across experimental waves, potentially reflecting time trends in beliefs about AI-based detection. Importantly, in the additional experiment conducted in 2025 under the same experimental procedure, participants relied more on linguistic experts than on ChatGPT. Moreover, performance improvements reflect the joint

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Can Language Models Identify Shadow Trading Targets? An NLP Evaluation of SEC Enforcement Theory

arXiv:2608.01322v1 Announce Type: cross Abstract: Shadow trading -- trading in a peer firm's securities on the basis of material nonpublic information (MNPI) about an "economically linked" company -- is a novel and contested theory of insider trading liability, first prosecuted in SEC v. Panuwat (2023). Enforcing it requires identifying economically linked firms ex ante, a determination the SEC makes only after the fact using mass market surveillance infrastructure. We ask whether NLP can do what the SEC's theory presumes insiders already know: identify peer firms ex ante from publicly mandated disclosures. Using a two-stage LLM pipeline applied to Item 7 (Management's Discussion and Analysis) sections of SEC 10-K filings, we score semantic similarity across 30 M&A events spanning five industries and relate similarity to announcement-day abnormal stock returns. On the Panuwat fact pattern itself the pipeline recovers Incyte among the closest peers, a sanity check on the one case with a

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

arXiv:2608.01193v1 Announce Type: cross Abstract: An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safely, or move faster while taking a risk that may remove its final reward. We use this repeated game to study strategic safety behaviour among large language model (LLM) agents in races with two to five players. However, a valid action does not show that an agent understands the game. We therefore place an audit gate before behavioural interpretation. We first verify the game engine, then test rule recall, state tracking, payoff calculation, and stability under different but equivalent task descriptions. We then compare LLM action sequences with an evolutionary game-theory benchmark and published human data, and explore differences across models, risk conditions, personas, and two- to five-player races. The audit shows that strong rule recall can coexist with weak state tracking and expected-payoff calculation. Providing verified arithmeti

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

EmergencyBias: Bias in Text-to-Image Models under Emergency Scenarios

arXiv:2608.00598v1 Announce Type: cross Abstract: Bias in Text-to-Image (T2I) generation has become an important problem in multimedia content creation and communication. However, existing studies have primarily focused on relatively static and explicit forms of bias, such as disparities in the representation of gender, race, and geo-cultural attributes. Less attention has been paid to behavioral bias in how different groups are portrayed acting, reacting, and occupying social roles. Emergency scenarios provide a revealing setting for studying such bias because they require models to depict not only who is present, but also who is at risk, who intervenes, and how responsibility is allocated. In this paper, we define EmergencyBias, a form of bias in T2I generation under emergency scenarios that includes both demographic bias and behavioral bias. We construct an evaluation framework to systematically study EmergencyBias across seven leading T2I models, six representative emergency scenar

Source ↗
Showing 751–800 of 1593 signals
← Prev Page 16 of 32 Next →