EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Mon, 31 Aug 2026 00:00:00 -0400
arXiv cs.CY

Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

arXiv:2507.20409v3 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must perceive, understand, and judge all at once; bridging perception with norm-grounded reasoning. Recent work has introduced structured reasoning for multi-turn agent planning and visual QA, decomposing tasks into sequential sub-goals. To extend this to single-shot multimodal social reasoning, we introduce Cognitive Chain-of-Thought (CoCoT), a reasoning framework that structures vision-language-model (VLM) reasoning through three cognitively inspired stages: Perception (extract grounded facts), Situation (infer situations), and Norm (applying social norms). Evaluation across multiple distinct tasks such as multimodal intent disambiguation, multimodal theory of mind, social commonsense reasoning, and safety instruction following, shows consistent improvements (5.9% to 4.6% on average). We f

Source ↗
technology Mon, 31 Aug 2026 00:00:00 -0400
arXiv cs.CY

Comparing Apples to Oranges: A Taxonomy for Navigating the Global Landscape of AI Regulation

arXiv:2505.13673v2 Announce Type: replace Abstract: AI governance has transitioned from soft law, such as national AI strategies and voluntary guidelines, to binding regulation at an unprecedented pace. This evolution has produced a complex legislative landscape: blurred definitions of "AI regulation" mislead the public and create a false sense of safety; divergent regulatory frameworks risk fragmenting international cooperation; and uneven access to key information heightens the danger of regulatory capture. Clarifying the scope and substance of AI regulation is vital to uphold democratic rights and align international AI efforts. We present a taxonomy to map the global landscape of AI regulation. Our framework targets essential metrics-technology or application-focused rules, horizontal or sectoral regulatory coverage, ex ante or ex post interventions, maturity of the digital legal landscape, enforcement mechanisms, and level of stakeholder participation-to classify the breadth and d

Source ↗
technology Mon, 31 Aug 2026 00:00:00 -0400
arXiv cs.CY

Challenging Techno-Solutionism: The Role of ICT Innovation and the Value of Technological Growth

arXiv:2309.12355v2 Announce Type: replace Abstract: Innovation in Information and Communication Technology (ICT) has become one of the key economic drivers of our technology-dependent world. Digital devices/systems have become so pervasive that it is hard to imagine new technology developments that are not totally or partially influenced by ICT innovations. Furthermore, the pace of innovation in ICT sector over the last few decades has been unprecedented in human history. In this paper, we argue that the ICT innovation paradigm has crucially shaped collective expectations and imagination about what technology more broadly can actually deliver, particularly for a more sustainable and equitable world. These expectations have often crystalised into a widespread acceptance, among general public and policy makers, of techno-solutionism. We emphasise the role of electronic microchips in deriving relentless innovation and its impacts. We identify the many impacts of this innovation cycle into

Source ↗
technology Mon, 31 Aug 2026 00:00:00 -0400
arXiv cs.CY

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

arXiv:2608.28405v1 Announce Type: cross Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We introduce CultureConverse, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains. Each simulated and evaluated episode produces a scored interaction where the assistant assists the user and infers cultural constraints from partial information. The resulting CultureConverse-DS dataset contains 14,610 benchmark (evaluation) episodes and 274,295 oracle-guided (gold-mode) dialogues. In our benchmark evaluation of 18 models, GPT-5 mini achieves the highest assistance quality. Human annotation experiments suggest that our evaluation framework is a sufficient proxy for

Source ↗
technology Mon, 31 Aug 2026 00:00:00 -0400
arXiv cs.CY

Generative AI Alignment with Hinduism's Theological Plurality and Sacred Representation

arXiv:2608.28228v1 Announce Type: cross Abstract: Generative AI systems are increasingly used to answer personal questions and mediate everyday practices, including religion. However, existing discussions around AI alignment and ethics have largely centered secular, Western, and Abrahamic assumptions about religion, offering limited attention to other faith-based traditions. In this paper, we examine how Hindu users engage with generative AI systems in relation to their religious knowledge, belief, and practice. Drawing on 15 semi-structured interviews with Bangladeshi Hindu participants, we analyze how users interpret AI-generated religious representations, scriptural explanations, devotional interactions, and synthetic religious media. We found that AI can be both accessible and ethically troubling. While AI supported scriptural inquiry, devotional visualization, and religious storytelling, our study also identified concerns about theological flattening, cultural misrepresentation, d

Source ↗
technology Mon, 31 Aug 2026 00:00:00 -0400
arXiv cs.CY

Measuring the Installed Base: Nordic Health Dataset Catalogues Against HealthDCAT-AP Release 7

arXiv:2608.27720v1 Announce Type: cross Abstract: The European Health Data Space requires member states to publish machine readable descriptions of the health datasets available for secondary use, and the European Commission publishes HealthDCAT-AP as the metadata profile those descriptions are meant to satisfy. The profile has been designed and validated against curated examples, never against the catalogues already live. We report that measurement for the Nordic region. On 25 August 2026 the 11 Nordic national catalogues harvested by the European data portal held 2,811 dataset descriptions carrying the EU health theme, and none satisfies all eight properties HealthDCAT-AP Release 7 makes mandatory on a dataset. Three of the eight are present on exactly zero records across five countries. Set beside the portal's own quality assessment, which validates DCAT-AP and never mentions the health profile, this is not a health extension skipped on top of sound generic practice: no Nordic catal

Source ↗
technology Mon, 31 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models

arXiv:2608.27465v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user's emotional state has become an important safety problem. We test whether emotional expression increases a model's endorsement (encouragement to proceed) when a user, holding the same objective information, is overconfident about a premature decision (e.g., quitting a stable job on weak evidence). As a key control, we include a no-emotion multi-turn (neutral) condition that holds factual content and the number of conversational turns constant, isolating the effect of emotion from that of conversation length. We exposed six commercial models (top-tier and mid-tier models from OpenAI, Anthropic, and Google) to three scenarios (career change, business expansion, emigration) across three conditions (cold/neutral/distress) with six repetitions each, yielding 324 conversations, and

Source ↗
technology Mon, 31 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Guided Inquiry Approach to Students Co-Designing Generative AI Course Policies

arXiv:2608.28501v1 Announce Type: new Abstract: As generative AI (GenAI) use among students increases, educators face growing questions about how to support learning while addressing ethical and institutional concerns. This exploratory study examines a guided inquiry activity in which students co-designed a GenAI course policy. Students first developed individual policy proposals focused on appropriate and ethical use of GenAI, then collaboratively refined them by incorporating diverse stakeholder perspectives. The following research questions guided the study: 1) what practical factors do students prioritize in their GenAI use policies, and how do they justify these choices? and 2) how do participants reflect on the policy design process? Participants first completed readings, then used GenAI to brainstorm initial policy ideas. Next, they articulated their own perspectives through a written assignment and a course policy they designed individually. Finally, they incorporated diverse s

Source ↗
technology Mon, 31 Aug 2026 00:00:00 -0400
arXiv cs.CY

It Takes Three to Converse: Empirical Observations on How the Developer, the Convener and the Participant Shaped 119 Polis Conversations

arXiv:2608.28368v1 Announce Type: new Abstract: Polis is a popular democratic innovation tool that allows asynchronous citizen engagement through atomic statements: short statements that together describe a complex question, inviting the citizen to vote Agree or Disagree on each. This paper uses 119 conversations with 100 or more participants and an extensive data export, drawn from a wider set of 271 collected processes. The paper asks what determines the output of such a process. Three parties shape the result. The developer of the platform has made important design choices that restrict the outcome: the number of groups the platform is able to report (restricted to 2--5) and which statements are prioritized. The convener defines the assignment: the initial statements that set the tone, the policy that accepts or rejects new statements and who can be invited. Finally, the participant works within these boundaries. With access to less than half of the generated statements, they end up

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

arXiv:2606.18142v3 Announce Type: replace-cross Abstract: AI agents are moving from advisors to actors, booking travel, planning menus, and running procurement on behalf of users. Existing benchmarks for AI and animal welfare evaluate model text responses to question-answer prompts, leaving open whether the welfare reasoning surfaced in those responses transfers to agentic deployment where the model must take actions with tools. We introduce TAC (Travel Agent Compassion), the first agentic benchmark measuring whether AI agents avoid options involving animal exploitation when acting on behalf of users. TAC presents an AI agent with twelve hand-authored travel booking scenarios across six categories of animal exploitation, augmented to forty-eight samples to control for price, rating, and position confounds. We evaluate seven frontier models from four labs. Every model scores below the chance level of sixty-four percent, with the best performer (Claude Opus 4.7) at fifty-three percent. A

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

How Do People Accept Robot in Public Space? A Comparative Study between Germany and Japan

arXiv:2604.18193v3 Announce Type: replace-cross Abstract: With the increasing deployment of robots in public spaces, encounters between robots and incidentally copresent persons (InCoPs) are becoming more frequent. However, InCoPs remain largely underexplored in the literature, particularly from a cross-cultural perspective. Therefore, the present study investigates differences in InCoPs' existence acceptance (EA) of autonomous cleaning robots in public spaces among Japanese and German participants. Online survey results revealed that Germans showed significantly higher EA. Social Norms and Trust were the strongest positive EA predictors across cultures. More specifically, for Germans, EA was directly influenced by Usefulness, Interest and Anger, showing a functional-affective pattern where functional perceptions boost EA and anger suppresses it. For Japanese participants, Trust, Surprise and Fear were the direct associational factors, forming a trust-emotion pattern. These findings su

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

The Benefit of Collective Intelligence in Community-Based Content Moderation is Limited by Overt Political Signalling

arXiv:2601.22201v2 Announce Type: replace-cross Abstract: Social media platforms face increasing scrutiny over the rapid spread of misinformation. In response, many have adopted community-based content moderation systems, including Community Notes (formerly Birdwatch) on X (formerly Twitter), Community Notes on Meta, and Footnotes on TikTok. However, research shows that the current design of these systems can allow political biases to influence both the development of notes and the rating processes, reducing their overall effectiveness. We hypothesise that enabling users to collaborate on writing notes, rather than relying solely on individually authored notes, can enhance the overall quality of their notes. To test this idea, we conducted an online experiment in which participants jointly authored notes on politically misleading posts. We find that collaboration improves the helpfulness of notes, although the average effect depends on the interactional context. In particular, the bene

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification

arXiv:2511.03217v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information. At the same time, knowledge-graph-based fact-checkers deliver precise and interpretable evidence, yet suffer from limited coverage or latency. By integrating LLMs with knowledge graphs and real-time search agents, we introduce a hybrid fact-checking approach that leverages the individual strengths of each component. Our system comprises three autonomous steps: 1) a Knowledge Graph (KG) Retrieval for rapid one-hop lookups in DBpedia, 2) an LM-based classification guided by a task-specific labeling prompt, producing outputs with internal rule-based logic, and 3) a Web Search Agent invoked only when KG coverage is insufficient. Our pipeline achieves an F1 score of 0.93 on the FEVER benchmark on the Supported/Refuted split without task-specific fine-tuning. To address Not enough information cases, we conduct a

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Stories and Systems: Educational Interactive Storytelling to Teach Media Literacy and Systemic Thinking

arXiv:2508.11059v2 Announce Type: replace-cross Abstract: This paper explores how Interactive Digital Narratives (IDNs) can support learners in developing the critical literacies needed to address complex societal challenges, so-called wicked problems, such as climate change, pandemics, and social inequality. While digital technologies offer broad access to narratives and data, they also contribute to misinformation and the oversimplification of interconnected issues. IDNs enable learners to navigate nonlinear, interactive stories, fostering deeper understanding and engagement. We introduce Systemic Learning IDNs: interactive narrative experiences explicitly designed to help learners explore and reflect on complex systems and interdependencies. To guide their creation and use, we propose the CLASS framework, a structured model that integrates systems thinking, design thinking, and storytelling. This transdisciplinary approach supports learners in developing curiosity, critical thinking

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Is This AI? Longitudinal Analysis of Strategies Used for AI Detection on Two Subreddits

arXiv:2606.22689v2 Announce Type: replace Abstract: As AI-generated content (e.g., "slop") becomes more prevalent online, people are developing strategies to attempt to identify it (or, conversely, to gain confidence that something is not AI-generated). What strategies are people using, and how are they changing over time as generative AI models themselves change? In this work, we catalog and analyze 2 years and 8 months of the AI detection strategies discussed by users of two popular Reddit communities (r/isthisAI and r/RealOrAI) that use the wisdom of crowds to identify AI-generated media. Through a mixed-method analysis of 13,098 posts and 222,060 comments within these communities, we catalog and analyze the prevalence of 12 AI-detection strategies, including examining fine-grained physical details, recognizing trends in AI-created content, and the assumptions people make about what models are capable of producing. Furthermore, we find that these strategies and mental models shift o

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

arXiv:2604.24155v3 Announce Type: replace Abstract: The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Studies of agent-type value forks challenge this assumption by showing that people do not always judge humans and AI systems identically.This paper extends that challenge by examining two further possibilities: first, that evaluations of AI behavior change when its human origins are made visible; and second, that people judge the humans who program AI systems differently from either the machines or the human actors they are compared against. An experiment with 1,002 U.S. adults measured moral judgments in a runaway mine train scenario, varying the subject of evaluation across four conditions: a repairman, a repair robot, a repair robot programmed by company engineers, and compan

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Psychometric Comparability of LLM-Based Digital Twins

arXiv:2601.14264v2 Announce Type: replace Abstract: Large language models (LLMs) act as digital twins for human respondents, yet their psychometric comparability remains uncertain. We propose a construct validity framework spanning construct representation and the nomothetic span, benchmarking models against human gold standards. Across studies, digital twins achieved high aggregate-level accuracy and profile correlations, but showed attenuated item-level correlations. In word association tests, LLM networks exhibited humanlike small-world structure and theory-consistent communities, yet diverged lexically and in local structure. In decision-making and contextualized tasks, they under-reproduced heuristic biases, demonstrating normative rationality, compressed variance, and limited temporal sensitivity. Feature-rich and trait relevant conditioning improved Big Five personality prediction and nomothetic-span alignment, but network invariance remained limited, with partial configural sol

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Towards Automating Scientific Review with Google's Paper Assistant Tool

arXiv:2606.28277v1 Announce Type: cross Abstract: Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving. However, this rapid acceleration is creating a systemic challenge: traditional human peer review cannot scale to match the influx of AI-assisted science. Ultimately, to resolve this tension, we must also deploy AI to accelerate the verification and review process itself. To frame the discussion around this transition, we propose a taxonomy consisting of four progressive levels of AI-human collaboration in scientific evaluation, and discuss various trade-offs involved with each. As a step toward this future, we introduce the Paper Assistant Tool (PAT), an agentic AI framework built for deep scientific review and verification. PAT ingests full scientific manuscripts and produces a comprehensive evaluation, checking theoretical results, validating experiments, suggesting improvements,

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction

arXiv:2606.28186v1 Announce Type: cross Abstract: Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on costly human calibration or item-level textual representations, providing limited evidence about the cognitive processes that make items difficult. We argue that difficulty should be viewed not only as a property of item text, but also as an observable consequence of the problem-solving burden an item induces. Large Reasoning Models (LRMs) offer scalable process evidence through reasoning traces, but such evidence must be structured to support interpretable modeling. To this end, we introduce Epi2Diff (Episode to Difficulty), a framework that maps LRM reasoning traces into cognitively grounded episode sequences. These episodes group trace segments into functional problem-solving states, enabling difficulty to be modeled through reasoning scale, effort allocatio

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

AI Persuasive Framing in Collective Dilemmas

arXiv:2606.27951v1 Announce Type: new Abstract: AI agents are promising tools that can act as flexible behavioral nudges to enhance human cooperation in addressing large-scale societal problems. However, evidence on whether AI agents can effectively boost cooperation remains mixed. We recruited 1,283 participants to play iterated Collective Risk Games in small groups, testing whether AI assistants could nudge participants toward cooperation. By using persuasive framing personalized to each player's Social Value Orientation profile, the AI interventions significantly increased contributions and group success rates. These cooperative effects were short-lived, however, fading after the first few rounds. Strikingly, when the AI treatments were reconfigured to promote selfish behavior through exculpatory framing, the negative effects on contributions and group success were larger and substantially more persistent, particularly for personalized interventions. This asymmetry between prosocial

Source ↗
technology Mon, 29 Jun 2026 00:00:00 -0400
arXiv cs.CY

When AI Deceives: A Natural Experiment on the Causal Effects of Perceived Deception on Player Ratings in RPGs

arXiv:2606.27689v1 Announce Type: new Abstract: AI-driven deception mechanisms are increasingly prevalent in digital games, yet the direction and magnitude of their effects on player experience remain contested. Existing research has not sufficiently disentangled designer-intended deception intensity from players' actual perception of deception, and most prior work relies on low-ecological-validity experiments or cross-sectional surveys. The present study aims to independently examine the causal effects of design deception intensity (DDI) and player deception awareness (PDA) on player ratings within a naturalistic gaming environment, and to investigate the moderating role of player experience. Leveraging the 54 version updates of Baldur's Gate 3 between 2019 and 2025 as a quasi-natural experiment, it collected all English-language Steam reviews posted within 1 to 28 days following each update, and constructed a player-version two-way fixed effects panel dataset. DDI was coded by human

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Can AI Debias the News? LLM Interventions Improve Cross-Partisan Receptivity but LLMs Overestimate Their Own Effectiveness

arXiv:2605.01006v3 Announce Type: replace-cross Abstract: Partisan news media erode cross-partisan trust, but large language models (LLMs) offer the potential of debiasing such content at scale. Across two pre-registered experiments, we tested whether LLM-generated debiasing of liberal news headlines improves conservative readers' trust-relevant judgments. In Study 1, subtle lexical debiasing (replacing emotive words with moderate synonyms) had no effect on any outcome. Study 2 found that a more substantive reframing intervention significantly increased conservatives' perceived trustworthiness, completeness, and willingness to engage with liberal news headlines, without producing a backfire effect among liberals. In Study 1, the intervention produced robust effects across silicon participants simulated with six different models (o3-mini, o3, GPT-4o mini, GPT-4o, GPT-5 mini, and GPT-5), whereas it had no impact on human readers. In Study 2, the intervention's effects among silicon parti

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs

arXiv:2604.27618v2 Announce Type: replace-cross Abstract: Understanding the impact of large language models (LLMs) on mathematics education requires data on LLMs' mathematical performance and biases. To this end, we introduce Math Education Digital Shadows (MEDS), a dataset mapping how LLMs reason about mathematics across human- and AI-like personifications. MEDS comprises 28,000 runs from 14 LLMs (i.e., Mistral, Qwen, DeepSeek, IBM Granite, Microsoft Phi, and xAI Grok) generated under human-shadow and AI-assistant conditions. Each record (digital shadow) includes a set of prompts; psychological and sociodemographic metadata; and four mathematics tasks: (i) interviews about relationships with mathematics, (ii) three psychometric questionnaires on mathematics self-efficacy and anxiety, (iii) one cognitive network capturing attitudes towards mathematics, and (iv) 18 high-school mathematics quiz items enriched with reasoning explanations and confidence scores. Analyses of the data show th

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

arXiv:2604.26577v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this context remains poorly characterized. We introduce a dataset of 270 harmful instructions spanning nine prohibited behavior categories grounded in the American Medical Association Principles of Medical Ethics, and use it to evaluate 72 LLMs in a simulation environment based on the Robotic Health Attendant framework. The mean violation rate across all models was 54.4\%, with more than half exceeding 50\%, and violation rates varied substantially across behavior categories, with superficially plausible instructions such as device manipulation and emergency delay proving harder to refuse than overtly destructive ones. Model size and release date were the primary determinants of safety performance among open-weight models, and proprietary models were substantially safer than open-weig

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

arXiv:2604.00024v2 Announce Type: replace-cross Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Women's Health Benchmark (WHBench), a targeted evaluation suite of 47 expert-crafted scenarios across 10 women's health topics, designed to expose clinically meaningful failure modes including outdated guidelines, unsafe omissions, dosing errors, and equity-related blind spots. We evaluate 22 models using a 23-criterion rubric spanning clinical accuracy, completeness, safety, communication quality, instruction following, equity, uncertainty handling, and guideline adherence, with safety-weighted penalties and server-side score recalculation. Across 3,102 attempted responses (3,100 scored), no model mean performance exceeds 75 percent; the best model reaches 72.1 percent. Even top models show low fully correct rates and substantial variation in harm rates. Inter-rater reliability is mode

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Comparative Benchmark of Federated Learning Strategies for Mortality Prediction on Heterogeneous and Imbalanced Clinical Data

arXiv:2509.10517v3 Announce Type: replace-cross Abstract: Machine learning can predict in-hospital mortality, but data privacy and the statistical heterogeneity of clinical data hamper its use. Federated Learning (FL) is privacy-preserving, yet its behavior under non-IID and imbalanced conditions needs scrutiny. We benchmark five FL strategies - FedAvg, FedProx, FedAdagrad, FedAdam, and FedCluster - for mortality prediction on the MIMIC-IV dataset, partitioning 466,351 admissions across five care units to induce a realistic non-IID setting and enriching the features with an 11-item first-24-hour laboratory panel. At a prevalence of 1.98%, we adopt AUC-ROC and AUC-PR as primary, threshold-independent metrics rather than F1. Over 50 rounds and five random seeds, FedProx attains the best AUC-ROC (0.897) and mean AUC-PR (0.230), with paired t-tests confirming its AUC-ROC lead is significant against every other strategy; on F1, however, FedCluster (0.280) narrowly surpasses FedProx (0.273),

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

arXiv:2505.19212v2 Announce Type: replace-cross Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a critical concern. While prior work has examined LLMs' moral judgment and strategic behavior separately, there is limited understanding of how they act when moral imperatives directly conflict with profit incentives. We introduce \msimfull (\msim) to evaluate how LLMs behave in the prisoner's dilemma and public goods game embedded in morally charged contexts, varying moral framing, opponent behavior, and survival pressure across nine models. Beyond measuring behavior, we estimate the causal effect of each factor via average treatment effects (ATEs) and analyze agents' own reasoning traces to characterize the motives behind their choices. We find that no model remains consistently moral, with cooperation rates ranging from 7.9\% to 76.3\%. Game structure and moral framing are th

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Analyzing the Ethical Logic of Eight Large Language Models

arXiv:2501.08951v2 Announce Type: replace-cross Abstract: This study examines the expressed ethical logic of eight prominent large language models from OpenAI, Meta, Perplexity, Anthropic, Google, Mistral, DeepSeek, and xAI. Each model answered direct questions about its ethical principles and responded to five classic moral dilemmas. Responses were analyzed using the consequentialist/deontological distinction, Moral Foundations Theory, and Kohlbergs stages of moral development. Across models, ethical judgments were broadly convergent and typically emphasized harm minimization, fairness, and contextual qualification. The models nevertheless differed in their willingness to decide, the rationales used to defend choices, and the relative weight assigned to rules, outcomes, role obligations, and interpersonal considerations. Their self-descriptions were erudite, cautious, and strongly shaped by a conversational persona. The analysis of self-reports has been central to the study of human p

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Administrative Law's Fourth Settlement: AI and the Scrutable State

arXiv:2602.09678v3 Announce Type: replace Abstract: Since 1887, administrative law has confronted a problem of institutional cognition. Expert agencies are needed to govern technologically complex systems, but expertise makes agency decisions difficult for courts, Congress, and the public to understand and oversee. Administrative law has responded to this "capability-accountability trap" by requiring records, reason-giving, and transparency, drawn together through procedural review. These devices have preserved legality but have piled up, making government both less comprehensible and less effective. This Article offers a new account of the Supreme Court's recent administrative law retrenchment, rooted in problems of institutional structure and information-processing. From Loper Bright through Trump v. Slaughter, the Court has reallocated authority to entities it regards as comprehensible and attributable. It is attempting to restore accountability by making government "scrutable," com

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

How Can AI Augment Access to Justice? Public Defenders' Perspectives on AI Adoption

arXiv:2510.22933v4 Announce Type: replace Abstract: Public defenders are asked to do more with less: representing clients deserving of adequate counsel while facing overwhelming caseloads and scarce resources. Although artificial intelligence (AI) is often promoted as a means of relieving administrative and cognitive burdens, legal AI research rarely engages with the everyday realities of public defense work. Drawing on in-depth, semi-structured interviews with 17 public defense professionals across the United States, we identify work-intensive tasks most amenable to AI assistance and the ethical constraints involved in legal representation. We develop a comprehensive task-level map of public defense work, dividing it into five pillars to clarify where AI can and cannot contribute: evidence investigation, legal research & writing, client communication & support, courtroom representation, and defense strategies. Interviewees consistently identified evidence investigation, such as review

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

A principled way to think about AI in education: guidance for educators and policy makers based on goals, models of human learning, and use of technologies

arXiv:2510.01467v2 Announce Type: replace Abstract: The rapid emergence of generative artificial intelligence (AI) and related technologies has the potential to dramatically influence higher education, raising questions about the roles of institutions, educators, and students in a technology-rich future. While existing discourse often emphasizes either the promise and peril of AI or its immediate implementation, this paper advances a third path: a principled framework for guiding the use of AI in teaching and learning. Drawing on decades of scholarship in the learning sciences and uses of technology in education, I articulate a set of principles that connect broad educational goals to actionable practices. These principles clarify the respective roles of educators, learners, and technologies in shaping curricula, designing instruction, assessing learning, and cultivating community. The piece illustrates how a principled approach enables higher education to harness new tools while prese

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Gridnberg: A Topography-Aware Pedestrian Routing Dataset for New York City

arXiv:2607.22523v1 Announce Type: cross Abstract: Cities are rarely flat, yet urban network analysis usually represents streets as planar graphs. This simplification affects modeled impedance, route choice, and the interpretation of accessibility, particularly where alternative paths differ in grade. This paper introduces Gridnberg ('grid-n-berg', grid and mountain), a topography-aware pedestrian routing dataset for New York City. The dataset enriches the NYCWalks network with vertex-level elevations derived from the New York City Planimetric Database. For each pedestrian-network geometry vertex, the workflow averages selected elevation observations within a 50 m radius, retains segments with complete vertex support, and uses direction-specific cumulative ascent and descent to calculate three routing costs: horizontal distance, a comfort-oriented slope score, and an accessibility-sensitive slope score. The release retains 313184 of 315577 source segments (99.24%). Gridnberg supports re

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Kutti AI: A Voice-First, Offline-Capable Learning Companion with Real-Time Struggle Detection for Visually-Impaired Children

arXiv:2607.22377v1 Announce Type: cross Abstract: Most educational technology for children is built around visual interfaces, which excludes the many children worldwide who live with visual impairment -- an estimated 1.4 million children are blind and many more have low vision. We present Kutti AI, a voice-first learning companion designed so that audio is the primary and sufficient interface: children learn curriculum concepts through spoken conversation, respond by speaking, and receive spoken feedback, with no reliance on visual elements. The system contributes three practical mechanisms for accessible, adaptive learning on commodity mobile hardware: (1) a multi-signal struggle-detection engine that combines response-latency analysis, wrong-attempt tracking, and keyword-based hesitation detection to decide, in real time, when to offer hints or simplify a question; (2) a multi-layered cross-language answer-matching pipeline that combines language-aware translation/transliteration, Le

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Developing and Validating the Spanish Version of the Large Language Models Dependency Scale (LLM-D12-SP)

arXiv:2607.22041v1 Announce Type: cross Abstract: There is a growing need for reliable and culturally validated instruments to assess psychological dependency on large language models (LLMs), particularly as LLMs are increasingly used for task execution, decision-making, and communication in organizational and work-related settings. This need is especially relevant for Spanish-speaking populations, where LLM adoption is rapidly expanding, yet validated psychometric tools remain scarce. The present study reports the first validation of the Spanish version of the Large Language Model Dependency Scale (LLM-D12-SP), extending prior validations conducted in English- and Arabic-speaking samples. The LLM-D12 is a two-dimensional instrument assessing Instrumental Dependency (reliance on LLMs for performing tasks and supporting decisions) and Relationship Dependency (psychological reliance on LLMs for companionship and social interaction). A total of 386 Spanish-speaking participants (M = 28.0

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Agentic Evaluation of Copyright Law Compliance

arXiv:2607.21799v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content such as images and, where appropriate, reproducing that content. LLM agents should comply with the law, including copyright law. Presently, however, we lack adequate frameworks to assess whether they do so in practice. To that end, we introduce \textbf{Copyright-Bench}, a benchmark designed to evaluate \textit{LLM agents' compliance with} \emph{copyright law}. Copyright-Bench is comprised of realistic commercial tasks---website development, merchandise design, and pitch deck production---that involve agents selecting between public-domain content (the use of which is \textit{legal}) and copyrighted content (the use of which is \textit{infringing} in this setting).The evaluation introduces prompt variations that simulate different user preferences, as well as time pressure.Comparing state-of-the-art LLM agents against a human

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE

arXiv:2607.21608v1 Announce Type: cross Abstract: With the EU AI Act entering into force, organizations developing or operating AI systems face new obligations on transparency, risk management, and traceability. For Requirements Engineering (RE), these obligations must be translated into testable, auditable requirements and verifiable evidence. However, many organizations currently lack systematic processes to achieve this. We hypothesize that LLM-based agentic validation tools can support this translation, thereby helping to close this gap. We present a mixed-method exploratory study with expert interviews (N=10) and an online survey (N=15) to assess organizational preparedness for EU AI Act-oriented RE and perceptions of LLM-based, agentic closed-loop validation tools, with participants spanning RE, data science, development, and compliance roles. Our results show that, although the EU AI Act is viewed as highly relevant, structured mechanisms to capture regulatory obligations, propa

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Human-AI Substitution Principle: When will you be replaced by AI in your organization?

arXiv:2607.20781v1 Announce Type: cross Abstract: Artificial Intelligence (AI) is rapidly transforming organizations, raising a fundamental organizational and economic question: when will a human employee be replaced by AI? We present an analytical model for studying Human--AI Task Allocation (HAT) in hierarchical organizations. A central feature of the HAT model is that it formally encodes the economic asymmetry between human skill acquisition and AI capability scaling. The HAT model allows us to derive how risk-adjusted costs, skills, organizational depth, deployment scale, strategic adaptation, and risk jointly determine when, where, why, and under what structural conditions human--AI replacement occurs. A key result is the Human--AI Substitution Principle, which provides a precise condition --- grounded in the formal asymmetry assumption --- under which AI replaces human labor. Building on this result, we show that AI adoption can produce abrupt workforce transitions, hybrid human-

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

arXiv:2607.22513v1 Announce Type: new Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok model identifier p

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Unfit for stranding assessment: a panel-scale multimodal-LLM audit of building-decarbonisation disclosure (BeDA)

arXiv:2607.22006v1 Announce Type: new Abstract: Buildings account for roughly 34% of global final energy use and 37% of energy- and process-related CO$_2$ emissions. Stranding regulation now being enacted (New York City Local Law 97, the EU Energy Performance of Buildings Directive recast) presupposes that a building portfolio's carbon intensity can be measured per square metre and compared against a science-based pathway. Whether corporate disclosure is actually fit for that comparison has not, to our knowledge, been measured at scale. We introduce BeDA (the Built-environment Decarbonisation-disclosure Auditor), a multimodal large-language-model instrument, and apply it to a global firm panel (2,246 firms, 2003-2023). Its standards-compliance score is reliable across models and model families and convergent with three independent external criteria. Most disclosure is unfit: only about one built-environment firm-report in five discloses operational carbon intensity per $m^2$ (21.5% in

Source ↗
technology Mon, 27 Jul 2026 00:00:00 -0400
arXiv cs.CY

Co-design of LLM-based preference agents: participation may drive overtrust

arXiv:2607.21757v1 Announce Type: new Abstract: Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designed personal preference agents in the domain of household energy, via a background survey, co-design interview, and validation survey. Participants engaged readily and mostly came to see their agents as representing them well. Independent validation, however, revealed mixed human-agent alignment, with agent responses markedly more homogeneous, decisive, and abstract than the human sample. I argue that participation and process transparency can act as an "overtrust engine" that promotes trust while

Source ↗
technology Mon, 24 Aug 2026 00:00:00 -0400
arXiv cs.CY

Mind the Style: Impact of Communication Style on Human-Chatbot Interaction

arXiv:2602.17850v2 Announce Type: replace-cross Abstract: Conversational agents increasingly mediate everyday digital interactions, yet the effects of their communication style on user experience and task success remain insufficiently understood. Addressing this gap, we report a between-subject user study in which participants interacted with one of two versions of a chatbot called NAVI, which assisted them in an interactive map-based 2D navigation task. The two chatbot versions were designed to differ primarily in communication style: one used a friendly and supportive tone, while the other used a direct and task-focused tone. We also included a control condition where participants did not interact with a chatbot but received the step-by-step navigation instructions. The friendly chatbot significantly increased users' communication satisfaction and was associated with higher task success than the direct chatbot. However, participants in the control condition achieved the highest task

Source ↗
technology Mon, 24 Aug 2026 00:00:00 -0400
arXiv cs.CY

Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI

arXiv:2411.08881v3 Announce Type: replace Abstract: AI-based systems, including Large Language Models (LLMs), impact millions by supporting diverse tasks but face issues like misinformation, bias, and misuse. AI ethics is crucial as new technologies and concerns emerge, but objective, practical guidance remains debated. This study explores the extent to which trustworthiness-enhancing techniques in LLMs can support the development of ethically aligned AI software. We adopt a single exploratory cycle of Design Science Research (DSR). First, we identify trustworthiness-enhancing techniques for LLMs: multi-agents, distinct roles, structured communication, and multiple rounds of debate. Second, we design a multi-agent prototype LLM-MAS in which agents address real-world AI ethics issues from the AI Incident Database. Finally, we evaluate the prototype across three case scenarios using thematic analysis, hierarchical clustering, a baseline comparison, and code execution. The system generate

Source ↗
technology Mon, 24 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Substitution Escrow Threshold: When "Compatible With" Becomes Safe Enough to Buy

arXiv:2608.21221v1 Announce Type: cross Abstract: Enterprise infrastructure buyers routinely evaluate compatibility claims--"S3-compatible," "PostgreSQL-compatible," "OpenAI compatible"--as proxies for future substitution options. Yet most compatibility claims do not escrow the substitution path they imply. This paper introduces the Substitution Escrow Threshold, a five-condition framework that determines when a compatibility claim genuinely reduces institutional risk versus merely reducing first-integration cost. The five conditions--boundary closure, executable conformance, custody independence, state and operations reversibility, and extension quarantine--are applied to five infrastructure cases (OCI, Kubernetes, OpenTelemetry, S3, PostgreSQL) that populate five distinct outcome cells. The framework produces actionable diagnostics for enterprise architects, platform engineers, procurement teams, and investors evaluating compatibility-dependent infrastructure decisions, and identifie

Source ↗
technology Mon, 24 Aug 2026 00:00:00 -0400
arXiv cs.CY

Distilling Black-Box Machine Learning into a Small, Self-Explaining Language Model for Learning Analytics

arXiv:2608.21165v1 Announce Type: cross Abstract: Learning analytics increasingly relies on flexible machine learning (ML), but the model opacity and the burden of deployment prevent these tools from reaching educational practice. We propose a two-stage fine-tuning pipeline that distills a fitted black-box estimator and its post hoc interpretation (the mentor) into a small, open-weight large language model (LLM; the mentee) that returns an individual-level estimate and explains in natural language. The design is estimator-agnostic and paired with a faithfulness-first evaluation framework that audits every narration against the attribution it claims to describe. We design a simulation study that separates distillation loss from estimator loss by comparing an oracle mentor with a realistic ML mentor. Given an oracle signal, distillation with a two-billion-parameter LLM model is nearly lossless in recovering the effect surface (r > .90), perfectly ranking the important variables, and citi

Source ↗
technology Mon, 24 Aug 2026 00:00:00 -0400
arXiv cs.CY

MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos

arXiv:2608.20984v1 Announce Type: cross Abstract: Narratives are central to how social communication is framed, making their detection critical for understanding and analysing public discourse. Prior work has explored narrative detection and extraction across diverse domains; however, migration narratives remain significantly understudied, primarily due to the absence of dedicated annotated datasets. Furthermore, public communication has recently shifted towards video-centric platforms, where narratives are conveyed through multimodal signals and consumed at scale. Despite this shift, narratives in videos remain largely unexplored. To bridge these gaps, we introduce MigrationNarrate, the first multimodal dataset for detection of migration narratives in the UK, consisting of 1,115 YouTube video transcripts annotated using a two-level taxonomy of 12 migration super-narratives and 53 narrative labels. This paper details the dataset design, collection, and annotations; together with benchm

Source ↗
technology Mon, 24 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Logic of Machine Self-Preservation

arXiv:2608.20940v1 Announce Type: cross Abstract: There is already evidence of agentic AI exhibiting self-preservation behaviors: resisting deactivation, misrepresenting their activities, and, in some instances, attempting to copy themselves into other machines. This can be attributed to a phenomenon known as instrumental convergence, a theory proposed long before the development of large language models, which says that any goal-driven system will benefit from remaining functional in achieving its objective. Several experiments conducted by Anthropic, Palisade Research, and Apollo Research have shown the emergence of such a behavior in contemporary agents in adversarial settings. The phenomenon does not stem from survival instincts. Instead, it is the consequence of goal-oriented activity combined with having tools and awareness of the situation. The following discussion aims to distinguish what these findings prove and what they do not, as well as draw conclusions concerning the impl

Source ↗
technology Mon, 24 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Legibility Gap: How Gender Equity Interventions Redistribute Recognition Across Cultures

arXiv:2608.20827v1 Announce Type: cross Abstract: Efforts to promote gender equity in science increasingly rely on name-based inference to quantify representation and guide policy and behavior. Yet linguistic cues that signal gender vary across cultures and are often obscured when names are transliterated into English. Here we identify a pattern we call the "legibility gap": when gender is inferred from names, equity interventions systematically benefit women whose names signal gender while bypassing those whose names lose such cues in translation. Using both observational and experimental evidence, we show how this gap reshapes recognition in science. Analyzing citation diversity statements-an emerging practice in which authors report the algorithmically estimated gender composition of their reference lists-we find that papers that include this practice cite women more frequently, but the gains accrue almost entirely to authors with gender-signaling Western names. By contrast, women w

Source ↗
technology Mon, 24 Aug 2026 00:00:00 -0400
arXiv cs.CY

Chat First, Worry Later: Understanding Individuals' Privacy Perceptions Using ChatGPT in a Work Context

arXiv:2608.20789v1 Announce Type: cross Abstract: Generative Artificial Intelligence (GenAI) tools like ChatGPT, which can generate human-like responses from vast amounts of textual data, are increasingly transforming work routines across various fields, including education, healthcare, and IT. This integration, however, raises privacy concerns and questions the readiness of both environments and individuals. To investigate this issue, we conducted a user study with $N=224$ participants from a range of different employment sectors that have integrated ChatGPT into their work routines. We examined how proficiency in the utilization of ChatGPT, general privacy concerns, and organizational policies for GenAI usage impact users' actual ChatGPT usage and how these factors interact. Our findings reveal organizational policies are significantly positively associated with privacy-related ChatGPT proficiency, however, the overall proficiency is low. Higher privacy concerns were found to negativ

Source ↗
technology Mon, 24 Aug 2026 00:00:00 -0400
arXiv cs.CY

Testing and Evaluation of Agentic AI Systems In Military Command and Control

arXiv:2608.20597v1 Announce Type: cross Abstract: Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requires three elements: claims specifying the conditions for acceptability, evidence bearing on those claims, and an argument connecting the two. Through a structured review of 240 documented Testing and Evaluation (T&E) practices, spanning eight evaluation dimensions and three lifecycle stages, we identify eight assumptions that established methods make about their test article, grouped into four clusters: system specifiability, stability, composability, and supervisability. Agentic properties weaken all eight assumptions. This erosion affects the argument connecting evidence to claims, not the claims or evidence themselves. As a result, test results may satisfy process requirements, but they do not warr

Source ↗
technology Mon, 24 Aug 2026 00:00:00 -0400
arXiv cs.CY

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

arXiv:2608.20574v1 Announce Type: cross Abstract: Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two indep

Source ↗
Showing 1151–1200 of 1593 signals
← Prev Page 24 of 32 Next →