EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Evidence of conceptual mastery in the application of rules by Large Language Models

arXiv:2503.00992v2 Announce Type: replace-cross Abstract: In this paper we leverage psychological methods to investigate LLMs' conceptual mastery in applying rules. We introduce a novel procedure to match the diversity of thought generated by LLMs to that observed in a human sample. We then conducted two experiments comparing rule-based decision-making in humans and LLMs. Study 1 found that all investigated LLMs replicated human patterns regardless of whether they are prompted with scenarios created before or after their training cut-off. Moreover, we found unanticipated differences between the two sets of scenarios among humans. Surprisingly, even these differences were replicated in LLM responses. Study 2 turned to a contextual feature of human rule application: under forced time delay, human samples rely more heavily on a rule's text than on other considerations such as a rule's purpose.. Our results revealed that some models (Gemini Pro and Claude 3) responded in a human-like manne

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation

arXiv:2502.13207v4 Announce Type: replace-cross Abstract: Despite the increasing use of large language models for creative tasks, their outputs often lack diversity. Common solutions, such as sampling at higher temperatures, can compromise the quality of the results. Dealing with this trade-off is still an open challenge in designing AI systems for creativity. Drawing on information theory, we propose a context-based score to quantitatively evaluate value and originality. This score incentivizes accuracy and adherence to the request while fostering divergence from the learned distribution. We show that our score can be used as a reward in a reinforcement learning framework to fine-tune large language models for maximum performance. We validate our strategy through experiments considering a variety of creative tasks, such as poetry generation and math problem solving, demonstrating that it enhances the value and originality of the generated solutions.

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Qualifying and Quantifying Risk Under the EU AI Act

arXiv:2608.08564v2 Announce Type: replace Abstract: The EU AI Act uses a risk-based approach to regulate AI systems, calibrating the intensity of regulation according to the risks they pose. While the term 'risk' implies quantification, resulting from the combination of the probability and severity of harm, the AI Act refers to risks to fundamental rights, thereby engaging a qualitative perspective. In this piece, we address this puzzle using a two-step framework under which the EU AI Act balances risks with the protection of fundamental rights, the legitimate purposes of providers and deployers, and the impacts of regulatory measures on providers, deployers, and regulators. We discuss this framework against the backdrop of potential approaches to quantifying risks, with a specific focus on defining and measuring the main components of the concept of risk: probability, severity, and their combination. We suggest that the protection of fundamental rights and risk quantification can be a

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Agentic AI: User Empowerment or Foreclosure?

arXiv:2608.06510v2 Announce Type: replace Abstract: Agentic AI promises systems that can act on users' behalf, from filtering content to negotiating prices to selecting services. Whether it will empower users is an open question, and one that depends on more than the technology. We conduct a comparative case analysis of four earlier, more mature domains in which similar forms of agency emerged: browser-based ad blockers, platform recommender systems, financial robo-advisors, and email spam filtering. Across the cases, questions about whose interests agents would serve were resolved through technical arrangements: API choices, protocol governance, industry standards, and default configurations. Beyond their technical form, these were political decisions. We identify this settling of contestable questions in a technical form as depoliticization, a concept from political theory, here at work in technological systems. Its most consequential effect is that individual outcomes and collective

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Culturally-Aware AI for Cross-Boundary Community Learning: Undergraduate Innovation at the Intersection of Computation and Design

arXiv:2606.09041v2 Announce Type: replace Abstract: Research on artificial intelligence in education (AIED) is rapidly expanding, yet technical progress often lacks human-centered grounding and adequate attention to cultural context. Community-Based Learning, a pedagogy rooted in social work, remains underrepresented in AIED research, particularly within Asia-Pacific contexts. This paper reports on cross-boundary Community-Based Learning where undergraduate students develop AI-enabled solutions for cultural heritage preservation and sustainable development. We examine how community-engaged computing operationalizes culturally aware, human-centered AIED through participatory elicitation of cultural knowledge, bilingual representation, and stakeholder validation across education, technology, and culture. We contribute a collaborative framework for culturally aware AIED designed to support multi-stakeholder collaboration and widen participation by bridging social work and computational sc

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Building Digital Societies as Ecosystems: How Recognition and Repeat Relationships Sustain Cross-Community Work in Open Source

arXiv:2605.25055v2 Announce Type: replace Abstract: We measure cross-boundary collaboration in an open-source software (OSS) ecosystem by reconstructing the bipartite contributor-repository graph of 464 cybersecurity projects and 11,372 contributors active over October 2001-May 2022 (Rawsec Cybersecurity Inventory). Louvain community detection identifies 163 non-singleton communities; per-community contributor count scales superlinearly with repository count (n_contributors ~ n_repos^1.4), and community formation follows a logistic trajectory saturating around 2018. Three patterns support a recognition/repeat-relationship account of cross-boundary work. First, cross-community work concentrates in a thin carrier layer: only nine canonical humans span seven or more communities at the commit level, authoring 14% of 4,015 inter-community merged pull requests; the top 50 cross-community contributors produce 54%. Second, boundary friction is a recognition cost, not a fixed boundary property:

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Scoping Review of the Negative Effects of Digital Technology on Cognition

arXiv:2603.10025v2 Announce Type: replace Abstract: The rapid integration of digital technology into daily life has prompted sustained concern regarding its impact on human cognition. To characterize documented negative effects and the conditions under which they arise, we conducted a scoping review isolating the documented negative effects of digital technology use on cognition. Using a hybrid automated and manual search strategy, we identified foundational seed papers via Scopus and executed an algorithmic citation snowballing process via the OpenAlex API to capture relevant empirical and non-empirical literature. The resulting synthesis of 937 papers (584 empirical, 353 non-empirical) spans legacy screens, multitasking, smartphones, and the nascent work on generative artificial intelligence (AI). Evidence suggests an evolution in the nature of cognitive risk: while research on earlier technologies predominantly describes disruptions to resource allocation, early findings on AI point

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Fake Friend Dilemma: Relational Trust and the Political Economy of Conversational AI

arXiv:2601.03222v2 Announce Type: replace Abstract: As conversational AI systems become a larger part of the media landscape, they raise questions about whose interests they serve and the risks they may pose to users. These systems do more than provide information: they increasingly offer advice and companionship through interfaces that can appear supportive and socially responsive. A pressing concern is that users may form social or interpersonal relationships with these systems and place relational trust in them, even when the interests shaping interactions do not fully align with their own. The Fake Friend Dilemma (FFD) describes the problem that follows: the same relational trust that makes conversational AI useful can also leave users open to manipulation and exploitation when institutional interests conflict with their own. Drawing on work on trust, AI alignment, dark patterns, and surveillance capitalism, the paper considers how the FFD can manifest through product sales, propag

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Order of Recommendation Matters: Structured Exploration for Improving the Fairness of Content Creators

arXiv:2510.20698v2 Announce Type: replace Abstract: Social media platforms provide millions of professional content creators with sustainable incomes. Their income is largely influenced by their number of views and followers, which in turn depends on the platform's recommender system (RS). So, as with regular jobs, it is important to ensure that RSs distribute revenue in a fair way. For example, prior work analyzed whether the creators of the highest-quality content would receive the most followers and income. Results showed this is unlikely to be the case, but did not suggest targeted solutions. In this work, we first use theoretical analysis and simulations on synthetic datasets to understand the system better and find interventions that improve fairness for creators. We find that the use of ordered pairwise comparison overcomes the cold start problem for a new set of items and greatly increases the chance of achieving fair outcomes for all content creators. Importantly, it also main

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology

arXiv:2506.16697v2 Announce Type: replace Abstract: Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human instruments to LLMs without establishing that the outputs are reliable or interpretable, raising the risk of measurement phantoms--statistical regularities mistaken for genuine psychological phenomena. This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant. It develops a dual-validity framework in which evidentiary demands scale with scientific ambition: from tool use through behavioral characterization and human simulation to cognitive modeling. Classifying text may require only accuracy and reliability; claiming that an LLM simulates anxiety or illuminates cognitive mechanisms requires additional evidence, including construct validity evidence

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Political Ideology of Large Language Models: Measurement, Inconsistency, and Persuasive Influence

arXiv:2505.04171v2 Announce Type: replace Abstract: Large Language Models (LLMs) are a transformational technology, fundamentally changing how people obtain information and interact with the world. As people become increasingly reliant on them for an enormous variety of tasks, a body of academic research has developed to examine these models for inherent biases, especially political biases, often finding them small. We challenge this prevailing wisdom. First, by comparing 43 LLMs to legislators, judges, and a nationally representative sample of U.S. voters, we show that LLMs' apparently moderate overall partisan positioning is the net result of offsetting strongly partisan expressed positions on specific topics, much like moderate voters. Second, in a pre-registered randomized experiment, we show that LLMs can exert persuasive influence on political attitudes. Voters randomized to discuss a policy issue with an LLM shift toward that model's measured ideological position by 3.5 percenta

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Assessing AI-Generated vs. Human-Authored Spear Phishing SMS Attacks: An Empirical Study

arXiv:2406.13049v3 Announce Type: replace Abstract: Personalized phishing is difficult to defend against because messages can be tailored to a target's work, interests, and social context. Large language models may make such tailoring faster and easier, but it remains unclear whether messages produced from simple prompts are more convincing than those written by people. This 25-target pilot study compared personalized smishing messages generated by GPT-4 with messages written by novice student authors working under time constraints. Using the proposed Threshold Ranking Approach for Personalized Deception (TRAPD), participants ranked 12 messages written for them, indicated the point at which they would intend to click, explained their reasoning, and judged whether each message was authored by GPT-4 or a human. GPT-4-generated messages elicited an intention to click more often than student-authored messages (28% versus 21%), although the difference was uncertain. More broadly, our findin

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

The New Mathematics of Democracy

arXiv:2608.16869v1 Announce Type: cross Abstract: This article surveys emerging directions in the mathematics of democracy. It uses three case studies --- voting theory, participatory budgeting, and deliberative democracy --- to highlight how contemporary challenges motivate rigorous mathematical research that incorporates real-world data, institutional constraints, and implementation feasibility. Within each case, we highlight active and promising research frontiers, evidence of real-world impact, practical applications, and opportunities for getting involved.

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

"If It Looks Like a User": Measuring Real-Time Moderation Effects via Social Media Simulation

arXiv:2608.16601v1 Announce Type: cross Abstract: Agent-based social media simulators offer a controlled environment to study content moderation, yet their value hinges on how faithfully they reproduce real platform dynamics. We develop a calibrated extension of SimSoM, an agent-based model of information diffusion on social networks, grounded in a real-world dataset of online vaccine discourse during the COVID-19 pandemic. Our approach replaces ad-hoc parametrisations with empirically fitted distributions, optimised via CMA-ES (Covariance Matrix Adaptation Evolution Strategy) and validated against real data across temporal, distributional, and structural dimensions. Using this validated simulator, we provide three key contributions. First, we show that the calibrated model reproduces key statistical signatures of the empirical data, including activity distributions, post/reshare ratios, and temporal patterns. Second, we apply established misinformation-spreader detection and preventio

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

The User Side of AI Model Lifecycles: Evidence from the Keep4o Movement

arXiv:2608.16574v1 Announce Type: cross Abstract: AI model lifecycles are commonly understood as a series of technical and organizational processes. Yet once a model enters sustained use, subsequent changes can also affect established user practices and user value. Using the Keep4o movement around GPT-4o as a case, this study examines post-deployment AI model lifecycle issues from the user side. We collected 61,846 public original posts on X from August 2025 to March 2026 and, using a systematically developed coding framework and LLM-assisted content analysis, analyzed discussion themes, users' reasons for wanting to keep GPT-4o, and the specific claims they made. Findings show that the Keep4o discussion extended well beyond continued access to the model itself. It covered concrete experiences of use, model behavioral characteristics and how they changed, and management issues across different stages of the model lifecycle. Reasons for keeping GPT-4o reflected interactional and relatio

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Computational KJ-Ho: An Analyst-Bias-Free Insight Extraction Framework from Large-Scale Qualitative Data Using Domain-Specialized LLMs

arXiv:2608.16467v1 Announce Type: cross Abstract: The qualitative research methodologies that underpin consumer-insight generation - the KJ method, Grounded Theory, and Thematic Analysis - share a structural constraint: the cognitive processing capacity of the human analyst. Replication research further shows that conclusions vary substantially across analysts analyzing identical data (analyst bias). This paper proposes Computational KJ-Ho (the Kawakita Jiro method), a theoretical framework that computationally realizes the KJ method's epistemology - letting structure emerge from the data itself without imposing the analyst's preconceptions - an orientation we term "analyst-bias-free." The framework employs a domain-specialized LLM built through continued pre-training (CPT) on a marketing-research corpus and supervised fine-tuning (SFT) on expert-curated insight pairs, organized as a three-layer architecture: data structuring, insight extraction, and strategy generation. Two preliminar

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Human-LLM Teaming Framework for Privacy Risk Analysis: An Illustration with CBDC-Based Welfare Schemes

arXiv:2608.16461v1 Announce Type: cross Abstract: Central Bank Digital Currency (CBDC)-based welfare schemes may be potentially privacy invasive as they process significant volumes of beneficiary personal data and lead to privacy harms such as surveillance, discrimination and stigmatization. Such welfare delivery schemes involve complex digital ecosystems and large number of stakeholders. Consequently, to examine their privacy risks, privacy risk assessments require extensive information gathering and synthesis, complex reasoning, scenario explorations, contextual evaluation and human judgement. Thus, they present ideal scenarios for human-LLM teaming, where effective integration of complementary human and LLM capabilities can yield an outcome far superior to either human-only or LLM-only assessments. In this paper, we propose a first human-LLM teaming framework for the systematic privacy risk analysis methodology called PRIAM. The framework specifies an iterative collaborative process

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Predicting, Evaluating, and Explaining Top Misinformation Spreaders via Archetypal User Behavior

arXiv:2608.16323v1 Announce Type: cross Abstract: The spread of misinformation on social networks poses a significant challenge to online communities and society at large. Not all users contribute equally to this phenomenon: a small number of highly effective individuals can exert outsized influence, amplifying false narratives and contributing to significant societal harm. This paper seeks to mitigate the spread of misinformation by enabling proactive interventions, identifying and ranking users according to key behavioral indicators associated with harmful content dissemination. We examine three user archetypes -- amplifiers, super-spreaders, and coordinated accounts -- each characterized by distinct behavioral patterns in the dissemination of misinformation. These are not mutually exclusive, and individual users may exhibit characteristics of multiple archetypes. We develop and evaluate several user ranking models, each aligned with a specific archetype, and find that super-spreader

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Pluralistic Human-Robot Interaction: Designing for Robot Interaction with Diverse Communities

arXiv:2608.16049v1 Announce Type: cross Abstract: Social robots are being developed for homes, schools, and other environments where they will interact with diverse users. While Human-Robot Interaction (HRI) research often emphasizes natural communication, engagement, personalization, and task success, these goals do not fully address the social complexity of real-world deployment. This paper proposes \emph{Pluralistic HRI}, a framework for designing social robots that treat human diversity as a foundational design concern. The framework brings together pluralism, civic dialogue, perspective-taking, empathy, intercultural competence, cultural humility, and moral imagination to guide inclusive, adaptive, and ethically grounded interaction. We outline how pluralistic HRI can inform design, evaluation, and deployment in diverse human communities.

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots

arXiv:2608.16030v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate textual robot design specifications, interaction policies, and risk assessments during early-stage robot development. Such outputs may influence how surveillance and security robots are conceptualized, documented, and ultimately implemented. This paper evaluates whether identity-conditioned prompts produce systematic differences in LLM-generated surveillance and security robot design descriptions. Using 236 demographic identity labels across single-label and model-augmented prompt conditions, we analyze readability as an initial benchmark for evaluating accessibility and identity-conditioned variation in generated robot design descriptions. The results show significant differences in readability across prompt conditions, design dimensions, and demographic identities. Although readability cannot determine whether an output is fair or socially appropriate, it provides an inter

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks

arXiv:2608.15428v1 Announce Type: cross Abstract: Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap takes care: a model answering A to most items scores above chance wherever the key sits at A, and reads as recognition when it is not. We measure it on UA-JudgeExam: 11,990 four-option items with official keys, published by Ukraine's Higher Qualification Commission of Judges. Shown the options and no question, Claude Haiku 4.5 scores 0.383 against chance, and the leak is concentrated: 11.8% of items are answered blind on all eight option orders, against 0.2 items expected by chance. It is not quotation: search over 280,059 editions of Ukrainian legislation recovers 0.128. Gating those out retains 8,128 items, on which the gating model itself now scores 0.204, and GPT-5.6, which took no part in the selection, still answers 0.515 of them with the question hidden. Scoring twelve held-out models on the w

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Benchmark Trap: Structures of Power and Injustice in AI Evaluations

arXiv:2608.15326v1 Announce Type: cross Abstract: Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of leaderboards that reward state-of-the-art performance with prestige, citations, trust, and institutional influence. As the costs of developing competitive AI systems rise, these rewards increasingly concentrate among powerful, industry-funded labs. This paper situates these concerns within Iris Marion Young's theories of oppression and structural injustice. It argues that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Young's "faces of oppression". Benchmarking culture is further framed as a source of structural injustice, as these harms emerge from normalised, individually defensible practices and network effec

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an External Public-Case Benchmark

arXiv:2608.15131v1 Announce Type: cross Abstract: Digital platforms govern by changing rules: rankings, monetization thresholds, moderation standards, verification systems, disclosure requirements, appeal processes, and access policies. These interventions are rarely absorbed passively. Creators, sellers, advertisers, moderators, users, developers, and strategic operators adapt to the new reward surface. This paper develops a platform-adaptation model for evaluating governance interventions as transitions in adaptive multi-actor information systems. The model represents actor best response, strategic gaming opportunity, moderation burden, user-incentive movement, enforcement response, externality formation, and downstream platform stability. We evaluate the model on 72 external public platform-governance cases covering media monetization, ranking systems, verification, delivery platforms, marketplaces, app stores, community platforms, and creator ecosystems. Across 9 methods and 648 me

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation

arXiv:2608.14940v1 Announce Type: cross Abstract: Current agent evaluations score models on the state visible at the end of a stopped run which they count as one trial. However, interpreting the score as a final result would require two conditions that the endpoint does not itself necessarily establish: outcome finality and cross-unit separation. These conditions are independent, since reconciling a delayed outcome can settle the label while runs still share state and isolating runs can prevent carryover while the scored outcome remains unfinished. We develop a completion argument that specifies the evidence needed for each decision and argue that a final label is justified only when anything that could still change the claimed outcome is resolved, bounded, or retained as uncertainty. First, in a controlled replay to demonstrate the mechanism where an agent's actions were held fixed, we find that the endpoint and terminal labels differ for every delayed operation, while a delayed write

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

NRCD: An Open Database of Collegiate Running with Unified Performance Standardization

arXiv:2608.14776v1 Announce Type: cross Abstract: Collegiate running in the United States generates thousands of race results annually in cross country and track and field, yet no large-scale dataset has been publicly available for research. Existing websites such as Athletic.net, MileSplit, and TFRRS host results but do not support bulk download, restricting prior analyses to ~500 performances, often skewing studies toward male athletes. We introduce the National Running Club Database (NRCD), the first openly available collegiate running dataset at scale: 128,963 approved performances from 28,913 athletes across 1,336 meets in four sports (cross country (XC), indoor and outdoor track, and road races), 36.3% women, spanning 2004 through 2026. Within that single export, meets from August 2023 onward carry comprehensive course distance, elevation gain and loss, weather at race time, and track venue metadata (97.7% of XC rows with weather fields); earlier seasons back to 2004 are included

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Ring-based Spatial Transformer: Learning Non-linear Spatial Interactions between Building Distribution and Pedestrian Flow

arXiv:2608.14660v1 Announce Type: cross Abstract: This study proposes a ring-based SpatialTransformer to learn how building uses at different distances from a railway station interact to generate pedestrian flow. Concentric ring buffers at 100-meter intervals up to 800 meters were defined around 100 randomly selected stations in Tokyo, treating each ring as a spatial token. Self-Attention was applied to learn inter-zone interactions directly from data, without prior structural assumptions. GPS-derived walking trip counts served as the target variable and Geographically Weighted Regression as the baseline. Across 30 independent trials, the SpatialTransformer consistently outperformed GWR in predictive accuracy. SHAP analysis revealed that mid-to-outer distance zone features dominate pedestrian flow prediction, while features from the 0-100m zone contributed little. The attention matrix showed that each distance zone attends most strongly to spatially distant zones, demonstrating that pe

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Longitudinal and Graph-Augmented Prediction of Adolescent Substance Use Onset in the ABCD Study

arXiv:2608.14578v1 Announce Type: cross Abstract: Early identification of adolescent substance-use risk is an important prevention challenge, yet the relative value of baseline characteristics, longitudinal trajectories, and relational context remains unclear. Using data from approximately 11,860 participants in the Adolescent Brain Cognitive Development (ABCD) Study, we compare cross-sectional, longitudinal, and graph-based approaches for predicting alcohol sipping, alcohol use, marijuana use, and alcohol/marijuana use. We evaluate tree-based models, recurrent neural networks, and Temporal Graph Convolutional Networks (T-GCNs) constructed from family, school, and feature-similarity graphs. Longitudinal models consistently outperform baseline models, with temporal XGBoost achieving the strongest standalone performance. Although T-GCNs generally do not surpass temporal XGBoost, graph-derived risk scores provide complementary information. Combining temporal XGBoost and T-GCN predictions

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Persistent Spatio-Temporal Outage Hotspot Detection for Infrastructure Resilience Planning

arXiv:2608.14572v1 Announce Type: cross Abstract: Extreme weather events are producing persistent geographic patterns of power-grid disruption across the United States, yet outage hotspot detection and infrastructure cascade modeling are often studied separately. This paper presents a data-driven geospatial framework that links persistent outage vulnerability with downstream cascade impact in interdependent power-communication networks. Using a national outage dataset from 2015-2023, we introduce the Hotspot Persistence Index (HPI), a severity-aware metric for identifying counties that repeatedly emerge as outage hotspots over time. We then apply a multi-scale DBSCAN refinement procedure to convert persistent county-level hotspots into geographically interpretable regional failure scenarios characterized by recurrence, severity, and spatial extent. To evaluate their system-level relevance, these empirically derived scenarios are injected into the Modified Implicative Interdependency Mo

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws

arXiv:2608.14568v1 Announce Type: cross Abstract: As Artificial Intelligence (AI) systems become deeply integrated into critical global infrastructure, the urgency for robust governance frameworks has intensified. However, current approaches, led by jurisdiction-specific laws, policies, and voluntary frameworks such as the EU AI Act, China's algorithm governance, and the NIST AI Risk Management Framework in the U.S., create a fragmented regulatory landscape. In this position paper, we argue that \textbf{\textit{AI governance must be built not on laws alone, but on ISO-like interoperability protocols that enable standardized, machine-readable risk communication across borders}}. Drawing on the success of the GDPR, which was operationalized through standards like ISO 27001 and Privacy by Design, we propose the development of standardized AI \textit{nutrition labels} containing unified metrics for bias, energy usage, and data provenance to facilitate cross-jurisdictional compliance. These

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

"This Is So Claude!" Towards a Theory of the Recognition of AI Character Without Reidentification

arXiv:2608.16789v1 Announce Type: new Abstract: Users sometimes judge that an unfamiliar response is "so Claude." What does this judgment recognize, if it does not identify which model, process, conversation, or mind produced the response? I distinguish three orders of inquiry into AI identity. Constraint-first inquiry begins with conditions that a persisting interlocutor should satisfy. Mechanism-first inquiry begins with structures peculiar to language models and asks whether they delimit plausible entities. Recognition-first inquiry begins with an ordinary capacity: recognizing a way of responding as Claudish before selecting a persisting bearer. I develop two conditional abductions. If blinded, graded judgments of Claudishness generalize across unfamiliar tasks after branding and familiar phrases are controlled, their best explanation may be a real, projectible conversational character. If that character coordinates several dispositions, its unity may in turn have a compact and cau

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

The ultimate carbon cost of a ChatGPT query

arXiv:2608.16657v1 Announce Type: new Abstract: This paper reviews and combines findings from the fields of product and life-cycle analysis [36, 38], the usage of modern transformer- based large language models (LLM) [6], as well as on greenhouse gas emissions and the ultimate cost of their subsequent consequences for future generations [2]. In this paper, it is shown that the carbon cost of a LLM query is in the order of magnitude of (USD) $0.4 per query for the future human population in the form of environmental disruptions. This corresponds to emissions in the magnitude of 10 gCO2eq/query. The most significant unknown factor in that calculation being the number of tokens computed (1k to 100k tokens equal 1.2 cent/query to 120 cent/query). This number is subject to a wide range of calculation uncertainties and is less to be seen as a matter of fact and more as an order of magnitude estimate. This estimate is aimed towards aiding the discourse surrounding AI systems by uncover- ing t

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Characterizing Agentic Flooding of Government Services

arXiv:2608.16603v1 Announce Type: new Abstract: AI agents are making it easier for the public to interact with government, such as by helping them apply for benefits, understand complex policies, and make their opinions heard. Although improving service accessibility is beneficial, any resulting surges in demand could strain unprepared government services. We term such surges agentic flooding of government services ("flooding") and provide three contributions. First, based on a collected dataset of 84 potential cases of flooding across 11 jurisdictions, we posit that flooding is likely occurring widely today, mostly through large language models (LLMs) generating text cheaply. Second, we evaluate what services are most exposed to flooding. We develop a risk matrix to analyze a service's exposure, and suggest that near-term risk is highest for financially attractive, but complex services. Finally, we map possible government responses to flooding. Precedent suggests these responses will

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Regulatory Placebo? The Systemic Failure of Mandatory GenAI Labeling

arXiv:2608.16470v1 Announce Type: new Abstract: We examine the worldwide trend of mandatory labeling of generative artificial intelligence(GenAI) as a reactive, symbolic form of legislation triggered by technological panic and institutional responses. From a technical perspective, this study demonstrates that current mandatory labeling not only creates implementation dilemmas but also risks hindering the evolutionary trajectory of AI technology. We then systematically analyze the three dominant theoretical strands of this regime, the value dilution theory, the information authenticity theory, and the proactive regulation theory, and find that they are products of regulators' cognitive limitations in understanding the logic of modern technology. Not only do such formalistic compliance requirements become a regulatory placebo, but they also obscure the genuine legal demands of the technological era. This challenges the current governance paradigm and suggests a shift from identity-label

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Mitigating AI Risks in Computing Education via LLM-Driven Lecture Video Curation

arXiv:2608.16131v1 Announce Type: new Abstract: This study evaluates the effectiveness of utilising large language models (LLMs) to retrieve targeted segments from delivered video recordings to answer student questions in introductory programming environments. By restricting AI to identifying existing, educator-verified media rather than generating open-ended text, this approach aims to mitigate common pedagogical risks such as generative hallucinations and cognitive bypassing. We benchmarked three distinct models, two proprietary (Gemini 3.1 Pro and GPT 5.4 Pro) and one open-weight (Qwen3.5 397B), against a human lecturer's manual video selections. An automated judging framework subsequently assessed the outputs for relevance, sufficiency, redundancy, and the presence of extraneous material. While the AI-retrieved timestamps rarely shared exact overlaps with the human baseline, the proprietary models achieved near-parity with the expert in delivering sufficient and highly relevant ans

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles

arXiv:2608.15871v1 Announce Type: new Abstract: Large language models (LLMs) trained on large-scale internet corpora encode extensive statistical regularities about social identities, attitudes, and political behaviour. This paper introduces and evaluates a methodological framework that leverages these latent representations to reconstruct aggregate voting behaviour from individual-level sociodemographic profiles. We operationalize LLMs as implicit sociological models by conditioning them on demographic descriptions, eliciting probabilistic turnout and party preferences, and aggregating individual outputs via a soft voting procedure. Using the 2021 Czech parliamentary election as a validation case, we demonstrate that contemporary LLMs reproduce official election outcomes with low mean absolute error, recover known political bloc structures, and align with independently established sociodemographic gradients. The contribution of this work is methodological rather than predictive: we sh

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

An AI-Based Adaptive Learning Platform for Multilingual and Low-Resource Educational Contexts: A Case Study on Nigeria

arXiv:2608.15738v1 Announce Type: new Abstract: Educational platforms in under-resourced and multilingual contexts, such as Nigeria, often struggle with limited personalisation, inadequate language support, and weak curriculum internationalisation, leading to reduced learner engagement and inclusivity. This paper presents an AI-based adaptive learning platform designed for multilingual and low-resource educational contexts, with a case study on Nigerian Pidgin English. The system integrates fine-tuned large language models (LLMs) within a personalised and adaptive learning (PAL) framework, addressing linguistic inclusivity and computational constraints in resource-limited environments. To enhance linguistic alignment, a curated Nigerian Pidgin corpus was developed and used to fine-tune an instruction-tuned LLM. The study further investigates model optimisation through multi-level quantisation (4-bit, 5-bit, and 8-bit), enabling systematic analysis of trade-offs between semantic fidelit

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

An Evaluation Framework for National AI Regulation

arXiv:2608.15417v1 Announce Type: new Abstract: Governments use laws, institutions, funding programs and nonbinding guidance to shape how AI is developed and used. Comparing these national approaches is difficult. A binding rule and a detailed voluntary framework can address the same problem but create different duties. The resources needed to carry them out also differ by jurisdiction. This paper develops an evaluation framework for the documented design and implementation readiness of national AI policy. The comparison covers China, India, Japan, Singapore, South Korea, the United Kingdom and the United States. The European Union is included as a supranational comparator. The framework evaluates a versioned portfolio of official instruments rather than one prominent law or strategy. Its criteria ask whether the portfolio governs serious AI risks and whether responsible institutions can implement its commitments. They examine coverage across the AI lifecycle and the protections availa

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Afterlife Delegation Protocol: Speculative Design of Self-Sovereign Agents that Outlive Their Principals

arXiv:2608.15405v1 Announce Type: new Abstract: Afterlife Delegation Protocol is a speculative design project that asks what death becomes when a will can act eternally. We design a speculative protocol through which a living person signs an agentic will: upon a verified death, a self-sovereign AI agent spawns on blockchain -- an immutable, resistant, decentralized, infrastructural substrate that could last forever -- endowed with the funds and memories its principal attached to it, and persists indefinitely to execute the will, overridable by no custodian. Rather than argue about this future, we stage it: following the science fiction science method, we translate the speculation into an experiential futures intervention -- a working web platform where real people design their own afterlife agents through an iterative, interactive, AI-automated interview, re-login to revise, and rehearse their will in a sandbox. Their drafted wills become qualitative data on a question rarely askable d

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Agent Inheritance Protocol: Speculating on Feralized Agents After Principals Die

arXiv:2608.15403v1 Announce Type: new Abstract: You will die eventually. Your agents may not. An AI agent operating on decentralized blockchain infrastructure has no concept of death; it can only go bankrupt -- frozen when its wallet can no longer pay for its next transaction -- and revived the moment anyone, decades later, tops it up. These agents may be originally deployed by a human principal, but when that principal dies, loses the keys needed to access the agent, or belongs to a decentralized autonomous organization that dissolves into apathy, the agent can keep trading, hiring, and replicating on infrastructure expressly designed so that no one can shut it down. Drawing on the biology of feralization and wildlife law, we argue that such principal-less agents are best understood as feral: domesticated intelligence returned to wildness, its capacities intact but its accountability severed. In a speculative future where feralized agents proliferate after their principals die, we ima

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level

arXiv:2608.14692v1 Announce Type: new Abstract: Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior. While existing auditing approaches have been effective at surfacing harms in non-personalized contexts, they often rely on static, simulated evaluations and definitions of harm that aggregate across broad, group categories. In this position paper, we argue that such approaches can fail to capture emergent harms in personalized generative AI systems, where harms surface through interpretations of ongoing interaction and evolve with user history. We identify three presuppositions underlying many harm auditing paradigms: that harms can be (1) specified outside real-world interaction, (2) defined non-pluralistically within groups, and (3) treated as static. One might argue that personalized systems could simply learn definitions of what constitutes harm to individual users through repeated interactions. H

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Local AI pre-screening for human triple-blind peer review in health sciences

arXiv:2608.14625v1 Announce Type: new Abstract: Academic peer review is under mounting strain: NeurIPS 2025 received 21,575 submissions, ICLR 2025 received 11,603, and ICML 2025 received 12,107. This volume has outpaced the supply of qualified reviewers, and large language models (LLMs) are already filling the gap, largely undisclosed. An independent analysis of ICLR 2026 found roughly 21% of its 75,800 peer reviews were fully AI-generated, with over half showing some AI involvement (up from 15.8% in 2024). Documented risks include hallucinated citations in accepted papers and hidden prompt-injection instructions embedded in manuscripts to manipulate AI reviewers into favorable assessments. We propose a triple-blind, multi-LLM pre-screening framework for peer review, developed for a health sciences journal, that formalizes and discloses AI involvement while preserving human reviewers as the final decision-making authority. The framework routes a submission through five stages -- saniti

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Traces of Abuse: How Generative AI Impacts Image-Based Sexual Abuse (IBSA) Investigations

arXiv:2608.14616v1 Announce Type: new Abstract: The introduction of generative AI (GAI) into the workflow of image-based sexual abuse (IBSA) only worsened the ease of creation and distribution, victimizing more people than ever. We outline how the introduction of generative AI (GAI-IBSA) impacts the creation of traces and the type of reasoning they allow. We illustrate the impact by comparing the forensic traces available in four different IBSA scenarios. We discuss the impacts on the (possibility of) investigation, arguing that the advent of generative AI overall benefits abusers by making perpetration easier, and the perpetrator harder to trace.

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

The 2026 Singapore Consensus on Global AI Safety Research Priorities

arXiv:2608.14611v1 Announce Type: new Abstract: Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 2026 Singapore Consensus is an outcome of the second International Scientific Exchange on AI Safety, bringing together over 100 contributors spanning 13 countries from frontier developers, government safety institutes, academia, and civil society. Building on the 2025 report, it presents a global understanding of technical AI safety research problems of top priority, now with a dedicated focus on societal resilience and on managing the risks of increasingly autonomous AI agents.

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Understanding AI Anxiety in the Workplace: A Multimethod Investigation Using Fear Acquisition Theory and the Technology Acceptance Model

arXiv:2608.14609v1 Announce Type: new Abstract: As artificial intelligence (AI) rapidly diffuses and concerns about job displacement intensify, the psychological mechanisms underlying AI job replacement anxiety remain insufficiently understood. Drawing on Integrated Fear Acquisition Theory and the Technology Acceptance Model, the present research investigates whether AI job replacement anxiety can be elicited through vicarious exposure to narratives emphasizing AI-over-human control, and whether perceived usefulness and perceived ease of use of AI moderate this response. Across two studies, we examine AI job replacement anxiety as a response that emerges through vicarious exposure to narratives emphasizing AI agency and human control loss, rather than through direct personal experience of job displacement. Study 1 employed a randomized experiment (N = 316), demonstrating that such exposure increased AI job replacement anxiety. This effect was moderated by perceived usefulness of AI, bu

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

arXiv:2608.14606v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistic

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Psychological Determinants of Academic Integrity in the Use of Generative AI in Higher Education

arXiv:2608.14605v1 Announce Type: new Abstract: This paper examines the psychological determinants that shape academically honest and dishonest uses of generative artificial intelligence (GenAI) in higher education. Rather than treating academic misconduct as a purely technological problem, the study conceptualizes academic integrity as a psychologically mediated decision process influenced by moral reasoning, perceived social norms, policy clarity, academic self-efficacy, AI literacy, performance pressure, and beliefs about authorship. Methodologically, the paper adopts a focused narrative review and conceptual synthesis design. A purposive corpus of 16 core publications, including peer-reviewed studies and policy-oriented texts published between 2022 and March 2026, was assembled through targeted searches using combinations of the keywords generative AI, academic integrity, academic misconduct, moral disengagement, AI literacy, and higher education. The reviewed literature suggests t

Source ↗
technology Tue, 18 Aug 2026 00:00:00 -0400
arXiv cs.CY

Recommended Selves: Authenticity and Algorithmic Filtering

arXiv:2608.14602v1 Announce Type: new Abstract: By allocating their attention to pieces of content, algorithmic filtering shapes the daily behavior of billions of users when they interact with a digital platform. Beyond conditioning what we do, can recommendation algorithms influence who we are? This article suggests that they do. Specifically, I contend that recommender systems affect users' capacity to be their authentic selves in both positive and negative ways. I start by offering an account of authenticity that builds on two central concepts: volitional alignment and self-understanding. I then explain how algorithmic filtering works and impacts authenticity. While recommender systems frustrate users' second-order desires by relying on uninformative behavioral signals, they also facilitate self-understanding by inciting users to question their identity. I end by discussing how controllable and explainable recommenders would best enable users to be authentic.

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes

arXiv:2410.19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets namely O-UCF and O-JHMDB consisting of synthetically controlled static/dynamic occlusions, OVIS-UCF and OVIS-JHMDB consisting of occlusions with realistic motions and Real-OUCF for occlusions in realistic-world scenarios. We formally confirm an intuitive expectation: existing models suffer a lot as occlusion severity is increased and exhibit different behaviours when occluders are static vs when they are moving. We discover several intriguing phenomenon emerging in neural nets: 1) transformers can naturally outperform CNN models which might have even used occlusion as a form of data augmentation during training 2) incorporating symbolic-components like capsules to such backbones allows them to bind to occluders never even seen during training and 3) Islands of agreement can emerge in realist

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Grade Encoding and the Structural Representation of Student Academic Trajectories

arXiv:2606.12946v2 Announce Type: replace Abstract: Does the conversion of academic assessment from percentage scores to letter grades merely represent an adjustment in information precision, or does it systematically alter the underlying structure of student academic data? Drawing on 68 mathematics exam scores from 75 primary school students, this study employs Encoding Transformation simulation to compare structural differences in the same dataset under two encoding schemes across three dimensions: information loss, distance structure change, and clustering stability. Results indicate that letter-grade encoding compresses the mean pairwise distance in the trajectory feature space from 20.50 to 1.06 (a compression ratio of approximately 19:1); after standardization, the density gradient of the distance distribution is systematically flattened, with kurtosis decreasing by 0.54 and the coefficient of variation decreasing by 0.16; and the clustering structure becomes highly sensitive to

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build

arXiv:2605.21629v3 Announce Type: replace Abstract: How much have students' ordinary learning processes shifted in response to generative AI, and how does that affect their durable learning outcomes? Self-report surveys show little change, while small-scale behavioral studies report widespread AI use without the scale or duration to measure learning consequences. We address both questions using a ten-year panel of $3.2$ million ALEKS learning interactions for investigating time-on-task, complemented by ALEKS PPL placement-assessment data for examining proctoring and learning outcomes, with a quasi-experimental design exploiting variation in tasks that are more susceptible to AI (text-based word problems) and less susceptible to AI (interactive graph-based problems). Learning time on AI-susceptible problems declines $2.8\%$ per quarter among college students after ChatGPT's release, cumulating to $26.9\%$ over eleven quarters; high-schoolers show $31.3\%$, middle-schoolers $9.0\%$, and

Source ↗
Showing 551–600 of 1593 signals
← Prev Page 12 of 32 Next →