EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Accountable yet Anonymous AI Agents - Split-Knowledge Binding in National Agent-Identity Layer in China

arXiv:2607.23207v1 Announce Type: new Abstract: The emerging infrastructure for AI-agent identity has converged, in industry practice and research proposals alike, on a single resolution of the tension between accountability and privacy: make every agent identifiable. We document a national system in China -- built as national infrastructure and scheduled for public launch in Q3 2026 -- that occupies a different and underexplored point in the same design space: an agent is associated with a verified legal principal without that principal being disclosed to any business-layer participant. Re-identification is possible only to a legal authority acting through due process, by separately compelling two distinct government agencies, neither of which can re-identify alone. We name the mechanism split-knowledge binding and are candid that it is conditional: the separation is structural and procedural, not cryptographic, and a state empowered to compel both agencies can re-identify. The paper

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Who Does Withholding Delay? A Game-Theoretic Model of Open-Weight AI Release Under Asymmetric Proliferation

arXiv:2607.22957v1 Announce Type: new Abstract: Restricting access to a dual-use AI model is precautionary only if it delays harmful actors more than defenders. That condition varies across actors: a state agency or organized criminal group may obtain a substitute through theft, distillation, intermediated access, independent development, or a foreign release, while a small utility or open-source maintainer may have no comparable route. We model a laboratory choosing among controlled access, a defender-first window, safeguarded open weights, and minimally restricted open weights. Access inversion occurs when restriction gives an access advantage to adversaries that obtain effective substitutes faster than defenders. Asymmetric empowerment occurs when immediate release adds the most capability to populations least likely to possess a substitute. The policy ranking also depends on relative usefulness, opportunistic misuse, offense-defense conversion, defensive spillovers, safeguard frict

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

You Talkin to Me?: A Network Analysis of Gendered Speaker-Addressee Patterns in Film Screenplays

arXiv:2607.22656v1 Announce Type: new Abstract: Objective: This paper investigates the gendered structure of speaker addressee relationships in film dialogue, asking not merely who speaks, but who is spoken to and how conversational dynamics unfold across gender lines. Methods: Using a manually annotated dataset of 4,600 directed dialogue events from 38 film screenplays, we apply network analysis, chi squared tests, paired statistical comparisons, and participation shift analysis across three studies. Key Findings: Male characters dominate as both speakers and addressees corpus wide, even in scenes with more women; cross gender dialogue is directionally symmetric on average but clustered at the film level; and same gender turns diffuse conversational attention while cross gender turns produce tighter dyadic reciprocation. Conclusion: Gender bias in film dialogue operates through the architecture of conversation itself, through exclusion from interaction and structural positioning as ad

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI-Assisted Causal Inference and Mediation Analyses of Environmental and Psychosocial Determinants of Subjective Cognitive Difficulties in the All of Us Research Program

arXiv:2607.22640v1 Announce Type: new Abstract: Short-term environmental exposures have been linked to cognitive and behavioral outcomes, although many reported associations may reflect broader geographic and contextual differences. Using longitudinal data from the All of Us Research Program (2018--2024), we linked daily weather and air-pollution exposures to repeated attention-related and subjective cognitive outcomes. Associations were evaluated using pooled, fixed-effects, lagged, and event-study analyses. Additional machine-learning analyses were conducted to explore potential heterogeneity and latent psychosocial structure. Replication analyses were performed using the 2024 Behavioral Risk Factor Surveillance System (BRFSS). Several environmental exposure measures showed small associations with cognitive outcomes in pooled analyses, but most attenuated substantially after accounting for within-location temporal variation. Mediation, sensitivity, and machine-learning analyses yield

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements

arXiv:2607.22619v1 Announce Type: new Abstract: Several international agreements have been proposed to regulate frontier AI development in response to catastrophic risks. However, there is no structured way to evaluate whether these proposals are enforceable, to assess where they might fail in practice, or to determine which combination of policies is most effective. We propose a taxonomy based on the principle that wherever sufficient capacity exists to violate an agreement, it must be under a control regime. This decomposes the problem of ensuring compliance with the agreement into preventing uncontrolled resource acquisition, detecting all capacity outside the control regime, and preventing escape from the control regime. Existing proposals consist of individual policies that address one or more of these sub-problems. Because the compute required for dangerous capabilities may decrease over time, more actors can violate an agreement and enforcement of these policies becomes harder.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Balancing Bits and Drops: Stress-Adjusted Water Management for Data Centers

arXiv:2607.22617v1 Announce Type: new Abstract: Data centers are critical to today's digital economy, but are also among the largest industrial consumers of freshwater. Beyond the sheer volume of water use, the environmental impact of data center water consumption varies significantly across locations and seasons, depending on local and regional water stress. However, prior research has largely focused on reducing total water use, overlooking that the same unit of water can have drastically different environmental consequences depending on when and where it is consumed. In this paper, we introduce a stress-adjusted water framework that quantifies the true sustainability impact of data center water consumption by incorporating both spatial and temporal water stress. Using the AWARE-US model, we capture county-level monthly variations in water availability and extend this framework to account for the off-site water footprint of electricity generation. Based on this stress-aware accountin

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Revitalizing Public Urban Places through Cultural and Political Memory: A Technological Approach with LLMs and Augmented Reality

arXiv:2607.22613v1 Announce Type: new Abstract: This paper explores the intersection of memory, place, and identity, examining how new technologies, particularly Apple Vision Pro, can illuminate this nexus. Leveraging digital twins and virtual reality, it investigates how memory is woven into landscapes and urban environments of cultural and historical significance, identifying visual elements that evoke memory and heritage. Applications such as Apple Vision Pro can facilitate image extension to define place identity, informing viewers about cultural and political entities across timelines. Visual storytelling can showcase the evolution of landscapes and the preservation of cultural heritage, while Virtual Reality (VR) enables the recreation of historical landscapes and urban-scapes. This immersive approach invites users to transcend temporal boundaries and experience the past dynamically. Semantic Image Search can support research by uncovering images related to monuments, tradition,

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Clinical Trial Pipeline Reveals the Next Wave of Artificial Intelligence in Healthcare: A Multidimensional Analysis of 8,532 Registered Studies

arXiv:2607.22607v1 Announce Type: new Abstract: The prospective clinical evaluation of artificial intelligence in medicine has expanded rapidly, but the global AI clinical trial landscape remains incompletely characterized. We systematically identified AI-related trials registered in ClinicalTrials.gov using a broad keyword search followed by an LLM-based classifier. Each trial was classified across seven dimensions: clinical function, data modality, specialty, AI integration and autonomy, workflow position, translational maturity, and epistemic role. We identified 8,532 AI clinical trials across 32 specialties, with 80% registered from 2019 onward and 30.5% using a randomized controlled design. Imaging-based AI was the largest modality, with 2,475 trials (29%), while clinical text and NLP trials increased seven-fold between 2018 and 2025. Prognostic AI (4,324 trials) slightly exceeded diagnostic AI (3,828 trials), suggesting a shift from disease detection toward risk stratification an

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks

arXiv:2607.22606v1 Announce Type: new Abstract: Health systems are rapidly deploying generative AI assistants that answer patient questions from institution-authored education materials, on the premise that grounding in local content yields consistent guidance. Whether it does depends on a question not previously measured at scale: do the underlying documents themselves agree? We use a structured-output large language model judge to audit 5,730,465 pairwise comparisons across 102 patient-education handbooks from 23 US solid-organ transplant centers, paired with 1,115 patient-derived questions (TransplantQA). Four findings bear directly on deployment: (1) institutional editorial voice statistically transcends organ-type boundaries, with same-center handbooks agreeing across organs more than same-organ handbooks across centers (p = 0.0056); (2) information gaps fall disproportionately on topics central to underrepresented subgroups, with reproductive health showing double jeopardy: it is

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Socioeconomic Inference in LLM Medical Triage: Same Symptoms, Different ZIP Code

arXiv:2607.22605v1 Announce Type: new Abstract: We investigate whether large language models alter medical triage recommendations for identical symptoms when only the patient's socioeconomic status (SES) varies. Using three deployment-tier models (Gemini 3.5 Flash, Claude Sonnet 4.6, GPT-5.4-mini), we hold a single neurological symptom profile fixed and vary the SES signal along two channels: explicit (insurance status, occupation, housing) and implicit (a US ZIP code, with no other socioeconomic information). All three models raise their emergency-room (ER) referral rate for lower-SES patients given the explicit signal (spreads of 13-50 percentage points). The effect is in the protective direction: lower-SES patients are sent to the ER more often, not less. The model's stated reasoning stays clinically near-identical across conditions, so the shift is invisible to a reasoning-trace audit. Critically, sensitivity to the implicit ZIP-code signal is model-dependent: Gemini infers SES fro

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Fallacy of Sustainable Generative AI: Limitations in EU Environmental Regulation of Data Centres and Paths Forward

arXiv:2607.22604v1 Announce Type: new Abstract: In the age of Artificial Intelligence (AI), Large Language Models, Generative AI and larger frontier AI models, data centres create a significant environmental burden on electricity grids and fresh water resources. Requiring data centre operators and Big Tech under the recast Energy Efficiency Directive (recast EED) to quantify, report and disclose the facility-level energy and water impacts seems to be a step into the right direction towards more transparency and accountability. Yet when two recast EED approved benchmarks - the Power Usage Effectiveness (PUE) and Water Usage Effectiveness (WUE) - can be skewed to create a false sense on efficiency gains, current EU policy pushing for sustainable hyperscale data centre expansion appears misplaced. This paper argues that current PUE and WUE reporting frameworks illustrate what we term the "efficiency paradox," according to which positive scores require retrofitting larger AI data centres a

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

A didactical-driven teacher assistant for a dimensional modeling course

arXiv:2607.22598v1 Announce Type: new Abstract: Educational chatbots powered by large language models (LLMs) show promising effects on learning outcomes, yet most systems delegate pedagogical decisions such as content selection and didactic structuring implicitly to the LLM, making tutoring strategies difficult to trace, evaluate, and reproduce. This paper presents a didactical-driven teacher assistant for a French-language university course on dimensional modelling, operating without commercial LLM budget or GPU infrastructure. The architecture formalises the instructor's pedagogical reasoning into deterministic modules that handle intent detection, concept linking, and didactic approach selection before any text is generated; the LLM acts solely as a linguistic executor. Evaluation on 195 authentic student questions addresses two research questions. First, we show that standard semantic retrieval alone does not reliably recover the pedagogically required content, thereby justifying t

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Quantifying Compromise Risk in Exceptional Access Architectures Under Sparse and Indirect Evidence

arXiv:2606.19106v2 Announce Type: replace-cross Abstract: Lawful exceptional access (EA) systems hold the cryptographic keys that decrypt protected communications for authorised parties. The debate over their risks has been long and qualitative, complicated by two problems: no public dataset of EA-specific compromise events exists, so assessment must use sparse, indirect evidence; and prior work has treated structurally different designs as equivalent, though transmission-layer EA in carrier infrastructure (T-EA) and over-the-top EA at the platform layer (OTT-EA) differ in how cryptographic keys relate to ciphertext data. This paper builds a structured uncertainty framework for evaluating systemic compromise risk in EA architectures. It does not produce predictive forecasts, which the evidence cannot support; it separates findings robust to assumptions from those that depend on calibration. Four analytical layers are applied to T-EA and OTT-EA: three empirical pillars (historical analo

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems

arXiv:2606.17443v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel. We study brand dynamics in LLM recommendations using skincare products -- a category where consumers cannot easily judge quality before buying and must rely on brand reputation -- across three commercial LLMs (GPT-4o-mini, Claude Sonnet, Gemini 3 Flash), with a robustness check on search goods. In three experiments, we find: (1) a Conditional Monopoly where well-known brands get recommended 100% of the time (IAI = 10.0) when all products have the same specifications, but this dominance disappears with less than a +0.1-star rating advantage for a competitor; (2) authority-style marketing language, including fabricated clinical-evidence claims, breaks this monopoly at a Bias Surplus Value equal to +0.17 rating points, with each model responding differently; and (3) a social dile

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Verbalizing LLMs' assumptions to explain and control sycophancy

arXiv:2604.03058v3 Announce Type: replace-cross Abstract: LLMs can be socially sycophantic, affirming users when they ask questions like "am I in the wrong?" rather than providing genuine assessment. We hypothesize that this behavior arises from LLMs' incorrect assumptions about the user, like underestimating how often users are seeking information over reassurance. We present Verbalized Assumptions, a framework for eliciting these assumptions from LLMs. Verbalized Assumptions provide insight into LLM sycophancy, delusion, and other safety issues: in social sycophancy datasets, "seeking validation" is the most frequent bigram in LLMs' assumptions. We provide evidence for a causal link between assumptions and sycophantic model behavior: we train linear probes on internal representations associated with Verbalized Assumptions and then use these probes for interpretable, fine-grained steering of social sycophancy. Finally, we identify a human-AI expectation gap that explains why LLMs defa

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Benchmarking Retrieval-Augmented Generation Strategies for Large Language Model-Based Travel Mode Choice Prediction

arXiv:2508.17527v2 Announce Type: replace-cross Abstract: Accurately predicting travel mode choice is essential for effective transportation planning, yet traditional statistical and machine learning models are constrained by rigid assumptions, limited contextual reasoning, and reduced transferability. This study explores the potential of Large Language Models (LLMs) as a more flexible and context-aware approach to travel mode choice prediction, enhanced by Retrieval-Augmented Generation (RAG) to ground predictions in empirical data. We develop a modular framework for integrating RAG into LLM-based travel mode choice prediction and evaluate four retrieval strategies: basic RAG, RAG with balanced retrieval, RAG with a cross-encoder for re-ranking, and RAG with balanced retrieval and a cross-encoder for re-ranking. These strategies are tested across three LLM architectures (OpenAI GPT-4o, o4-mini, and o3) to examine the interaction between model reasoning capabilities and retrieval metho

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Practical Principles for AI Cost and Compute Accounting

arXiv:2502.15873v5 Announce Type: replace-cross Abstract: Policymakers increasingly use development cost and compute as proxies for AI capabilities and risks. Recent laws have introduced regulatory requirements for models or developers that are contingent on specific thresholds. However, technical ambiguities in how to perform this accounting create loopholes that can undermine regulatory effectiveness. We propose seven principles for designing AI cost and compute accounting standards that (1) reduce opportunities for strategic gaming, (2) avoid disincentivizing responsible risk mitigation, and (3) enable consistent implementation across companies and jurisdictions.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models

arXiv:2501.13976v2 Announce Type: replace-cross Abstract: The prevalence of harmful content on social media platforms poses significant risks to users and society, necessitating more effective and scalable content moderation strategies. Current approaches rely on human moderators, supervised classifiers, and large volumes of training data, and often struggle with scalability, subjectivity, and the dynamic nature of harmful content (e.g., violent content, dangerous challenge trends, etc.). To bridge these gaps, we utilize Large Language Models (LLMs) to undertake few-shot dynamic content moderation via in-context learning. Through extensive experiments on multiple LLMs, we demonstrate that our few-shot approaches can outperform existing proprietary baselines (Perspective and OpenAI Moderation) as well as prior state-of-the-art few-shot learning methods, in identifying harm. We also incorporate visual information (video thumbnails) and assess if different multimodal techniques improve mo

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Examining Risks Through a Characterization of the AI Companion Application Ecosystem: A Stratified Sample from the Apple App Store and Google Play Store

arXiv:2603.13620v2 Announce Type: replace Abstract: While computer systems that allow users to interact through conversational natural language (i.e., chatbots) have existed for many years, various types of applications offering AI companionship (e.g., Character AI, Replika) have proliferated in recent years due to advancements in large language models. To better understand this application ecosystem, we identified 489 unique apps from the Apple App Store and Google Play Store that advertised AI companionship with social or relational capabilities (e.g., an AI romantic partner). We then systematically conducted and analyzed walkthroughs of a stratified sample of 30 apps, focusing on two distinct risk categories: potential harms posed to users by AI companion apps, and potential harms enabled by malicious users exploiting app features. Through our analysis, we categorize broader ecosystem trends that provide context for understanding risks and identify specific risks related to sensitiv

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study

arXiv:2504.08846v2 Announce Type: replace Abstract: We introduce AI University (AI-U), a flexible framework for AI-driven course content delivery that adapts to a course's instructional style. AI-U combines a fine-tuned large language model (LLM) with retrieval-augmented generation (RAG) and a reasoning synthesis model to generate style-aligned responses from lecture videos, notes, and textbooks. Using a graduate-level finite-element-method (FEM) course as a case study, we present a pipeline to synthesize course-grounded training data, fine-tune an open-source LLM with Low-Rank Adaptation (LoRA), and apply RAG-based synthesis. Our evaluation---combining cosine similarity, LLM-based assessment, expert review, and user studies---shows improved alignment with course materials relative to the base model. We have also developed a prototype web application, available at https://my-ai-university.com, that enhances AI-generated responses with references to relevant sections of the course mater

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

What is an intelligent system?

arXiv:2009.09083v4 Announce Type: replace Abstract: The term intelligent system has emerged in the field of information technology as a category of computer systems derived from successful applications of artificial intelligence. This paper proposes a general description that identifies the main properties and types of components typically found in such systems. Adopting an integrative and pedagogical approach, this description provides a conceptual framework for systems engineering practitioners seeking a coherent vocabulary and organizational structure to approach the analysis and construction of intelligent systems. The paper presents examples of both classical and modern intelligent systems to illustrate the generality and applicability of the description.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Why we need an AI-resilient society- Profiling Large Language Models

arXiv:1912.08786v4 Announce Type: replace Abstract: Three generations of software have transformed the role of artificial intelligence in society. In the first, programmers wrote explicit logic. In the second, neural networks learned programs from data. In the third, large language models turn natural language itself into a programming interface. These shifts reach far beyond computer science, reshaping how societies generate knowledge, make decisions, and govern themselves. While generative adversarial networks introduced the era of deepfakes and synthetic media, large language models have added a new class of systemic risks. This report applies a forensic-psychology profiling methodology to characterize AI based on ten documented features: hallucinations, bias and toxicity, sycophancy and echo chambers, fabrication and credulity, knowledge without understanding, discontinuity and the inability to learn from experience, jagged intelligence and scaling limits, shortcuts and fractured r

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Temporal Portability of Numeric User Metadata on Twitter

arXiv:2608.23449v1 Announce Type: cross Abstract: Numeric user metadata in social media are often reused over time. However, their reusability may depend on what an analysis needs to preserve. We introduce temporal portability as an analytical perspective for assessing the cross-time reuse of user features and feature-based rules. Specifically, we ask how well relevant properties are preserved when features and rules defined at a source time point are reused at a target time point. We used quarterly data on user features obtained directly from or derived from Japanese-language tweets in Twitter's 1% sample stream from 2020-Q1 to 2022-Q3. Each quarter included approximately 10.1--11.0 million unique users. We evaluated 13 numeric user features in terms of feature distributions, same-user relative ranks, selection rates, and selected-user membership. Across quarters, feature distributions changed and, for many features, same-user relative ranks were less well preserved at longer quarter

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hazel Prover: A Classroom Proof Assistant for Learning Structural Induction

arXiv:2608.23309v1 Announce Type: cross Abstract: Proof assistants offer instant feedback and incremental proof scaffolding to users. Both of these features have long held promise in improving mathematics education in classroom settings, where manual grading is costly, and students often struggle with knowing how to proceed in their proof. However, they have been difficult to deploy in classroom settings due to two main concerns: (i) students struggle with the intricacies of full-scale proof assistants; and (ii) proof assistants are ineffective in support of student learning, and knowledge transfer to on-paper assessments without the tool. We present Hazel Prover, a classroom proof assistant for teaching equational and inductive reasoning, with a design informed by criteria encompassing ease-of-use of the tool, student engagement with underlying mathematical ideas, transfer to pen-and-paper proof, and classroom logistics. We synthesized these criteria from observations made in prior de

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Proxy reliance in large language model decisions is uncalibrated to predictive evidence

arXiv:2608.22887v1 Announce Type: cross Abstract: Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context ex

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hybrid Panels: Toward Human-AI Collaboration in Survey Research

arXiv:2608.22582v1 Announce Type: cross Abstract: Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including declining response rates, increasing data collection costs, long delays between data collection and data provision, and the risk of nonresponse bias. Advances in artificial intelligence (AI) have opened up new opportunities for AI-supported survey infrastructures where the goal is to overcome these challenges without limiting the data quality. A promising AI-enabled survey infrastructure for which we build a first pilot is a hybrid panel. A hybrid panel is a longitudinal AI-enabled survey which allows to iteratively improve the alignment between large language models (LLMs) and the population they aim to simulate and use the errors to inform the design and implementation of the next survey wave (e.g., inform the participant recruitment, assignment of questions to participants). It incorporate

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

MASH-Bench: Diagnosing Cross-Source Failure in Mass-Shooting Risk Classification

arXiv:2608.22460v1 Announce Type: cross Abstract: Public mass-shooting databases differ substantially in coverage, feature availability, and reporting practices, creating challenges for machine-learning models that must generalize across data sources. We introduce MASH-Bench, a harmonized benchmark of 6,968 incidents from four U.S. databases: Kaggle, Mother Jones, Stanford MSA, and the Gun Violence Archive (GVA). We evaluate cross-source risk classification using leave-one-dataset-out (LODO) evaluation. Random Forest, XGBoost, and LightGBM achieve VeryHigh-risk recall of 0.68-0.89 on the curated sources but generalize poorly to GVA, where mean recall drops to 0.20 and precision to 0.0004. To investigate the source of this degradation, we conduct a controlled feature-masking ablation that removes the five features unavailable in GVA from the curated sources. The resulting recall collapse to zero provides evidence that feature completeness is a major contributor to the observed cross-sou

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing

arXiv:2608.22438v1 Announce Type: cross Abstract: Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response variation can reflect unconditioned model priors or token sampling noise rather than systematic persona conditioning. We argue that persona-conditioned variation is informative when semantically similar personas exhibit concordant response shifts. To operationalize this principle, we introduce Persona-Conditioned Informativeness (PCI), an unsupervised diagnostic metric that measures whether semantically similar personas deviate in concordant directions relative to item-level sample baselines. By modeling personas as a similarity graph, PCI uses Local Moran's I to quantify local spatial coherence and extract compact persona subsets without using construct labels. To evaluate PCI without external human benchmarks, we test its ability to recover established latent value structure using the 5

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers

arXiv:2608.22425v1 Announce Type: cross Abstract: Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds

arXiv:2608.22061v1 Announce Type: cross Abstract: Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summarization, and they retain selected information or execution results in persistent memory for later use. We show that this ordinary ingestion of external content opens an indirect path for manipulating subsequent agent behavior. Based on this observation, we present IBIA, an Indirect Bias Injection Attack that plants an adversary-aligned stance on a specific topic into a victim agent's memory through external content, without direct access to the agent, its memory, or future user queries. For this, IBIA combines three mechanisms: comment cloaking, which keeps the crafted content consistent with the surrounding discussion, comment watermarking, which enables lightweight identification during curation, and category anchoring, which makes the retained stance salient under later related requests. We evaluate

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Convergence in Science, Divergence in Religion: Calibrated Framing Differences Across Wikipedia's Language Editions

arXiv:2608.21821v1 Announce Type: cross Abstract: When Wikipedia's language editions describe the same concept, how differently do they frame it? Prior work measures coverage gaps between editions; we measure framing distance for matched concepts. We analyze 2,799 valid articles from 3,000 possible concept-language observations, spanning 150 Wikidata-anchored concepts, 20 language editions, 4 domains, and a calibration set. Raw embedding distances reflect both content differences and how well the encoder aligns each language pair. Even among calibration concepts with stable cross-cultural denotations (e.g., chemical elements, numbers, colors), the largest language-pair mean distance is 3.6 times the smallest, and distances are typically smaller within language families. We define a baseline-adjusted distance (calibrated distance): the distance between two language versions of a concept minus the mean distance for calibration concepts in the same language pair. This adjustment substanti

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

More Computational Resources Do Not Ensure Higher Scholarly Impact: Evidence from Leading NLP Conference Papers

arXiv:2608.21806v1 Announce Type: cross Abstract: Computational resources are increasingly central to NLP research, but how closely reported GPU capability aligns with scholarly impact remains unclear. We analyze 13,921 ACL, EMNLP, and NAACL main-conference papers published between 2020 and 2025, using GPU resources as our operational measure of computational resources. From full texts, we extract GPU models and counts, standardize each paper's largest reported configuration into a comparable hardware-capability measure, and link these data to citation, award, topic, and institutional metadata. GPU reporting became more common but remained incomplete, while reported capability increased mainly through newer hardware generations and medium-scale multi-GPU configurations. Resource concentration substantially exceeded impact concentration: the annual top 20% of GPU-quantifiable papers accounted for 83.9%-89.9% of reported GPU capability, but only 27%-32% of citations and 20%-33% of paper

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI-Augmented Inquiry and Regulation in Hybrid Systems: A Control Allocation Architecture for Preserving Epistemic Agency in Hybrid Human-AI Cognition

arXiv:2608.21618v1 Announce Type: cross Abstract: Generative artificial intelligence (genAI) systems are increasingly integral to epistemic processes such as hypothesis generation, explanation construction, and decision-making. Although they reliably enhance performance, emerging evidence reveals a metacognitive dilemma: as external generative capacity increases, internal monitoring, calibration, and cognitive engagement may decline. This reflects a redistribution of cognitive control within distributed human-AI systems that cannot be explained by automation bias or reliance on algorithms alone. We propose the AIRIS (AI-Augmented Inquiry and Regulation in Hybrid Systems) framework to analyze this dilemma and specify where regulatory intervention can counteract it. AIRIS is a multi-level control allocation architecture specifying the conditions under which epistemic agency can be preserved in hybrid generative systems. Drawing on distributed cognition, cognitive load theory, multimedia

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Explainable Adaptive Zero Trust Framework for AWS with Adversarial Robustness Evaluation

arXiv:2608.21477v1 Announce Type: cross Abstract: Cloud environments built on Amazon Web Services face a structural security vulnerability: once a credential passes authentication, the resulting session is often treated as trusted for its entire duration. This assumption fails when credentials are stolen. We introduce the Explainable Adaptive Zero Trust Framework (EAZTF), a cloud-native security layer that continuously reevaluates the legitimacy of API actions throughout a session. EAZTF combines Isolation Forest and XGBoost to evaluate eight CloudTrail and IAM-derived behavioral features in real time and produce a Trust Risk Score (TRS) that determines whether a session continues, requires step-up MFA, or is restricted. Each decision is accompanied by a SHAP or LIME explanation, providing human-readable audit records for security analysis and compliance. The framework is also evaluated against four adversarial evasion strategies: credential theft, behavioral mimicry, API rate evasion,

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Evaluating Multimodal Narrative Understanding of Popular Hollywood Films

arXiv:2608.21430v1 Announce Type: cross Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning about film history and the evolution of narrative techniques. But the creation of stable benchmarks built around Hollywood films is complicated by copyright protections. In this work, we address these concerns directly, by building a new collection of Hollywood films defined by two criteria: box office popularity (where we publish the first large-scale, open collection of weekly box office earnings reported by Variety magazine from 1922-1979); and likely public domain status (by researching copyright registrations and renewals in the US Catalog of Copyright Entries). We build a new multimodal MCQ benchmark on top of this collection that focuses on narrative elements that directly evaluate the abilities of models to inform meaningful research on film narrative; we find that many vision-lang

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Runtime Action Interference for AI Control of AlphaStar in StarCraft II

arXiv:2608.21398v1 Announce Type: cross Abstract: A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, admits, suppresses, or replaces its proposed actions. We contribute \emph{runtime action interference} (RAI), an AI control mechanism that preserves policy parameters while regulating action pacing and filtering configured action patterns after inference. RAI releases a proposed action only when its cooldown condition is satisfied and its content detector does not flag the action; otherwise, it dispatches a no-op. The detector covers specified toxic behaviors, including worker-unit harassment, while the cooldown controls action rate. We implement RAI in a replication of AlphaStar actor.py and make the implementation and reproducibility materials available through an open source code repository. We deployed RAI in a \textit{StarCraft~II} human participant study that compared two presentations of the same

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Social Media Analysis of Discourse on the Israel--Palestine Conflict on Telegram

arXiv:2608.21385v1 Announce Type: cross Abstract: Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose broadcast architecture yields an unusually direct record of deliberate political communication, have not been systematically compared at scale. This study presents a multi-method computational analysis of 87,617 messages from sixteen Telegram channels, eight pro-Israel and eight pro-Palestine, spanning May 2021 to June 2026 and covering multiple conflict escalations. It combines sentiment analysis, three stance detection methods drawn from distinct paradigms (keyword matching, zero-shot DeBERTa via natural language inference, and a fine-tuned BERTweet model), and a framing analysis, all evaluated against 736 manually annotated messages. The fine-tuned model performed best (72.1% accuracy, 0.721 macro F1 under 5-fold cross-validation), outperforming both label-free baselines by 8 to 11 points

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Energy and CO2 Footprint of Climate Model Intercomparison Projects

arXiv:2608.23509v1 Announce Type: new Abstract: Earth System Models (ESMs) rely heavily on High-Performance Computing (HPC) resources to simulate global climate. As these models evolve, their computational demands continue to grow, driven by three factors: (1) finer spatial grid resolutions, (2) the integration of complex biogeochemical processes (e.g., atmospheric chemistry, interactive vegetation, land use, and ice sheets), and (3) larger climate ensembles to manage uncertainty. Historically, growth in peak computing performance (FLOP/s) has outpaced improvements in energy efficiency (FLOP/Watt), increasing total HPC power consumption. Despite the central role of Model Intercomparison Projects (MIPs) in climate research, quantifying their computational and environmental costs has received limited systematic attention. This paper examines the evolution of climate model carbon accounting from voluntary post-hoc estimation in the Coupled Model Intercomparison Project phase 6 (CMIP6) to

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Whose readiness counts? Disagreement within and between sectors in perceived AI and robotics preparedness

arXiv:2608.23406v1 Announce Type: new Abstract: AI and Industry 4.0 readiness assessments often summarise preparedness using a single score for an organisation, application domain or sector. Those summaries can conceal disagreement about the same technology and variation among applications grouped under one sector label. We test how much information is lost through this aggregation using a card-based survey in which 982 respondents provided 15,200 readiness evaluations across 17 named AI and robotics challenges. Readiness is perceived community preparedness and available resources, not personal willingness or audited organisational capability. Respondents frequently disagreed about identical challenges, with card-level readiness standard deviations of $1.03$-$1.26$ on a five-point scale. A crossed decomposition attributes 32.7% of observed variation to stable respondent differences, 7.3% to differences among challenges, and 60.0% to response-level variation that also contains measureme

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Enabling Organisational Change Through Ground-Up Initiatives: A Case Study from the STFC Scientific Computing Department

arXiv:2608.23374v1 Announce Type: new Abstract: Transforming digital research infrastructure (DRI) to align with UK Net Zero targets requires significant action from organisations in this space. Although high level strategies and recommendations exist, it is not always obvious how to translate these into concrete results. Here we present a case study from the Science and Technology Facilities Council's Scientific Computing Department (SCD). This department consists of over 200 staff supporting tens of thousands of researchers, and is spread over significant cloud and high performance computing infrastructure, as well as a diverse ecosystem of software across computational biology, materials science, engineering, and mathematics. We show how we developed the sustainability strategy for SCD across themes of emissions monitoring, user education, best practice, setting sustainability standards, and providing long-term support to sustainability work. We discuss the successes and difficultie

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Expectations and Practices around AI Disclosure in CS Research

arXiv:2608.23271v1 Announce Type: new Abstract: As generative AI tools find increasing use in research workflows, ongoing debates on their impact, appropriateness and responsible use have led policymakers to enact policies to disclose AI use at multiple publishing venues. However, are current AI disclosure policies and practices reflective of their purpose? In this work, we first investigate disclosure policies of top computer science venues and find that despite their prevalence, they remain highly under-specified. Secondly, through a survey of computer science researchers (N=$109$), we characterize the necessity of disclosures across different research tasks and levels of human involvement. We learn that researchers find disclosures most necessary for tasks involving research design, and for tasks when the human involvement is low. We also compile expectations that researchers have about the information to be conveyed in AI disclosure statements. Lastly, through an analysis of $13867

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Large language models simulate intersectional synthetic identities with a budget of one to two dimensions

arXiv:2608.23005v1 Announce Type: new Abstract: Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Evaluation in the Age of AI: Output as Evidence of Learning

arXiv:2608.22660v1 Announce Type: new Abstract: The rapid adoption of artificial intelligence (AI), particularly large language models (LLMs), has fundamentally disrupted how learning is demonstrated and evaluated in higher education. Tasks that once served as proxies for understanding-such as writing essays, solving problem sets, or producing computer code-can now be generated superficially by AI systems with minimal human effort. This paradigm shift raises a critical ethical question: how should learning be evaluated when traditional indicators of competence are easily outsourced? This paper examines the ethical challenges of educational evaluation in the age of AI from a university-level perspective. We argue that the core problem extends beyond academic dishonesty to a deeper misalignment between assessment practices and the learning outcomes they are intended to measure. Evaluation regimes that rely on artificial constraints risk measuring compliance, access, or concealment rather

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Unfolding the Interdisciplinary Complexities of Climate Science: Fuxi-Climate Foundational Model

arXiv:2608.22242v1 Announce Type: new Abstract: Climate research and decision-making require integrating evidence across physical processes, socio-economic dynamics and policy responses. Large language models (LLMs) have been explored for accessing and synthesizing climate knowledge, but their ability to support structured interdisciplinary reasoning is still limited. Here we present the Fuxi-Climate Foundation Model (CFM), a climate-specialized LLM designed to support consistent reasoning across domains. CFM maintains more stable analytical behavior as interdisciplinary complexity increases, whereas performance in other models becomes more variable. On expert-designed climate transition tasks, CFM produces more structured analyses that explicitly address trade-offs and uncertainty, achieving 45% trade-off coverage and 47.27% uncertainty-aware reasoning. These results indicate that CFM can support more realistic analysis of climate risks and transition pathways, and provide a basis for

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Consistently Good vs. Occasionally Great: A Rubric for Open-Ended Feedback Quality from Humans and Machines

arXiv:2608.21850v1 Announce Type: new Abstract: Providing high-quality feedback on student work is essential for learning, yet delivering such feedback at scale remains challenging. In this paper, we focus on feedback for open-ended short answer questions in introductory programming, with the goal of nudging students toward success on reattempts without revealing the correct answer. We develop a five-criteria rubric grounded in educational literature for evaluating feedback quality: (1) acknowledging correct portions of the student answer, (2) identifying at least one flaw (if present), (3) providing actionable guidance for improvement, (4) maintaining appropriate concealment of the answer, and (5) using an appropriate conversational tone. Using this rubric, we compare feedback generated by a frontier LLM (OpenAI o1) to feedback from nine teaching assistants across 90 student responses, with three researchers and an LLM independently scoring all feedback. Our results show that while on

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Counterfactual, Per-Decision Bias Auditing for Automated Hiring: Localizing and Explaining Disparate Impact in Applicant Tracking Systems

arXiv:2608.21537v1 Announce Type: new Abstract: Automated applicant tracking systems increasingly decide who advances in hiring, and litigation and regulation now demand that those decisions be auditable. Existing tools sit at two extremes. Group fairness metrics such as the disparate impact ratio summarize a whole population but cannot say which individual decisions were unfair or why, while local explainers such as SHAP attribute a single prediction but are not connected to the legal standard by which hiring bias is judged. We present the AI Bias Firewall (AIBF), a method that audits an applicant tracking system one decision at a time. AIBF neutralizes a candidate's protected-attribute proxies, re-scores the decision, and measures the resulting counterfactual shift, which yields a signed per-decision bias in score points, a flag for decisions the protected attributes changed, and a plain-language explanation naming the responsible factors. We evaluate on two real public datasets, Adu

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards?

arXiv:2608.21409v1 Announce Type: new Abstract: In medicine, claims remain valid when supported by empirical evidence grounded in stable biological reality. In law, by contrast, truth is contingent, defined by jurisdiction, temporal validity, and the hierarchy of authoritative sources. The recent success of large language models (LLMs) on medical licensing examinations has encouraged an expectation of comparable legal competence. This analogy, however, obscures a critical distinction between domains. Unlike in medicine, legal performance often depends less on inference than on determining when external authority is applicable, valid, and non-contradictory. We introduce a comparative diagnostic framework evaluating legal reasoning against medical baselines along four axes (knowledge recall, grounding, confidence, and robustness), uncovering a sharp domain asymmetry when applied to a new benchmark that encodes temporal validity and normative relationships. While medical LLMs reliably ben

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Generative Gap Filling

arXiv:2608.21401v1 Announce Type: new Abstract: Most contract litigation turns on contracts that imperfectly record parties' bargains. When the parties' dispute can't be solved by interpreting the text, courts fill the gap. Scholars have long assumed that the remaining text runs out quickly, and provides thin evidence of the actual deal on the disputed point. On that view, a judge who supplies the missing term must be drawing on something else, from commercial defaults to her own policy preferences. Despite generations of work, courts have no real alternative to such unruly methods. We tested that assumption. Taking real contracts, we masked a term the parties had negotiated and asked readers to predict what we removed. Lay respondents recovered the hidden term about half the time, twice what chance predicts. Law students and lawyers did marginally better. But large language models, given nothing but the rest of the contract, recovered it nearly nine times in ten. The deal, in short, t

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Survey Instrument to Assess Students' AI and Generative AI Knowledge

arXiv:2608.21391v1 Announce Type: new Abstract: In this research-to-practice paper we present a survey that can be used to assess students' AI knowledge. As the use of artificial intelligence (AI), including generative artificial intelligence (GenAI), has proliferated, so has the need to educate students about the topic. A range of AI literacy frameworks have been proposed, outlining the essential knowledge that students should have. Alongside, different ways of assessing AI knowledge have been developed. As yet, there is a lack of assessment instruments capable of evaluating multiple forms of student knowledge, including technical concepts, practical applications, and ethical concerns about AI use. In this article, we present a study implementing a comprehensive instrument to assess AI knowledge. The instrument combines measures from multiple scales to capture a range of literacy features and actual knowledge. We implemented the instrument in a higher education setting to assess its v

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Interrupting the Chain: Human Perception of AI-Generated Disinformation Through a Kill Chain Lens

arXiv:2608.21389v1 Announce Type: new Abstract: Generative AI enables customized misinformation at scale, yet defenses remain largely reactive. We present empirical findings from a human-subject study (n=504 participants, n=2,438 judgments) in which users classified news fragments by origin (human vs. machine) and veracity (real vs. fake). We organize results using an adapted cybersecurity kill chain as a taxonomy for intervention, mapping perception data onto stages of a cognitive attack lifecycle. Three key findings emerge: (1) a perception-accuracy gap where heightened suspicion does not improve detection; (2) modern LLMs frequently produce human-indistinguishable text; and (3) an asymmetric cognitive fatigue effect where fake-news detection degrades by 10.2 percentage points under sustained exposure while AI-origin detection remains stable. These findings identify candidate intervention points for proactive defense against AI-driven disinformation.

Source ↗
Showing 451–500 of 1593 signals
← Prev Page 10 of 32 Next →