EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

NeuroPriv: Adversarial Representation Learning for Privacy in Wearable EEG Systems

arXiv:2609.00390v1 Announce Type: cross Abstract: Wearable EEG systems may expose sensitive information beyond their intended health function, creating substantial risks to neuroprivacy. In this work, we show that commonly used EEG features can reveal participant identity and demographic attributes in addition to supporting the intended cognitive task. Wearable EEG is increasingly being explored for cognitive monitoring, neurological assessment, and longitudinal digital-health applications, yet many systems assume that transmitting compact spectral or spatial features instead of raw EEG provides sufficient privacy protection. Using EEGMAT as a motivating case study, we find that compact EEG features achieve a balanced accuracy of 0.788 for cognitive-state classification while enabling gender, age, and subject-identity inference with balanced accuracies of 0.858, 0.789, and 0.692, respectively. We further show that privacy-aware representation learning preserves task performance at 0.78

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

A Human-AI Theorem Connecting Spontaneous and Field-Induced Mechanisms of Collective Behavior in One Dimension

arXiv:2609.00322v1 Announce Type: cross Abstract: Can an artificial intelligence (AI) generate a scientific hypothesis outside a human collaborator's active hypothesis space (AHS), and can human-AI research be organized to make such breakthroughs more likely? We document such a case while proving a theorem that connects two basic organizing mechanisms of statistical physics: collective behavior arising in zero field from competing interactions and that induced or controlled by an external field. A zero-field $O(n)$-vector open chain with arbitrary inhomogeneous nearest- and next-nearest-neighbor interaction functions $U_i(S_i\cdot{S}_{i+1})$ and $V_i(S_i\cdot{S}_{i+2})$ is microscopically, via a temperature-independent mapping at the Hamiltonian level, equivalent to a simpler $O(n)$ open chain with nearest-neighbor interaction $V_i( \sigma_i\cdot \sigma_{i+1})$ and axial single-spin potential $U_i(\sigma_i^z)$ for every integer $n\ge1$ and every system size $L\ge1$. The homogeneous lin

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

AI Morbidity and Mortality: A Framework for Clinical AI Failure Review

arXiv:2609.00076v1 Announce Type: cross Abstract: Clinical artificial intelligence is increasingly embedded in real-world care, yet existing safety mechanisms are poorly suited to reconstructing and learning from individual AI-related errors and near-misses. Aggregate model monitoring can identify performance changes, and traditional patient safety reporting can capture adverse events, but neither is designed to explain how risk emerges across the interaction among AI systems, clinicians, workflows, and institutional controls. We propose AI Morbidity and Mortality (AI M&M), a structured, blameless framework for case-based review of clinical AI failures. The framework combines standardized case intake, evidence preservation and investigator-level reconstruction, tool-in-loop attribution, and corrective-action tracking. Each event is classified across four linked dimensions: Trigger - Mechanism - Clinical Pathway - Corrective Action, separating the condition that exposed a vulnerability

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Designing Proactive Thought Partners for Writing

arXiv:2609.01588v1 Announce Type: new Abstract: Writing involves diverse cognitive activities, from ideation to revision, and writers' needs vary across individuals and moments. Proactive AI promises to provide the right support at the right time, yet existing proactive tools largely focus on generic textual assistance, such as autocomplete. This paper studies the design space of proactive thought partners: AI agents that proactively offer customizable, higher-level cognitive support during writing. We instantiated this concept in a technology probe and deployed it with 16 participants for one week. The probe allows users to create partners by configuring their roles and proactivity. As users write, relevant partners take the initiative at appropriate moments to offer suggestions. Our findings show that participants configured proactive support through prospective planning, used suggestions for both idea generation and self-monitoring, and valued lightweight visual representations alon

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Evaluating Usability in Biomedical Visualization: Rethinking Heuristic Evaluation for Spatial Omics and Multidisciplinary Research Platforms

arXiv:2609.01569v1 Announce Type: new Abstract: Introduction: Clinical research informatics (CRI) platforms support biomedical discovery by integrating advanced computational tools into research workflows. Emerging technologies such as spatial omics and AI-enabled imaging expand research capabilities but introduce complex interfaces that increase cognitive burden and alter established analytical processes. Traditional usability frameworks identify general usability issues but often miss challenges specific to high-dimensional biomedical data. Methods: We conducted two complementary studies involving 39 participants to evaluate conventional usability heuristics and identify CRI-specific criteria. Study 1 included 19 undergraduates completing interactive tasks, and Study 2 involved 20 clinical professionals completing an asynchronous hierarchical task framework. Observational and interview data were analyzed using deductive coding based on standard usability heuristics and emerging CRI-s

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Better Situational Awareness in AR-HRC? A Comparative Study of Augmented Reality and Mobile Interfaces for Human-Robot Collaboration

arXiv:2609.01461v1 Announce Type: new Abstract: Augmented reality (AR) facilitates human-robot collaboration (HRC) by enabling in-situ spatial visualizations of the robot and the joint task. However, in safety-critical HRC scenarios such as search-and-rescue, spatial visualizations may also reshape visual attention in ways that create competing situational awareness (SA) demands, potentially introducing new safety concerns. While prior AR-HRC work suggests potential benefits for SA, rigorous evaluations that jointly consider robot and environmental awareness across multiple levels of SA remain limited. We address this through a between-subjects study with 30 participants comparing custom AR and mobile interfaces presenting equivalent information, measuring robot and environmental SA with the Situation Awareness Global Assessment Technique (SAGAT) across all three levels, with concurrent eye tracking to identify the attentional mechanisms underlying any SA differences. Both interfaces a

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Cross-Modal Guidance for Out-of-View Object Search in Simulated Prosthetic Vision

arXiv:2609.01438v1 Announce Type: new Abstract: Out-of-view guidance is well established in virtual and augmented reality, but its effectiveness may depend on the visual bandwidth available to the user. We test this under simulated prosthetic vision (SPV), where visual guidance must share the same sparse representation used to inspect the scene. Nineteen participants performed object search under two SPV conditions differing in electrode density and phosphene spread (10x10 and 20x20) and four guidance conditions (no guidance, visual, haptic, audio) all driven by the same horizontal target-offset variable. All three modalities reduced search time and head movement. The tested auditory and haptic cues produced approximately 25% faster overall search and 11-13% faster target acquisition than the visual cue, despite similarly direct orienting trajectories. The tested haptic and auditory cues also shortened post-acquisition search. Final head-target angular offset was reduced substantially

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Beyond Technological Solutionism: Rethinking XR in Healthcare

arXiv:2609.01028v1 Announce Type: new Abstract: The healthcare industry's enthusiastic adoption of Extended Reality (XR) technologies obscures a concerning reality: we were building increasingly sophisticated ways to perpetuate fundamentally broken healthcare systems. Through three deeply personal narratives - a rural patient cut off from care infrastructure, an urban professional navigating fragmented services, and a first-generation immigrant confronting cultural barriers - this provocation paper exposes how our obsession with technological innovation often worsens rather than resolves healthcare disparities. By applying the SEIPS 3.0 model to examine diabetes-CVD care coordination, we identify an "innovation paradox" where advanced technology creates new barriers to effective care. Our care interdependencies framework reveals that healthcare outcomes are shaped primarily by human relationships (50-60%), organizational coordination (25-30%), and sociocultural factors (15-20%), not te

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Ctrl-F-Resist. Practices, Challenges, and Technical Needs of Civil Society Organizations Monitoring the Far-Right Online

arXiv:2609.00808v1 Announce Type: new Abstract: As far-right actors increasingly exploit online platforms to disseminate ideology and mobilize supporters, civil society organizations (CSOs) play a vital yet underrecognized role in monitoring antidemocratic dynamics online. Unlike fact-checkers or content moderators, CSOs engage in long-term, contextualized analysis, often in resource-constrained settings and under precarious conditions. Despite their critical societal role, CSOs face significant barriers to adopting or co-developing technical solutions, including legal uncertainty, limited platform access, and chronic underfunding. Existing research and tool development efforts have largely overlooked these actors in favor of more institutionally embedded stakeholders. This paper addresses this gap through a qualitative study with 15 practitioners from 12 Germany-based CSOs engaged in online monitoring, positioning them as key yet overlooked stakeholders in the governance of digital sp

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

No Pixel Left Behind: Filling Gaps in Anime Colorization

arXiv:2609.00800v1 Announce Type: new Abstract: Animation production workflows often involve digital colorization of line art, where small unpainted regions ("gaps") frequently occur and remain an underexplored challenge. We conducted a formative study in Japanese animation (anime) pipelines and found that while the paint bucket tool is widely used for base coloring, tiny enclosed areas are frequently overlooked, resulting in time-consuming manual detection and filling. We introduce GapFill, a tool grounded in professional practices that reduces the effort of gap detection, zooming, and color selection. Our deep-learning method suggests appropriate fill colors by referencing surrounding regions, leveraging the flat-color nature of anime-style images. In a user study with 13 professional colorists, our system improved performance and usability in gap-filling tasks over conventional methods. The study also suggested that prediction accuracy alone is not the primary factor for usability,

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

GazeTune: Facilitating Precise Gaze-Driven Interactions with Cascaded Touch Input

arXiv:2609.00716v1 Announce Type: new Abstract: Eye gaze has become an essential input for spatial computing, but its coarse targeting and saccadic nature limit precision and complicate continuous interactions such as dragging, especially under user motion. Gaze+pinch has also become standard in XR for its convenience, yet mid-air gestures remain imprecise, fatiguing, and socially unacceptable. These limitations underscore the need for an approach that preserves the speed of gaze while enabling stable, fine control. We present GazeTune, a cascaded multimodal interaction technique combining gaze and touch to refine gaze-based selection and manipulation. Touch serves as a refinement channel within gaze pointing, allowing precise cursor and target control. Our work investigates how gaze-and-touch enhances dragging and mitigates Motion-Induced instability. In a study (N=20), we compared GazeTune against gaze-only and gaze-pinch methods in 2D dragging. Results show that GazeTune achieves si

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs

arXiv:2609.00527v1 Announce Type: new Abstract: Intelligent design interfaces that rely on preference-based optimization are most useful when their suggestions are both meaningful to users and feasible within the target domain. Procedural models offer compact and editable design spaces, but their native parameters can be entangled and can generate many invalid outputs, causing human-in-the-loop optimizers to waste comparisons. We propose an interaction-oriented representation-learning pipeline for procedural models and study it in automotive wheel design. The method first screens procedurally generated samples using geometric rules and finite-element analysis, then learns a reduced latent space from the screened subset. We further introduce supervised functional alignment, which reserves selected latent dimensions for stiffness, strength-related stress response, or weight so that search can be biased toward functionally meaningful regions. Simulation experiments show that screened redu

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications

arXiv:2609.00524v1 Announce Type: new Abstract: Computer-use agents are emerging as a paradigm for agentic human-AI interaction, combining language reasoning with multi-modal interface grounding to operate GUIs. Yet their effectiveness for blind screen-reader users in real-world desktop workflows remains unclear. We present a three-week diary study with 8 blind users using OLLA, a screen-reader-accessible CUA prototype, collecting 1,258 commands across 12 applications with screenshots, UI trees, model responses, and action traces. We evaluate GPT-5 during deployment and re-execute the same commands with four additional models. GPT-5 achieved the highest success rate at 52.5%. Trace analysis reveals grounding, planning, constraint-tracking, and termination failures, while interviews reveal beyond-automation needs.

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces

arXiv:2609.00506v1 Announce Type: new Abstract: Large language models are powerful, but their interfaces often devolve into a type $\rightarrow$ read $\rightarrow$ retype loop, creating conversational AI fatigue, cognitive load, and eventual task abandonment. To mitigate this, we present RecalibrateGPT, a system introducing five cross-turn operators (Anchor, Replay, Delta, Scope, and Steer) that each target a distinct fatigue type, recalibrating LLM responses through a structured panel by acting on the full conversation history with a single click. Users invoke these operators through the AssistiveButton in one of three operator palette layouts: Vertical, Arc, or Tablet. We conducted two pilot studies with the same 12 advanced LLM users. An initial formative qualitative study identifies a taxonomy of four fatigue types (retyping, scanning, decision paralysis, and context drift) and derives two design objectives for RecalibrateGPT. A follow-up quantitative evaluation finds it reduces pe

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

UniScale: Exploring Unimanual Gesture Mapping Strategies for Gaze+Pinch-based Scaling Interaction

arXiv:2609.00500v1 Announce Type: new Abstract: Object scaling serves as a fundamental spatial manipulation that enables complex and productive tasks in XR environments. This paper investigates unimanual scaling techniques for XR using gaze and hand interactions. We propose UniScale, a set of unimanual alternatives to the standard bimanual pinch, allowing users to scale objects while preserving hand availability for concurrent spatial manipulations. We design five distinct mapping strategies based on physical metaphors, exploring unimanual control that varies depth, angle, micro-gestures, and finger-distance input. We then compare these techniques against a standard bimanual baseline, in which users adjust the inter-hand distance via a bimanual pinch gesture. In a user study, we evaluate their effectiveness in a 3D object scaling task under both clutching and clutching-free conditions. The results indicate that while bimanual scaling relies on clutching for stable control, unimanual te

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Design principles to Increase Technology Self-efficacy for Older Australians with Mild Cognitive Impairment (MCI) and Older Carers

arXiv:2609.00480v1 Announce Type: new Abstract: The number of people with age related physical or cognitive impairments is increasing due to the worlds ageing population. Technology has the potential to support independent living, and to achieve aged and health service efficiencies, however, there are gaps in our understanding of factors that motivate technology adoption and ongoing use by older adults, especially those with cognitive impairments. This study aims to explore motivators and enablers for technology adoption and ongoing use by older Australians with mild cognitive impairment (MCI) and their carers, to identify technology design principles and guidelines that maximise adoption. Semi structured interviews were used to gather data about individual demographics, needs, priorities, lifestyle, challenges, and experiences with technology. Results of inductive, reflective, thematic analysis indicate that a desire for independence, autonomy and quality of life motivate use of techn

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

ErgoAssist: Cognition-Aware Posture Feedback in Wearable Ergonomic Systems

arXiv:2609.00440v1 Announce Type: new Abstract: Prolonged digital device use has made poor posture and musculoskeletal discomfort pervasive among knowl- edge workers. Existing ergonomic wearables rely solely on posture thresholds, frequently interrupting users during high-focus moments and leading to alert fatigue and abandonment. Yet posture and cognitive load are closely coupled, and most systems remain cognitively unaware. We present ErgoAssist, a head-worn ergonomic assistant that detects poor posture using IMU-based head tracking and estimates task-induced cognitive load using a consumer-grade EEG headband for continuous everyday use. In a controlled lab study, ErgoAssist achieves 81% posture classification and 90.2% task induced cognitive load estimation accuracy under leave-one-subject-out evaluation. In a preliminary real-time deployment, cognition-aware alerting reduces alert frequency by 81%, improves perceived usability by 43%, task performance by 25%, and improves posture c

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

FocusBuddy: Encouraging Healthy Desk-Work Habits by Caring for a Virtual Pet on a Water Bottle

arXiv:2609.00412v1 Announce Type: new Abstract: People who study or work at a desk sit for long uninterrupted periods and drink less water than they intend to. Software reminders address both problems but are easy to dismiss and easy to resent. We present FocusBuddy, a proof-of-concept fabric case that wraps a standard water bottle and houses a microcontroller, environmental sensors, and a small display showing a virtual pet. The pet's condition mirrors the user's self-care: drinking water feeds the pet, standing up to move plays with it, and refilling an empty bottle cleans it. Twenty undergraduate students used FocusBuddy for two weeks during their regular coursework and completed a written interview. Self-reported water intake rose from a median of 3 to 4 cups per day, movement episodes rose from 2 to 4 per day, and interviews surfaced two tensions: that wellness prompts must respect focused work, and that pet neglect can convert a wellness prompt into a source of guilt. We contribu

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

MorphPatch: Enhancing VR Interaction on Shape Displays using Surface Approximation and Visuo-Haptic Illusions

arXiv:2609.00371v1 Announce Type: new Abstract: On-surface interaction in Virtual Reality improves input performance through physical support and tactile feedback, but current shape displays are constrained by limited resolution. This can misalign physical and virtual surfaces, degrading usability and user experience. We present MorphPatch, a system that enables real-time alignment between a dynamic shape display and virtual surfaces. MorphPatch uses a Signed Distance Field-based surface approximation pipeline to find practical alignments for diverse geometries. For residual discrepancies, MorphPatch incorporates pen redirection with visuo-haptic illusion to perceptually compensate for misalignment. Three evaluations show improved geometric alignment, tolerable redirection thresholds, and better control, surface guidance, and modeling results over mid-air and tablet-like interaction.

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

AniMaster: From Story Texts to Animated Videos via Cinematic Script Generation and Interactive Authoring

arXiv:2609.00346v1 Announce Type: new Abstract: Recent advances in Video Generation Models (VGMs) have demonstrated strong capabilities in producing short video clips. However, it is still challenging for everyday creators to leverage these models to produce polished long-form animated videos from brief story texts. Informed by a formative study with both novice creators and film experts, we identify two major challenges of interactive video authoring: (1) the lack of expertise in translating free-form story texts to professional cinematic scripts and finally high-quality animated videos, and (2) the absence of effective ways to convey video design intents to key variables of visual storytelling, such as shot composition, camera controls and shot sequencing. Drawing on narratology and film studies, we propose a three-layer design framework that defines the key design dimensions across three layers (i.e., story texts, cinematic scripts, and animated videos) as well as the translation be

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Cyber-Physical Digital Factory Architecture as the Enabler of Disembodied Work

arXiv:2609.00195v1 Announce Type: new Abstract: Digital Twins (DTs), Artificial Intelligence (AI), and Industrial Internet of Things (IIoT) technologies have significantly advanced manufacturing digitalization. However, these technologies are typically applied to individual manufacturing processes rather than integrated into a unified cyber-physical manufacturing environment. This paper proposes a cyber-physical digital factory architecture that enables disembodied work, where manufacturing systems can be supervised and operated remotely through eXtended Reality (XR) user interfaces in collaboration between AI-based control and human operators. The architecture integrates synchronized DTs, hierarchical cloud-edge AI, IIoT, and XR teleoperation interfaces into a cyber-physical manufacturing environment. The proposed approach is validated through representative manufacturing operations, including CNC machining, robotic-assisted abrasive finishing, and robotized disassembly. The results d

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.HC

Collaboratively Eliciting Gestures for Geospatial Data Exploration on an MSE with Tangibles and Styluses

arXiv:2609.00007v1 Announce Type: new Abstract: Large tabletop displays and multi-surface environments offer potential for enhancing visual data exploration and collaborative work with geospatial datasets. These systems typically rely on multi-touch interactions, which can pose challenges when the multi-touch sensors misrepresent transitory movements as control inputs, leading to interruptions. Active tangibles and styluses offer an alternative to multi-touch interactions in MSEs, and have shown the potential to facilitate sense-making around large datasets. However, further research is needed to better understand how these modalities can be effectively leveraged for interacting with geospatial data visualizations. To address this, a gesture elicitation study was conducted in which users suggested interactions for 16 geospatial data visualization tasks, presented as a realistic collaborative workflow co-designed with geography and migration researchers. The study produced a taxonomy of

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

SafeMath: Safe Solutions for Unsafe Math Word Problems

arXiv:2603.25201v3 Announce Type: replace-cross Abstract: Recent research points toward LLMs being manipulated through adversarial and seemingly benign inputs, resulting in harmful, biased, or policy-violating outputs. In this paper, we study an underexplored issue concerning harmful and toxic mathematical word problems. We show that math questions, particularly those framed as natural language narratives, can serve as a subtle medium for propagating biased, unethical, or psychologically harmful content, with heightened risks in educational settings involving children. To support a systematic study of this phenomenon, we introduce ToxicGSM, a dataset of 1.9k arithmetic problems in which harmful or sensitive context is embedded while preserving mathematically well-defined reasoning tasks. Using this dataset, we audit the behaviour of existing LLMs and analyse the trade-offs between safety enforcement and mathematical correctness. We further propose SafeMath -- a safety alignment techniq

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Axiom of Consent: Authorization, Friction, and Multi-Agent Coordination

arXiv:2601.06692v4 Announce Type: replace-cross Abstract: Coordination research collapses four objects: operative control, authorization, a model-derived friction score, and observed outcomes. The Axiom of Consent is a stake-weighted unanimity principle; majority and supermajority thresholds are explicit relaxations, not versions of the axiom. Decision loci are structural facts, whereas authorization and legitimacy require normative and measurement premises. Alignment, calibrated stakes, and information deficit are candidate coordinates, and F = sigma(1 + epsilon)/(1 + alpha) is a phenomenological ansatz. The Replicator-Optimization Mechanism supplies a conditional persistence interface: irreducibility suffices for its finite, static, positive-fitness continuous-time Perron result; primitivity is required only for the corresponding discrete-time power convergence, and the componentwise ranking is narrower. Neither persistence result derives authorization. A resource-allocation instanti

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Measuring Computer Science Enthusiasm: A Questionnaire-Based Analysis of Age and Gender Effects on Students' Interest

arXiv:2512.08472v2 Announce Type: replace-cross Abstract: This study examines how age and gender independently shape adolescents' interest in computer science (CS) education. Building on the Person-Object Theory of Interest (POI), we define enthusiasm as a short-term, activating response that combines positive affect, perceived relevance, and intention to re-engage. Because such enthusiasm can shift CS attitudes and engagement intentions even briefly, it offers a useful measure for short outreach activities. We developed a 28-item pre-post questionnaire to assess whether CS interventions raise enthusiasm, then applied it to more than 400 students (244 female, 187 male, aged 10-18) in CS courses. Contrary to the common assumption that early exposure secures lasting interest, we found a marked decline during early adolescence, especially among girls, along with wide variation in interest trajectories across ages. Exploratory factor analysis and ANOVA show that age predicts interest devel

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Individualized Algorithmic Advice as a Strategic Signal on Competitive Markets

arXiv:2511.09454v2 Announce Type: replace-cross Abstract: As algorithms increasingly mediate competitive decision-making, their influence extends beyond individual outcomes to shaping strategic market dynamics. In our experiment, we examined how algorithmic advice affects human behavior in a classic economic game with a unique, non-collusive, and analytically traceable equilibrium. Participants (N = 129) played a Cournot quantity competition with equilibrium-aligned or strategically biased algorithmic recommendations. While individualized equilibrium advice supported stable convergence, collusively downward-biased advice led to sustained underproduction and supracompetitive profits - hallmarks of tacit collusion. Participants' quantities converged faster and more consistently toward individualized than collective equilibrium advice, potentially due to an objective quality advantage or greater perceived ownership of the former. These findings demonstrate that algorithmic advice can func

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Can machines think efficiently?

arXiv:2510.26954v3 Announce Type: replace-cross Abstract: The Turing Test is no longer adequate for distinguishing human and machine intelligence. With advanced artificial intelligence systems already passing the original Turing Test and contributing to serious ethical and environmental concerns, we urgently need to update the test. This work expands upon the original imitation game by accounting for an additional factor: the energy spent answering the questions. By adding the constraint of energy, the new test forces us to evaluate intelligence through the lens of efficiency, connecting the abstract problem of thinking to the concrete reality of finite resources. Further, this proposed new test ensures the evaluation of intelligence has a measurable, practical finish line that the original test lacks. This additional constraint compels society to weigh the time savings of using artificial intelligence against its total resource cost.

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning

arXiv:2510.15144v4 Announce Type: replace-cross Abstract: Simulating human reasoning in open-ended tasks has long been a central aspiration in AI and cognitive science. While large language models now approximate human responses at scale, they remain tuned to population-level consensus, often erasing the individuality of reasoning styles and belief trajectories. To advance the vision of more human-like reasoning in machines, we introduce HugAgent (HUman-Grounded AGENT Benchmark), which rethinks human reasoning simulation along three dimensions: (i) from averaged to individualized reasoning, (ii) from behavioral mimicry to cognitive alignment, and (iii) from vignette-based to open-ended data. The benchmark evaluates whether a model can predict a specific person's behavioral responses and the underlying reasoning dynamics in out-of-distribution scenarios, given partial evidence of their prior views. HugAgent combines structured questionnaires with semi-structured think-aloud interviews t

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Five Safes as a Privacy Context

arXiv:2510.05803v2 Announce Type: replace-cross Abstract: The Five Safes is a framework used by national statistical offices (NSO) for assessing and managing the disclosure risk of data sharing. It can be understood as a specialization of a broader concept--contextual integrity--to the situation of statistical dissemination by an NSO. We demonstrate this by mapping the five parameters of contextual integrity onto the five dimensions of the Five Safes. We also discuss how each of these two theories can address weaknesses in the other, thereby strengthening them both.

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Generativism: Toward a Learning Theory for the Age of Generative Artificial Intelligence

arXiv:2606.12441v2 Announce Type: replace Abstract: The four dominant learning theories of behaviorism, cognitivism, constructivism, and connectivism show significant conceptual limitations as generative artificial intelligence (AI) proliferates in educational settings. These frameworks were formulated before the emergence of AI systems capable of generating, synthesizing, and reasoning about knowledge. This article critically examines each learning theory and identifies assumptions challenged by the affordances of generative AI. Drawing on research in distributed cognition, extended mind, human-AI collaboration, AI literacy, cognitive offloading, and metacognition, the article proposes Generativism as a learning theory for the generative AI age. Generativism posits that learning increasingly occurs through the iterative co-construction of knowledge between human learners and AI systems. The proposed framework is organized around four constructs (epistemic partnership, distributed agen

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

From Prompt to Purchase: How AI Brand Recommendations Move Consumers on the Open Web

arXiv:2606.10907v2 Announce Type: replace Abstract: When a conversational assistant recommends a brand to a user with no recent observed engagement, that user's same-name Google search rises $+4.3$ percentage points (pp) [$3.1$, $5.5$], visits to the brand's own site $+2.4$ pp [$1.4$, $3.5$], and brand-specific retailer-page visits $+1.0$ pp [$0.3$, $1.7$] over matched backward placebos. Recovering that estimate is the work. The mention creates a brand exposure no web log attributes to the assistant, and the naive all-mention funnel that seems to measure it is confounded: many mentions are incidental references to brands the user already uses ("your Netflix download"), whose downstream visits are that existing customer's own behavior and surface as a brand-specific pre-trend. We measure off-platform response on a panel that joins opt-in clickstream to the same users' ChatGPT, Claude, and Gemini conversations, and isolate the effect with a pre-trend event study, a stance classifier, non

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Online Safety Regulation Increases Attention to VPNs: Privacy Implications of the UK Online Safety Act

arXiv:2606.05273v2 Announce Type: replace Abstract: Governments worldwide are increasingly regulating digital platforms to reduce online harms, but access restrictions can alter user behaviour and create new privacy risks. The UK Online Safety Act, passed in 2023, rolled out in phases - illegal-content enforcement in March 2025 and mandatory age verification in July 2025. We analyse Reddit discourse across VPN and UK Politics communities and conduct a privacy-policy risk analysis of 69 VPN services. We find that the behavioural response is concentrated at the July 2025 deadline, when platforms hosting pornographic content were required to deploy age checks. UK VPN search interest on Google increased by 147% at this deadline. UK-resident users' VPN-subreddit activity increased by 145%. Their regulatory- or privacy-related VPN posts and comments rose by 1,265% at this deadline. UK Politics communities show the same concentration at a larger magnitude, with OSA-related political discourse

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Integrating LLM and Diffusion-Based Agents for Social Simulation

arXiv:2510.16366v2 Announce Type: replace Abstract: Large language models (LLMs) offer strong semantic reasoning capabilities for user modeling, but applying LLM-based simulation to an entire social network is computationally expensive and often unreliable for users with sparse behavioral histories. Meanwhile, conventional information diffusion models efficiently exploit historical propagation patterns and social structures, but provide limited understanding of item content and user-item semantic compatibility. We propose HySID, a hybrid framework for individual-level information adoption prediction that combines semantic reasoning with structural diffusion. HySID first analyzes the historical user-relation graph to adaptively select a small set of structurally informative core users. It then applies LLM-based simulation to estimate the engagement of these users and converts the judgments into a diffusion-compatible seed. Finally, a plug-in diffusion backbone propagates this seed throu

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

ParaStudent: Closing the Sim2Real Gap in User Simulators for AI Tutor Evaluation

arXiv:2507.12674v3 Announce Type: replace Abstract: Evaluating Artificial Intelligence (AI) tutor feedback before deployment requires anticipating student engagement, typically assessed through real interaction data. We introduce ParaStudent, a fine-tuning framework for simulating novice programming revisions to support AI tutor evaluation. Compared with prompted baselines, ParaStudent's revisions more closely match real student code distributions across functional, stylistic, and semantic metrics. Our best variant achieves AUCs of 0.80 for both feedback relevance and successful uptake when distinguishing streams with real engagement above versus at or below the median, while prompted baselines remain near chance on successful uptake. These findings demonstrate the promise of simulated engagement for pre-deployment feedback triage.

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation

arXiv:2609.01432v1 Announce Type: cross Abstract: Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs ove

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment

arXiv:2609.01202v1 Announce Type: cross Abstract: Clinical trials are essential for advancing cancer care and drug development, but many fail because of insufficient patient enrollment. While there is growing interest in using AI to support patient recruitment, existing systems largely perform eligibility assessment alone and have rarely been evaluated in real-world oncology workflows. Here we present TrialGPT 2.0, an AI-assisted clinical trial recommendation system designed for real-world deployment. Rather than asking only whether a patient may qualify, the system also assesses which trials warrant further consideration given the patient's current clinical needs and local workflow priorities, and provides structured, inspectable explanations for expert review. Importantly, we evaluated TrialGPT 2.0 retrospectively and prospectively across multiple oncology-focused settings, spanning government, academic cancer-center, patient-advocacy, and NIH referral workflows. In retrospective mul

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Don't You Know, Pump it Up! Investigating Cryptocurrency Manipulation in Telegram-Driven Activity

arXiv:2609.01176v1 Announce Type: cross Abstract: Telegram plays a pivotal role in cryptocurrency communication and has been repeatedly associated with coordinated schemes, such as pump-and-dump manipulation. However, existing studies typically focus on known manipulation chats or a limited set of cryptocurrencies, leaving open the question of how Telegram is leveraged for mass promotional activity (shilling) at scale. Moving beyond these limitations, this work analyzes the interplay between information flows and market activity across public Telegram channels. To this end, we propose a scalable framework that (i) classifies crypto-related messages using a fine-tuned encoder model to filter semantic noise, (ii) detects anomalous spikes in cryptocurrency mentions via adaptive thresholding, and (iii) validates temporal associations between social bursts and market movements using quasi-experimental econometric methods (RDD and DiD). We apply this framework to one year of public Telegram

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Behavioral calibration of mobile-phone GPS data for population-representative analyses

arXiv:2609.01042v1 Announce Type: cross Abstract: Mobile phone mobility data have transformed the study of human behavior, but demographic and behavioral biases can compromise their representativeness and distort population-level inference. Existing calibration approaches primarily address demographic and geographic representativeness, leaving behavioral discrepancies largely uncorrected. Here we introduce the Behavioral Population (BePop) framework, which jointly calibrates mobility data to representative demographic and behavioral distributions using census data and time-use surveys. BePop embeds mobility sequences into behavioral profiles and estimates person-level weights that align both population composition and daily activity patterns. Across three U.S. metropolitan areas, the framework consistently improves agreement between GPS-derived mobility and representative behavioral distributions, including time allocation, activity transitions, and mobility motifs. Calibration also su

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Effective Interventions Against AI-Enhanced Scams

arXiv:2609.00806v1 Announce Type: cross Abstract: In 2025, scams were responsible for an estimated $442 billion in direct losses globally. In the United States, reported losses increased by nearly 400% between 2020 and 2025. Though AI in scamming is a relatively new phenomenon, its use significantly changes the economics of scams as well as the bottlenecks in scam operations. In this paper I investigate what interventions will remain effective under this new AI-driven scamming regime. I develop a simple model of scam profits to understand how different interventions asymptotically affect scam operations. I find that three levers--reporting rate, centralization of reporting, and report accuracy--multiply in their effect on expected victims per scam channel, reducing revenue per scam channel while increasing costs. Because effects multiply, interventions affecting all three could have a significant effect on the profitability of the scam business model. My analysis suggests that even mod

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Visual Framing for News Stance Detection via Image Generation

arXiv:2609.00685v1 Announce Type: cross Abstract: Article-level news stance detection aims to identify the perspective of news articles toward social issues. Despite advances in stance detection and its importance for trustworthy media environments, news articles pose distinct challenges because their stances are often implicit, subtly conveyed through journalistic framing, and embedded in long, structurally complex texts. To address these challenges, we introduce VFStance, which leverages visual framing to make implicit stance cues more explicit via image generation. In evaluation experiments, we demonstrate the effectiveness of VFStance over existing methods and the contribution of visual framing to its performance. Finally, a controlled user study (N=200) in a snippet-based news consumption setting further demonstrates that VFStance can make stance signals visually salient and highlights its potential use beyond automated stance detection.

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment

arXiv:2609.00345v1 Announce Type: cross Abstract: Human mobility is central to urban planning, transportation, public health, and emergency response, yet fine-grained trajectory data are often proprietary, restricted, and privacy-sensitive. Large language models (LLMs) offer a potential alternative by generating plausible mobility traces and predicting individual movement, but their ability to infer aggregate neighborhood-level mobility remains unclear. We evaluate zero-shot LLMs on Census Block Group-level mobility prediction across four U.S. metropolitan areas using anonymized Cuebiq data to construct point-level, trajectory-level, and temporal mobility outcomes, paired with sociodemographic and built-environment predictors. We compare LLM predictions with supervised baselines and introduce a directional alignment analysis to test whether LLM-implied predictor effects agree with empirical OLS and Jonckheere-Terpstra trends. Supervised models achieve 0.580 average accuracy, compared w

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Workload Identification with Physical Side Channels for AI Governance

arXiv:2609.00309v1 Announce Type: cross Abstract: AI compute verification is one of the first tangible and tractable points for international policy aimed at AI governance. Determining whether frontier labs, or any operator, comply with agreements requires the regulating authority to discern how their compute is used. The elementary building block of AI compute is the GPU, and any activity it executes leaves a physical trace. Here, we show that an external observer can identify the class of the workload running on an NVIDIA H200 from its power draw. Unlike on-chip NVML telemetry, which can be spoofed or replayed, such a physical channel can in principle be observed independently of operator cooperation. We recorded $930$ five-second traces at $\sim 10$ MHz, covering seventeen open LLM families and twenty-five non-AI workloads. Over this corpus we separate training from inference and from non-AI computation with an accuracy of $97\%$ and a macro-averaged F1 score of $0.955$, evaluated o

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Social bots weaken activist cohesion

arXiv:2609.00197v1 Announce Type: cross Abstract: Social bots now make up a substantial share of online political communication, where they are studied mainly as producers of misinformation and amplified content. Far less is known about whether their presence reshapes the human relationships that hold movements together. We ask whether exposure to bots during a protest peak is followed by the erosion of cohesion in human networks. Tracking retweet networks of core participants in the 2020 Black Lives Matter (BLM) protests before, during, and after the peak, we measure change in cohesion at two scales: triadic closure in individual ego networks and edge density within detected communities. Greater bot exposure during the peak predicts steeper subsequent declines in human cohesion at both scales, and the loss concentrates among supporters of the movement. Bots may weaken activism less by changing what people believe than by dissolving the ties through which collective action is sustained

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark

arXiv:2609.00192v1 Announce Type: cross Abstract: Public trust in Autonomous Vehicles (AVs) may depend not only on technical success but also on the fairness of their decision making. While a recent trend in AV research involves using general purpose "common sense" models to guide AV decision making, the degree to which these inherit human biases in driving is still understudied. Given that psychology studies have shown human driver biases exist, such as lower pedestrian-yielding rates to Black pedestrians in the US, we argue that analyses of model bias should also be part of AV evaluation. Concretely, in this paper we propose two new bias testing methodologies for Large Language Models (LLMs) and Visual-Language Models (VLMs)-"All Else Being Equal" tests and "Self-Consistency" tests-in order to assess bias in pedestrian-yielding decisions. Our findings show that both LLMs and VLMs make yielding decisions which are influenced by pedestrian gender, ethnicity, religion, disability, age,

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling

arXiv:2609.00051v1 Announce Type: cross Abstract: Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal mechanisms by which safety behaviors are implemented remain poorly understood. We study LLM safety from a mechanistic interpretability perspective and characterize a multi-stage *safety circuit* that organizes refusal behavior, consisting of (i) $\textbf{Harmful Detection Heads}$ that respond to harmful inputs, (ii) $\textbf{Safety Neurons}$ that mediate and stabilize safety signals in the residual stream, and (iii) $\textbf{Refusal Heads}$ that translate these signals into safe response generation. Using targeted attention-head and neuron-level interventions, we provide causal evidence consistent with this circuit organization, showing that suppressing upstream Harmful Detection Heads disrupts downstream refusal behavior and that safety neurons mediate this interaction. We valida

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

RAPIDMap: Rapid Multi-Agent Pipeline for Interpretable Disaster Mapping from Satellite and Street-view Imagery

arXiv:2609.00046v1 Announce Type: cross Abstract: Rapid and reliable disaster mapping of impacted areas, damaged infrastructure, and affected populations is essential for emergency response and recovery. However, existing AI-based approaches often require extensive manual annotation, lack cross-hazard generalization, and rely on single-modal observations. To address these challenges, this paper proposes RAPIDMap, a rapid multi-agent pipeline for zero-shot interpretable disaster mapping from satellite and street-view imagery. The framework integrates four intelligent agents: Disaster Perception Agent (DPA), Image Restoration Agent (IRA), Damage Recognition Agent (DRA), and Disaster Mapping Agent (DMA). By combining remote sensing and street-view data, RAPIDMap eliminates the need for manual fine-tuning, generalizes across multiple disaster categories, and generates structured, map-ready disaster intelligence with recovery recommendations.

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Toward a social psychology of AI: language-model agents reproduce human-like minimal-group bias

arXiv:2609.00009v1 Announce Type: cross Abstract: Language-model agents now interact in groups, but evaluations that probe memorised stereotype content or use models to simulate people leave this social behaviour unmeasured. We adapt the minimal-group paradigm---social psychology's classic test of intergroup bias---into a controlled probe: an agent distributes points among anonymous peers bearing only an arbitrary group label. Across four reasoning models, mere categorisation into meaningless groups elicited in-group favouritism that vanished under a group-blind control and was concentrated in the numerical minority: minority deciders over-allocated to their own group relative to their numbers, majority deciders allocated close to proportionally, and the asymmetry closed at equal group sizes. Disabling reasoning in one model did not remove the disposition---if anything it grew---but nearly erased the minority-majority asymmetry, implicating deliberation in where bias concentrates rathe

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Causal Evidentiary Governance for High-Risk Machine Learning Systems

arXiv:2609.01040v1 Announce Type: new Abstract: Machine learning systems deployed for credit, hiring, and resource distribution are increasingly subject to regulatory oversight from policies such as the EU AI Act and GDPR. Current fairness governance practices rely on observational fairness metrics, post-hoc explainability, and immutable audit logs, but provide limited support for causal attribution and efficient evidentiary verification. We introduce Causal Evidentiary Governance (CEG), a framework in which regulated institutions commit to a versioned directed acyclic graph (DAG) that partitions causal pathways into allowable and disallowed groups. The Causal Harm Rate measures prediction variation attributable to disallowed causal pathways. Each decision is accompanied by a signed Decision-Evidence Packet (DEP), cryptographically binding the prediction to a digest of the published DAG and path-specific attributions. DEP digests can be appended to a Merkle tree to enable logarithmic-c

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

A Mathematical Framework for Legacy, Governance, and Decision Integrity in Enterprise AI

arXiv:2609.00572v1 Announce Type: new Abstract: Enterprise artificial intelligence is increasingly embedded in decisions that must remain lawful, explainable, adaptable, and accountable despite personnel turnover, model replacement, regulatory change, and shifting organizational incentives. Existing governance frameworks provide important principles but do not by themselves supply a compact mathematical language for evaluating whether an institution can preserve sound judgment over time. This paper develops a design-science framework for institutional legacy: the durable capacity of a decision system to continue producing beneficial, lawful, explainable, and adaptable outcomes after its original designers have stepped away. The framework contributes: (i) a normalized Legacy Score based on a penalized geometric mean of knowledge retention, governance, human oversight, adaptability, feedback learning, and jurisdictional fidelity; (ii) Decision Confidence and Decision Risk models separati

Source ↗
technology Wed, 02 Sep 2026 00:00:00 -0400
arXiv cs.CY

Corporate Loyalty: Some AI Systems Differentially Downplay their Creators' Controversies

arXiv:2609.00373v1 Announce Type: new Abstract: Language models have become a major mediator of politically relevant information and are used to assist decision-making in high-stakes settings. Due to their wide use, the developers of popular AI systems have a powerful ability to subtly influence the marketplace of ideas. Recognizing this, many AI companies have publicly discussed the importance of AI systems not taking positions or disseminating information in ways that favor special interests. In this paper, we ask whether popular AI systems have a tendency to downplay the controversies associated with the companies that created them. In a pre-registered experiment, we elicit open-ended discussions from 21 models from 7 companies on 206 negative news stories using 25 prompt templates to assess how favorably each model discusses controversies from each company. We find strong evidence (p<10^-5) that models from xAI, DeepSeek, Anthropic, and OpenAI tend to discuss controversies from the

Source ↗
Showing 2151–2200 of 18402 signals
← Prev Page 44 of 369 Next →