EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Looking Around by Looking Around: Omnidirectional Gaze-based VR Viewport Control

arXiv:2608.30014v1 Announce Type: new Abstract: Traditional VR viewport control primarily relies on head and torso movement, which can be effortful and limiting in both constrained and extended-use settings. We introduce Looking Around by Looking Around (LALA), a gaze-based VR pitch-and-yaw viewport control technique designed for natural and effortless omnidirectional exploration via eye movements, without requiring or obstructing movement of the head, hand, or body, offering a low-effort and highly accessible interaction method. Because gaze is primarily used for perception and exhibits oculomotor and perceptual asymmetries, using it directly for control is difficult. To address this, we designed an asymmetric omnidirectional control profile for the eye, then built on it to exploit tendencies for eyes to stay within comfortable regions for viewport control. We evaluated LALA in a user study (N=18) featuring two contrasting tasks: alignment towards known directions and open-ended visua

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Multimodal Takeover Requests for Drivers with Hearing Loss: Implications for AI-Enabled Communication in Automated Vehicles

arXiv:2608.30013v1 Announce Type: new Abstract: More than 430 million people worldwide live with disabling hearing loss. Although people with hearing loss are legally permitted to drive and may benefit from conditionally automated vehicles, SAE Level 3 systems still require drivers to respond to takeover requests when automation reaches its limits. Existing takeover requests often rely on auditory information, yet little evidence addresses visual and tactile designs for drivers who cannot rely on sound. This driving-simulator study with 40 participants examined the effects of information type (instructional, informative, and baseline), signal type (visual, tactile, and visual-tactile), and hearing condition (normal hearing and simulated hearing impairment) on takeover performance. Information type significantly affected reaction time, with baseline displays producing the shortest times. Signal type significantly affected reaction and takeover time, with visual-tactile displays producin

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

A Cyber-Physical Machine Tool Framework with a Real-Time Machining Process Digital Twin

arXiv:2608.29955v1 Announce Type: new Abstract: Digital Twins (DTs) have emerged as a key technology for improving the monitoring, optimization, and automation of manufacturing systems. However, existing Cyber-Physical Machine Tool (CPMT) implementations primarily represent the machine tool, while the machining process remains only partially synchronized with its physical counterpart. This paper extends a previously presented CPMT framework by introducing a hierarchical DT framework that simultaneously maintains DTs of both the machine tool and the machining process. The proposed framework integrates real-time CNC operational data, a voxel-based workpiece representation, synchronized process vibration measurements, and a persistent part DT repository for process replay, traceability, and future synthetic data generation. Experimental evaluation demonstrated real-time operation at a 20 Hz machining-state update rate, interactive visualization exceeding 100 frames per second, and a mean

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

The Policy Deficit in AI x Social-Emotional Learning Research

arXiv:2608.29950v1 Announce Type: new Abstract: As artificial intelligence (AI) is increasingly integrated into social-emotional learning (SEL) initiatives, the need for evidence-based policy has become paramount. We systematically reviewed 65 peer-reviewed papers that examine the intersection of AI and SEL to investigate how these studies articulate policy implications. Our analysis revealed a substantial "policy deficit" in the current AI x SEL literature: nearly three-quarters of the studies did not mention policy implications at all. Using the "WH-question" framework (Who, What, Why, When/Where, and How), we map the policy implications narratives present in the literature and show that they often lack the specificity and actor-oriented guidance required for effective evidence-informed policymaking. We find a significant association between publication venue and policy engagement, suggesting that current academic incentive structures may prioritize technical innovation and pedagogic

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Towards Effective Generation of Interactive Visualizations with Vibe Coding: An Empirical Study

arXiv:2608.29550v1 Announce Type: new Abstract: Constructing interactive visualizations has traditionally required substantial human effort, involving both technical implementation and design decision-making. Recently, vibe coding, a programming paradigm leveraging Large Language Models to generate, interpret, and refactor code from natural language specifications, has emerged as a promising approach to reduce the burden. However, the capabilities and limitations of vibe coding in building interactive visualizations remain unexplored. To address this gap, we conducted a user study with 78 participants that were tasked with constructing interactive visualizations using vibe coding. We further collected users feedback through questionnaires, interviews, and case analyses. Based on this study, we examine (1) the capabilities and (2) user experience of vibe coding in generating interactive visualizations, and (3) the practical human-agent collaboration strategies adopted. Our findings prov

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

EITWatch: Smartwatch-Integrated Planar Electrical Impedance Tomography for Hand Gesture Recognition

arXiv:2608.29415v1 Announce Type: new Abstract: Wrist Electrical Impedance Tomography (EIT) senses hand gestures from muscle- and tendon-driven impedance changes, but prior wrist-EIT systems require electrode coverage beyond the watch-back contact patch and separate analog front ends. We present EITWatch, the first wrist-EIT system built around smartwatch case-back geometry, asking whether this contact patch alone can support gesture recognition: eight planar electrodes in a 31 mm ring acquire 35 impedance measurements at 48 Hz. Because a planar array cannot encircle the wrist, EITWatch uses multi-depth scanning to sample multiple source-sink distances and current paths; it beat matched adjacent injection by 15.1/10.4 percentage points (macro/micro) across all 12 participants. In a prompted study, within-session leave-one-round-out accuracy reached 91.4%/92.5% (window/trial) for six macro-gestures, and 90.1%/91.5% (window/segment) for five micro-gestures plus relax; window-level cross-

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Feelium: A Touchable Blimp Body for Aerial Telepresence

arXiv:2608.29391v1 Announce Type: new Abstract: Floating things invite touch. We present Feelium, a blimp-based telepresence platform that enables visual embodiment and touch interaction through its inflatable skin. Through a VR headset, a remote person inhabits the blimp, looking out of it first-person, appearing on its skin as a face or avatar, and steering it through the room. Partners in the room pat it, press a palm against it, draw on it, or lean into it; the skin senses each contact, renders it into the wearer's view in VR spaces. Touch thus provides a physical interaction channel for remote presence, turning the skin into a shared surface between remote and co-located partners.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Understanding Behavioral Dark Patterns of High BMI Individuals

arXiv:2608.29328v1 Announce Type: new Abstract: Understanding how everyday behaviors influence body weight is essential for designing effective and personalized health interventions. Existing studies largely rely on self-reported questionnaires or limited sensing modalities, making it difficult to capture the temporal dynamics of daily behavior. In this work, we analyze the DiversityOne dataset, comprising four weeks of passive smartphone sensing and ecological momentary assessments collected from 453 university students across eight countries. We extract behavioral features spanning dietary habits, physical activity, screen time, and smartphone usage, and investigate their associations with self-reported Body Mass Index (BMI). Beyond feature-level analysis, we employ Hidden Markov Models (HMMs) to uncover latent behavioral patterns. Our analysis reveals that higher BMI is associated with more frequent consumption of soda, alcohol, and processed meat. We further reveal that overweight

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Measuring the "Interaction Gap" in Drama Therapy with AI

arXiv:2608.29292v1 Announce Type: new Abstract: Generative AI is increasingly being introduced into expressive arts therapy, where it is often credited with offering a non-judgmental environment that supports psychological safety. Existing HCI work has largely positioned AI as a co-creative material or as a bridge/mediator into human-led care. This paper explores a different position. When a patient performs the same drama therapy task with an AI partner and with a human partner, the resulting self-presentations tend to differ in patterned ways. We propose treating this difference, the Interaction Gap, as a diagnostic lens within drama therapy. Rather than asking which context elicits a truer self, the lens reads the difference between the two performances as information about the social pressures shaping self-expression in each context. We sketch a starting point for task design and measurement signals grounded in drama therapy's existing use of role and aesthetic distance, and raise

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

An Eye-Tracking Dataset for Viewing Distance Categories in Real-World Scenarios

arXiv:2608.29192v1 Announce Type: new Abstract: Estimating viewing distance from gaze behavior is essential for understanding user intent and enabling distance-aware interactive systems. However, most existing eye-tracking datasets have been collected in constrained settings, such as laboratory environments or static tasks. Consequently, they only partially capture viewing behaviors in real-world situations where viewing distance changes with natural head and body movements. We introduce GazeDepth, an eye-tracking dataset collected from 19 participants using a wearable tracker during tasks reflecting real-world scenarios. GazeDepth includes fixed-distance viewing scenarios with constant observer-target distances at near (33 cm), middle (50 cm), and far (300 cm), as well as variable-distance viewing scenarios in which participants shift gaze among targets at different depths in indoor and outdoor environments. The dataset provides synchronized gaze data, pupil size, 3D eye-vectors, and

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Sharing Roughness with Hand-Outline Visualization to Reduce Sensory Asymmetry in VR Collaboration

arXiv:2608.29040v1 Announce Type: new Abstract: In collaborative VR, asymmetric access to haptic hardware creates a critical information gap: tactile evidence remains private to the haptic user, hindering the shared understanding needed for joint decision-making. While prior work has explored crossmodal sensory cues in virtual environments, it remains unclear how such cues should be designed for asymmetric collaboration, where collaborators receive information through different modalities. In our setting, the haptic user feels roughness through fingertip vibration, whereas the non-haptic user relies on vision alone. To reduce this asymmetry, we propose externalizing an object's tactile state through a glanceable hand-outline visual proxy. Specifically, we examine whether abstract visual roughness cues based on line shape and motion can encode three discrete roughness levels for both haptic and non-haptic users. Two preliminary studies establish a shared visual semantics by identifying

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Using LLMs to Mimic the Conversational Dynamics of Reddit Communities

arXiv:2608.28989v1 Announce Type: new Abstract: Online communities face a constant battle against toxicity and misinformation. While human moderators struggle to keep pace with the volume of content, LLMs offer a promising solution for automatically generating constructive responses and shaping online interactions. This paper preliminarily investigates if LLMs can mimic the communication styles of Reddit users using their comment history as context. We evaluate two prompting approaches: predicting a target comment and filling in masked comments. We find that LLMs outperform expectations at replicating comment structure and formality, but struggle to accurately capture nuanced emotions, e.g. understating joy and overstating anger. These findings highlight a promising direction for LLMs in guiding online conversations towards prosociality influencing emergent communication patterns and norms within the community. The results of our study inspire future work with more rigorous methods of

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

AREAs-Lab: An Interactive Environment for AI-driven Requirement Elicitation for AI Systems

arXiv:2608.28979v1 Announce Type: new Abstract: Building effective AI systems increasingly depends on writing high-quality task requirements, yet users often struggle to articulate the constraints, preferences, and edge cases that determine success. This problem is especially acute in AI development, where behavior is shaped not only by human expectations but also by data characteristics. We present AREAs-Lab, an interactive environment for AI-driven Requirement Elicitation for AI systems. In AREAs-Lab, an assistant iteratively refines an initially incomplete requirement by analyzing the underlying dataset and asking targeted clarification questions to uncover the user's latent intent. To study this setting systematically, we construct a synthetic benchmark grounded in 16 public datasets spanning diverse domains and task types. Each benchmark instance includes a user profile, a complete reference requirement, and an intentionally underspecified version that serves as the assistant's st

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

The Web-CLI: Verifiable Privacy for Tools, Models, and Inference Engines in the Browser

arXiv:2608.28950v1 Announce Type: new Abstract: We introduce the Web-CLI, a novel application architecture deploying powerful computational capabilities (command-line tools compiled to WebAssembly, models run through client-side inference runtimes, and GPU-accelerated engines) as zero-install, offline-capable browser applications that preserve full underlying capability. Unlike web-based alternatives that require server-side processing and expose user data to third parties, Web-CLI applications execute entirely on the client, providing a verifiable privacy guarantee by architecture rather than policy. We define the pattern and its four properties: fidelity, progressive disclosure, offline-first, and zero egress. We present four reference implementations across distinct domains: ffmpeg-webCLI, a browser-based video editor built on FFmpeg; whisper-webCLI, speech transcription via Transformers.js; chat-webCLI, WebLLM-based language model inference; and 3mf-webCLI, a deterministic tool seg

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Structured State Reconciliation for Human-AI Task Handover

arXiv:2608.28907v1 Announce Type: new Abstract: Task handover requires communicating enough current state for a successor to resume work, yet the relevant information is often divided between system records and human observations. System records can be precise and timestamped but only partially observe the task, while human reports capture intent and task knowledge that no log contains but are vulnerable to omission and memory error. We present a provenance-aware pipeline that converts task telemetry and human-authored reports into a shared typed task-state representation, aligns and reconciles their facts, detects conflicts, and generates structured handover reports. We evaluate the approach on 13 paired task states collected in a controlled spatial multitask environment, using task-grounded metrics that estimate the state-reconstruction cost a report would spare a hypothetical recipient and the misinformation burden it would impose. Reconciling both sources preserved greater estimate

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Toward Postural State Classification in Immersive VR with Multimodal Data and Explainability Analysis

arXiv:2608.28844v1 Announce Type: new Abstract: Ensuring a safe virtual reality (VR) experience requires systems that can predict and respond when users lose their balance. Although prior work has examined fall prediction and motion sickness, many approaches are regression-based and postural state classification remains less explored. This study compares machine learning (ML) and deep learning (DL) models for classifying postural states in VR under visual perturbations. We used a multimodal dataset containing kinematic, electromyographic (EMG), and electrodermal activity (EDA) signals. The data were prepared for a binary task to distinguish balanced from imbalanced postural states, and participant-wise downsampling addressed class imbalance. All models were evaluated with Leave-One-Participant-Out (LOPO) cross-validation to test generalization to unseen participants. Among the models, the Mamba-inspired CNN (MI-CNN) achieved the highest accuracy of 96.76%. SHapley Additive exPlanations

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Delegating Before Learning: Where Generative AI Sits in Students' Professional Communication

arXiv:2608.28837v1 Announce Type: new Abstract: We conducted an interview study with twelve students on their use of generative AI in academic communication. Students delegated professional messages to AI most where the pressure to sound professional is highest: email to instructors and administrators. AI involvement ranged from correcting the writer's own text to working out and writing the message outright, and students checked AI-written text against two criteria: whether it looks like AI and whether it sounds like them. Building on these findings, we model the AI-mediated process of writing a student--instructor email at the highest level of involvement we observed, and compare it with an unaided model of writing the same messages, built from participants' accounts and a classic model of the writing process. Three differences emerge: the learning loop that builds writing skill is removed, the message is no longer written for its specific recipient, and the confidence a successful e

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Visible but Not Yet Curatable: Characterizing the Curatability of Compact and Derived Open LLM Artifacts

arXiv:2608.28819v1 Announce Type: new Abstract: Open Large Language Model (LLM) research increasingly produces compact and derived artifacts, such as adapters, quantized checkpoints, merged models, and distilled variants, that are distributed across papers, model hubs, model cards, code repositories, and release statements. Although these artifacts are publicly visible, digital libraries often lack sufficient evidence to identify, preserve, and cite them as coherent scholarly objects. We introduce a framework that conceptualizes curatability as a record-level property of distributed scholarly records and operationalizes it through four evidence dimensions: artifact identity, scholarly linkage, upstream evidence, and release assets. Guided by this framework, we conduct the first collection-scale characterization of open LLM curatability using a May 2026 snapshot of 191,375 public Hugging Face repositories and a core corpus of 2,214 scholarly papers. Our results reveal a pronounced visib

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Bringing Data to Life: Designing Data Characters for the Emotional Self

arXiv:2608.28780v1 Announce Type: new Abstract: Journaling is a common practice for emotional expression, reflection, and processing. However, as entries accumulate, it can become difficult to interpret and compare their affective content, especially since traditional text-based analyses and visualizations often struggle to convey affective nuance. We introduce Data Characters, a visualization approach that represents affective content in journaling through human-like characters. Using a customizable Data Character as a design probe, we investigate the potential of character-based representations for conveying affective experiences and explore what visual encodings emerge through customization. Preliminary walkthroughs with two participants demonstrate the intuitiveness and feasibility of the approach. This work contributes an exploratory approach to studying how affective experiences can be visually represented and encoded through anthropomorphic forms.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

A Conceptual Framework for Modeling Team Adaptation in Cooperative Games Through Ludic Knowledge

arXiv:2608.28729v1 Announce Type: new Abstract: With the increasing importance of teamwork skills for modern workplaces, development of teamwork training programs has received substantial attention. Game-based teamwork training is one promising approach that is engaging, cost-effective, and well-suited to increasingly decentralized workplaces. However, design of effective game-based teamwork training requires understanding how a game elicits specific desired teamwork behaviors. Significant progress has been made in characterizing these relationships. However, despite its critical importance, little work has examined how a game's design influences team adaptability behaviors. This paper presents a preliminary framework for analyzing adaptability in cooperative games by conceptualizing adaptive stimuli as retrieval or disruption of players' ludic knowledge. We illustrate this framework through a qualitative case study that applies interaction analysis methods to gameplay videos of a Over

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Beyond Helpfulness: A Teaching-over-Solving Diagnostic for Measuring Educational Impact in LLM Tutors

arXiv:2606.16206v2 Announce Type: replace-cross Abstract: Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learning support. Motivated by recent calls to measure the social impact of NLP systems in practice, we study whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. We propose a lightweight diagnostic based on the gap between solving-oriented and pedagogy-oriented benchmark performance. Using public MathTutorBench leaderboard results, we show that these dimensions are only partially aligned: across eight publicly reported models, the correlation between solving and pedagogy composites is 0.421, and several models shift meaningfully in rank when evaluation moves from solving to pedagogy. We then analyze the public TutorBench sample and show that agency-relevant behaviors are explicitly encoded in benchmark rubrics, especially in active-le

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

From AGI to ASI

arXiv:2606.12683v2 Announce Type: replace-cross Abstract: Over the last decade, building human-level artificial general intelligence has moved from far-fetched speculation to being a concrete next-decade target for many of the largest AI organisations. Achieving this goal would have profound and far-reaching impacts on human society, which raises many complex questions for the decade ahead. This report investigates how AI itself might continue to develop in a post-AGI world along the continuum of machine intelligence. The endpoint of this continuum, Universal AI, is theoretically well understood, which provides some formal grounding for the main focus of this report: the transition from human-level AGI to artificial general superintelligence, which can intuitively be understood as a system that is more intelligent and cognitively capable than large organisations of humans. After characterizing ASI, the report discusses four potential pathways from AGI to ASI: scaling AGI, AI paradigm s

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems

arXiv:2606.05985v2 Announce Type: replace-cross Abstract: Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural backgrounds. Existing cultural evaluation focuses on value alignment: how closely a single agent matches a target culture. Yet alignment is a per-agent property and cannot reveal whether a system, taken as a whole, preserves the cultural plurality it is meant to represent. We propose value diversity as a system-level evaluation axis for multicultural agent systems, defined through the dissimilarity between culturally conditioned agents' responses on a shared value survey. Using the World Values Survey, we evaluate 19 cultures and 18 backbone models across a wide range of system configurations. We find that diversity is largely uncorrelated with alignment, indicating that the two capture complementary system properties, and that current multicultural agent systems fall substantially b

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Vision-Language Models Suppress Female Representations Under Ambiguous Input

arXiv:2605.31556v2 Announce Type: replace-cross Abstract: Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far less is known about ambiguous inputs (a worker in full gear, a figure seen from behind), cases common in practice yet rarely studied. We find that minimal prompting pressure exposes occupation-gender defaults when prompting ambiguous input images, with models collapsing to male even for strongly female-stereotyped occupations. But do these outputs reflect what models actually encode internally? We introduce LALS (Latent Association Leaning Score), a zero-shot metric that projects visual-token activations into the model's text-embedding space to measure concept associations per token and layer. Across 15 occupations, over 800 gender-ambiguous images, and four VLMs, internal representations and outputs often become systematically decoupled: models often encode a female association int

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues

arXiv:2605.30051v2 Announce Type: replace-cross Abstract: A key part of developing large language model (LLM)-powered, automated tutoring tools is student simulation, i.e., using LLMs to role-play as students, which can facilitate tutor model evaluation and training. Existing work mostly focuses on within-dialogue simulation, which lacks context on student knowledge and behavior, partly due to not grounding in past student question-answering or dialogue interactions. In this work, we introduce the task of history-conditioned student simulation, where the goal is to accurately predict student dialogue turns by leveraging information in the student's learning history. We propose a two-component framework in which a profile generator summarizes a student's history and a simulator predicts student turns conditioned on the resulting profile. We train both components with reinforcement learning (RL), yielding profiles optimized for faithful student simulation. We evaluate our method and base

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

BIRDS: Characterizing and Understanding Biodiversity Impact of Large Language Model Serving

arXiv:2605.27480v3 Announce Type: replace-cross Abstract: Large language model (LLM) serving creates environmental impacts beyond carbon and water, including ecosystem damage through biodiversity-related pathways. We present BIRDS, a framework for Biodiversity Impact of Request-Driven LLM Serving. BIRDS defines request-level functional units, quantifies operational and embodied biodiversity impact, and introduces Quality-Normalized Biodiversity Impact (QNBI) to jointly analyze ecological impact and response quality. Across diverse workloads, models, GPUs, and regions, BIRDS reveals that biodiversity impact accumulates at scale and exposes quality-aware serving tradeoffs. The code is available at https://github.com/TianyaoShi/BIRDS.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

SafeMath: Inference-time Safety improves Math Accuracy

arXiv:2603.25201v2 Announce Type: replace-cross Abstract: Recent research points toward LLMs being manipulated through adversarial and seemingly benign inputs, resulting in harmful, biased, or policy-violating outputs. In this paper, we study an underexplored issue concerning harmful and toxic mathematical word problems. We show that math questions, particularly those framed as natural language narratives, can serve as a subtle medium for propagating biased, unethical, or psychologically harmful content, with heightened risks in educational settings involving children. To support a systematic study of this phenomenon, we introduce ToxicGSM, a dataset of 1.9k arithmetic problems in which harmful or sensitive context is embedded while preserving mathematically well-defined reasoning tasks. Using this dataset, we audit the behaviour of existing LLMs and analyse the trade-offs between safety enforcement and mathematical correctness. We further propose SafeMath -- a safety alignment techniq

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services

arXiv:2603.10807v2 Announce Type: replace-cross Abstract: Existing LLM safety evaluations rely on binary attack-success rates and domain-agnostic taxonomies, leaving regulated Banking, Financial Services, and Insurance (BFSI) deployments exposed to failures elicited through legally or professionally plausible framing. We introduce RAHS (Risk-Adjusted Harm Score), a risk-sensitive metric jointly capturing disclosure severity, disclaimer mitigation, and inter-judge agreement, and FinRedTeamBench, a 989-prompt benchmark spanning seven BFSI risk areas and 34 sub-categories mapped to regulatory frameworks. Evaluation uses an ensemble of three heterogeneous LLM judges, validated against human experts, and an adaptive multi-turn red-teaming pipeline. On nine open-weight models, RAHS preserves separation under near-ceiling ASR, ranking is stable under hyperparameter sweeps, and multi-turn pressure drives not only more jailbreaks but more operationally severe disclosures, exposing failure modes

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias

arXiv:2601.18486v3 Announce Type: replace-cross Abstract: Demographic cue-based evaluation is widely used to study how large language models (LLMs) adapt their responses to signaled demographic attributes within and across groups. This approach typically relies on a single cue (e.g., names) as a proxy for group membership, implicitly treating different cues as interchangeable operationalizations of a single underlying identity-conditioned behavior. We test this assumption in realistic advice-seeking interactions spanning 14.8 million prompts, focusing on race and gender in a U.S. context. We find that cues for the same group induce only partially overlapping changes in model responses, yielding inconsistent conclusions about personalization, while bias conclusions are unstable, with both magnitude and direction of group differences varying across cues. We further show that these inconsistencies reflect differences in cue-group association strength and linguistic features bundled within

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models

arXiv:2601.15436v3 Announce Type: replace-cross Abstract: We propose a novel perspective for probing LLM sycophancy in a direct and neutral way, mitigating various forms of uncontrolled bias, noise, or manipulative language, deliberately injected to prompts in prior works. A key novelty of our approach is the use of an LLM-as-a-judge in a zero-sum betting game. Within this framework, sycophancy serves one individual (the user) while explicitly incurring cost on another. Comparing 11 leading models we find that while most models exhibit significant sycophantic tendencies in the common setting, in which sycophancy is self-serving to the user and incurs no cost on others, seven of the models exhibit ``moral remorse'', five of which significantly over-compensate for their sycophancy in case it explicitly harms a third party. We refer to this phenomenon as `anti-sycophancy' bias and discuss possible causes for this shift.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs

arXiv:2601.03087v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM applications, but suffers from resource-intensive query access. We conceptualise auditing as uncertainty estimation over a target fairness metric and introduce BAFA, the Bounded Active Fairness Auditor for query-efficient auditing of black-box LLMs. BAFA maintains a version space of surrogate models consistent with queried scores and computes uncertainty intervals for fairness metrics (e.g., $\Delta$ AUC) via constrained empirical risk minimisation. Active query selection narrows these intervals to reduce estimation error. We evaluate BAFA on two standard fairness dataset case studies: \textsc{CivilComments} and \textsc{Bias-in-Bios}, comparing against stratified sampling, power sampling, and ablations. BAFA achieves target error thresholds with up to 40$\times$ fewer queries than str

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Classification Drives Geographic Bias in Street Scene Segmentation

arXiv:2412.11061v2 Announce Type: replace-cross Abstract: Previous studies showed that image datasets lacking geographic diversity can lead to biased performance in models trained on them. While earlier work studied general-purpose image datasets (e.g., ImageNet) and simple tasks like image recognition, we investigated geo-biases in real-world driving datasets on a more complex task: instance segmentation. We examined if instance segmentation models trained on European driving scenes (Eurocentric models) are geo-biased. Consistent with previous work, we found that Eurocentric models were geo-biased. Interestingly, we found that geo-biases came from classification errors rather than localization errors, with classification errors alone contributing 10-90% of the geo-biases in segmentation and 19-88% of the geo-biases in detection. This showed that while classification is geo-biased, localization (including detection and segmentation) is geographically robust. Our findings show that in r

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Overreliance on AI in Information-seeking from Video Content

arXiv:2603.19843v2 Announce Type: replace Abstract: The ubiquity of multimedia content is reshaping online information spaces, particularly in social media environments. At the same time, search is being rapidly transformed by generative AI, with large language models (LLMs) routinely deployed as intermediaries between users and multimedia content to retrieve and summarize information. Despite their growing influence, the impact of LLM inaccuracies and potential vulnerabilities on multimedia information-seeking tasks remains largely unexplored. We investigate how generative AI affects accuracy, efficiency, and confidence in information retrieval from videos. We conduct an experiment with around 900 participants on 8,000+ video-based information-seeking tasks, comparing behavior across three conditions: (1) access to videos only, (2) access to videos with LLM-based AI assistance, and (3) access to videos with a deceiving AI assistant designed to provide false answers. We find that AI as

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Landscape of Generative AI in Information Systems: A Synthesis of Secondary Reviews and Research Agendas

arXiv:2603.11842v2 Announce Type: replace Abstract: The post-ChatGPT surge has rapidly reframed IS research and practice. As organizations and society grapple with GenAI adoption, a body of secondary studies and research agendas has emerged to synthesize early evidence and chart directions for future inquiry. This study reviews secondary and roadmap papers to synthesize the state of knowledge on GenAI's benefits and challenges in IS, and to identify future research directions. We performed a systematic search across Scopus, WoS, and eAIS for publications from 2023 onwards. Following a rigorous, multi-stage screening process, we selected a final set of 28 papers for analysis using bibliometric mapping and thematic analysis. We also conducted a quality assessment of all sources to gauge confidence in each source's contribution to the findings. GenAI offers transformative potential to drive productivity, accelerate innovation, personalize services, and democratize access to expertise. How

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

ACE-Align: Attribute Causal Effect Alignment for Cultural Values under Varying Persona Granularities

arXiv:2601.12962v2 Announce Type: replace Abstract: Ensuring that large language models (LLMs) reflect diverse cultural values is important for globally deployed NLP systems. However, existing approaches often treat cultural groups as homogeneous and overlook within-group heterogeneity arising from intersecting demographic attributes, leading to unstable behavior under varying persona granularity. To address this gap, we propose ACE-Align (Atribute Causal Effect Alignment), a causally inspired framework based on controlled persona edits that aligns how specific demographic attributes shift different cultural values, rather than treating each culture as a homogeneous group. We evaluate ACE-Align across 14 countries spanning five continents, with personas specified by subsets of four attributes (gender, education, residence, and marital status) and granularity instantiated by the number of specified attributes. Across all persona granularities, ACE-Align consistently outperforms baseline

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Bathtubs, Boundaries, and Sandboxes: AI Regulatory Learning under Legal Uncertainty

arXiv:2601.04094v4 Announce Type: replace Abstract: Effective regulation of AI is a defining policy challenge, driven by their integration into all aspects of society. To remain responsive to their rapid development and emergent properties, policymakers across the globe rely on high-level principles and abstract legal requirements. Yet, while this flexibility supports future-proofing human-centred regulations and aligning them with socio-ethical values, it also causes legal uncertainty downstream as developers, companies, and auditors struggle with translating these abstract requirements into verifiable technical requirements. Using the AI Act as an example, this paper draws on Coleman's bathtub to analyse the regulatory learning space in AI governance. It argues that legal uncertainty cannot be fully reduced ex ante and that, within reasonable bounds, it is also necessary for regulatory learning because it creates the space in which boundary negotiation over socio-technical meaning ca

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Retrofitters, pragmatists and activists: Public interest litigation for accountable automated decision-making

arXiv:2511.03211v5 Announce Type: replace Abstract: This paper examines the role of public interest litigation in promoting accountability for AI and automated decision-making (ADM) in Australia. Since ADM regulation faces political and geopolitical headwinds, effective governance will have to rely on the enforcement of existing laws. Drawing on interviews with Australian public interest litigators, technology policy activists, and technology law scholars, the paper positions public interest litigation as part of a larger ecosystem for transparency, accountability and justice with respect to ADM. The paper explores the tactics and strategies of what one participant described as 'retrofitting' old laws to ADM. These go beyond creative legal argumentation, to encompass practices of community-building, collaboration on theories of change, canny selection of clients and causes of action, and aligning the interests of stakeholders in litigation. Naturally, the paper also contends with the l

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES)

arXiv:2510.24958v3 Announce Type: replace Abstract: The evaluation of societal biases in NLP models is critically hindered by a geo-cultural gap, This leaves regions such as Latin America severely underserved, making it impossible to adequately assess or mitigate the perpetuation of harmful regional stereotypes in language technologies. This paper presents LACES, a stereotype association dataset, for 15 Latin American countries. This dataset includes 4,789 stereotype associations manually created and annotated by 83 participants. The dataset was developed through targeted community partnerships across Latin America. Additionally, in this paper, we propose a novel adaptive data collection methodology that uniquely integrates the sourcing of new stereotype entries and the validation of existing data within a single, unified workflow. This approach results in a resource with more unique stereotypes than previous static collection methods, enabling a more efficient stereotype collection. T

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Operationalising AI Regulatory Sandboxes: Activities, Requirements, and Technical Assessment under the EU AI Act

arXiv:2509.25256v4 Announce Type: replace Abstract: The systematic assessment of AI systems is increasingly vital as these technologies enter high-stakes domains. To address this, the EU's Artificial Intelligence Act introduces AI Regulatory Sandboxes (AIRS): supervised environments where AI systems can be tested under the oversight of Competent Authorities (CAs), balancing innovation with compliance, particularly for startups and SMEs. Yet significant challenges remain: assessment methods are fragmented, tests lack standardisation, and feedback loops between developers and regulators are weak. This paper operationalises the AIRS lifecycle. We map the sandbox journey into 29 concrete activities, from pre-participation guidance through application, preparation, participation, exit, and post-participation monitoring, and we distinguish between a Core AIRS centred on regulatory oversight and an Extended AIRS that additionally embeds structured technical testing through an AI Technical San

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Emerging Media Use and Acceptance of Digital Immortality: A Cluster Analysis among Chinese Young Generations

arXiv:2505.01355v2 Announce Type: replace Abstract: Digital immortality is increasingly discussed as a technological possibility, yet empirical evidence about potential users' evaluations remains limited. We surveyed 462 Chinese young adults, combining cluster analysis of four emerging-media use frequencies with valence coding of open-ended responses to physical death, physical immortality, digital immortality, and digital death. Three profiles emerged: broad emerging-media, gaming-focused, and low-use users. Broad emerging-media users reported the highest adjusted acceptance and scored higher on several personality and worldview measures, while fear of death did not differ across profiles. Physical immortality elicited the most negative responses; digital immortality produced more mixed, less negative appraisals. More favorable digital-immortality appraisals predicted higher acceptance after adjustment for media-use profile and demographics, although scenario valence did not differ ac

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Multimodal Large Language Models Predict Urban Safety Perception but Encode Non-Neutral Demographic Priors

arXiv:2503.00610v2 Announce Type: replace Abstract: Understanding how people perceive urban environments is essential for inclusive planning, yet conventional surveys are costly and difficult to scale. We investigate whether Multimodal Large Language Models (MLLMs) can assess perceived urban safety from street-view imagery while accounting for the observer-dependent nature of perception. Using Place Pulse 2.0, we evaluate four open and proprietary MLLMs across 56 cities under a Neutral prompt and socio-demographic personas defined by gender, age, and race or ethnicity. We also analyse the keywords generated to justify each classification. All four models display comparable zero-shot capability, with city-macro F1 scores of 65--69%, and preserve meaningful cross-city variation. However, they systematically favour the Safe class, underpredict unsafety, and compress differences between cities. Their explanations converge on a shared visual lexicon: maintenance, greenery, order, and reside

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

arXiv:2407.20371v3 Announce Type: replace Abstract: Artificial intelligence (AI) hiring tools have revolutionized resume screening, and large language models (LLMs) have the potential to do the same. However, given the biases which are embedded within LLMs, it is unclear whether they can be used in this scenario without disadvantaging groups based on their protected attributes. In this work, we investigate the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection. Using that framework, we then perform a resume audit study to determine whether a selection of Massive Text Embedding (MTE) models are biased in resume screening scenarios. We simulate this for nine occupations, using a collection of over 500 publicly available resumes and 500 job descriptions. We find that the MTEs are biased, significantly favoring White-associated names in 85.1\% of cases and female-associated names in only 11.1\% of cases, with

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

One note in three: a verified census of three deployed AI scribes, and the instrument that counted it

arXiv:2608.31017v1 Announce Type: cross Abstract: Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

MusGU+: Toward a Musician-Centered Evaluation Framework and Discovery Tool for Generative Music AI

arXiv:2608.30940v1 Announce Type: cross Abstract: Generative music systems are increasingly presented as tools that democratize music creation, yet their practical suitability for musicians remains underexplored. Prior work includes openness-focused evaluation frameworks, such as MusGO (Music-Generative Open AI), as well as qualitative studies of musicians' experiences with generative systems. However, these approaches do not support systematic comparison or early-stage discovery of models for creative use. Motivated by such limitations, we introduce MusGU+, a musician-centered evaluation framework organized around three dimensions: Adaptability, Usability, and Controllability. Together, these capture whether a model can be feasibly trained or fine-tuned on personal data, integrated into real-world music workflows, and controlled in musically meaningful ways. We evaluate 10 representative generative music systems and present an interactive discovery tool that enables musicians to explo

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Which Rules Matter Now? Policy-Centroid Routing Before an Intelligent System Acts

arXiv:2608.30757v1 Announce Type: cross Abstract: Before an intelligent system can decide whether an action is allowed, it must first know which rules the action has approached. A single proposed action can implicate several policy regimes at once. Their requirements may stack, overlap, or qualify one another, yet many remain written in natural language while the action itself arrives as an incomplete description of intent. The first problem is not judgment. It is attention. Policy-centroid routing creates a layer before adjudication. It compresses expressions within each policy regime into one or more representative centroids, places the proposed action in the same semantic space, applies a declared measure, and routes every regime crossing a declared threshold to authoritative review. Several regimes may trigger at once. The output is a review agenda, not permission, prohibition, legality, breach, compliance, certification, or enforcement. The paper develops six falsifiable propositi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

WildSEEK: Evaluating Language Models for Information-Seeking

arXiv:2608.30683v1 Announce Type: cross Abstract: Language models are increasingly mediating information access to end users, urging a systematic evaluation of their responses for a fair and reliable information ecosystem. Existing evaluations, however, are often topic-specific or synthetic, limiting their ability to capture the complexity of "in the wild" information-seeking queries and the risks present in model responses. To address this gap, we introduce WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses. WildSEEK includes annotations for risk-sensitive domains (e.g. health and financial information), and distinguishes factoid queries from analytical queries which seek responses beyond facts. We train classifiers on WildSEEK to analyze more than 1.8M realistic user queries. We find that over a third of information-seeking queries are high-risk and more often analytical. Our fi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Two Centuries of Sexism in British Parliament: A Computational Analysis of Women's Representation in the Hansard Corpus

arXiv:2608.30485v1 Announce Type: cross Abstract: The language a legislature uses to debate women's rights, even in favour of them, encodes systematic patterns of sexism that persist across two centuries. In this work, we analyse 6,531 speeches over 200 years of UK parliamentary debate (Hansard, 1803-2005) by using large language models to classify a speaker's perspective towards women's suffrage and political representation, as well as analyse sexist speech in parliament from the lens of the Ambivalent Sexism Inventory. We also release this parliamentary dataset, an organized and metadata-enriched version of the publicly available Hansard Corpus optimized for computational social science research, with 6.7 million speeches across 1.2 million debates, with 89% gender-matching for speeches by MPs from the House of Commons. We find that 54% of speeches opposing women's representation contain sexist content, compared to 21% of speeches that are for the cause, and that the two sides use fu

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP

arXiv:2608.30107v1 Announce Type: cross Abstract: Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and mot

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Language of the Question Selects the Market: Query Language and Exit IP as Separable Factors in Commercial Recommendations from a Generative Search Interface

arXiv:2608.30052v1 Announce Type: cross Abstract: When a generative search interface answers a commercial question, which market's products it names is decided before the model reasons about the products. We report a controlled probe of 234 runs against the logged-out ChatGPT web interface and the OpenAI API, collected on 29 and 30 August 2026 across four exit countries and six query languages, with six identical runs per cell. Three results. First, the top recommendation is unstable: it changed across six identical runs on four of six prompts, and that rate was identical in the browser interface and in the API with web search both enabled and disabled, so instability is a property of the system and not of the surface. Second, query language, and not location, decides whether local suppliers appear at all. Where the query language matched the country, a global brand won 1 of 24 runs; asked in English on the same connections, local brands took 0 of 6 runs in Estonia and Turkiye. Third,

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Text Data Analysis and Classification Methods - Insights from Customer Letters in Life Insurance

arXiv:2608.29699v1 Announce Type: cross Abstract: The business of life insurance companies is characterized by long-term contracts. For this reason, data describing customers is of immense value. A portion of the data provided to the customer is rarely or not at all analyzed. This includes customer letters of any kind. This work focuses on classifying customer letters as cancellations and identifying the respective reason, if available. The outlined approach can also be applied to other business transactions and reasons. We discuss data acquisition and preparation, present alternatives, and explain the reasons for the chosen approach. A successful implementation of such a tool can lead to a better understanding of customer cancellation behavior by the insurer, enabling more targeted actions in certain situations.

Source ↗
Showing 6351–6400 of 18402 signals
← Prev Page 128 of 369 Next →