EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Proceedings of The First Reflection in Creative Experience (RiCE) Workshop

arXiv:2607.24558v1 Announce Type: new Abstract: Reflection and metacognition are central to the creative user experience. However, most HCI research on reflection focuses on clear, task-oriented goals such as to reflect on personal data or pedagogical outcomes. This contrasts with the open-ended and challenging to articulate goals of creative user experiences. For the first time, this workshop brings together interdisciplinary researchers, designers, educators, and artists across HCI, Cognitive Science, Design, AI, Learning Sciences, and Digital Art to examine reflection in creative interaction. The workshop will discuss themes, drawn from earlier discussions with HCI researchers and artists, on: how best to capture reflection in creative contexts, how to leverage the arts to support reflection for ethical change, and how to design creative AI that enhances - not hinders - critical thinking. By bringing interdisciplinary perspectives on reflection into discussion, the workshop will dev

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

RemiAssist: A Therapist-Supporting System for Photo-Based Reminiscence Therapy in Dementia Care

arXiv:2607.24536v1 Announce Type: new Abstract: Despite growing interest in applying AI to photo-based reminiscence therapy (PRT) for people with dementia (PwD), existing systems primarily focus on PwD-AI interaction and often overlook therapists' critical role in practical PRT delivery. We present RemiAssist, a system that supports therapist-in-the-loop PRT through AI-assisted planning and real-time facilitation. RemiAssist incorporates two core techniques: (1) a Memory Graph, which organizes key life events from a PwD's photo collection into a hierarchical graph to support theme-centered intervention planning; and (2) a Context-Aware Guiding Strategy, which provides real-time suggestions to help therapists guide reminiscence conversations and respond to sensitive situations. A field study with eight therapist-PwD dyads suggests that RemiAssist was associated with a 44% improvement in planning efficiency and a 54% increase in conversation duration, and provided timely support for hand

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis

arXiv:2607.24430v1 Announce Type: new Abstract: Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, facial expressions encode subtle and rich affective cues that are crucial for empathetic speech interaction, whereas existing approaches often overlook this important modality. In addition, the lack of large-scale natural conversational datasets with both speech and visual modalities also limits the development of visual affect understanding in conversational settings.To address these limitations, we propose FacialTalker, a facial-expression-aware CSS framework built upon a large language model backbone. To efficiently encode facial expressions, we propose AUTokenizer, a single-codebook visual tokenizer that discretizes each frame-level facial expression into a compact token, trained with supervision from combinations of facial Action Units. We further introduce

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Order-Bound Companionship: The Practice of Emotional Labor in Professional Game Companionship

arXiv:2607.24363v1 Announce Type: new Abstract: Labor in platform gig economy increasingly involves services involving relationship that demand significant emotional investment. Grounded in China's unique socio-cultural and multi-platform context, this study explores professional game companionship, an under-explored digital labor practice. Through interviews with 22 game companionship practitioners, we used a micro-level perspective to relational gig work to analyze how workers navigate intimate boundaries and stakeholder networks. We found that companions adopt an "order-bound" mechanism: performing immersive deep acting during paid sessions, followed by complete emotional disengagement post-order. We also identified a tripartite companion-centric network featuring scenario-based performances with clients, competitive-symbiotic peer relations, and interdependent governance with companionship clubs. Furthermore, significant identity fluidity exists, with individuals frequently transit

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

EEGForceFusion: Joint Tokenised-Continuous Representation Learning for Subject-Independent Grasp Force Decoding

arXiv:2607.24126v1 Announce Type: new Abstract: Brain-machine interfaces provide a link between neural activity and external devices, enabling restoration of motor function and advancing human-machine interaction using non-invasive electroencephalography (EEG). However, continuous grasp force decoding remains challenging due to complex temporal dynamics, high inter-subject variability, and limited generalisation of existing approaches. To address this, we propose a hybrid EEG decoding framework that jointly models continuous and tokenised representations, enabling capture of both fine-grained neural structure and long-range temporal dependencies. The proposed approach integrates convolutional-recurrent representation learning, quantisation-based tokenisation, and transformer-based temporal modelling within a unified fusion-based regression architecture. Experimental evaluation on the WAY-EEG-GAL dataset under strict leave-one-subject-out conditions achieves $R^2$ = 0.817 in offline set

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Towards simultaneous decoding of kinetic and kinematic movement parameters during grasp and lift task by noninvasive brain imaging

arXiv:2607.24081v1 Announce Type: new Abstract: Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke survivors or amputees. One of the key challenges in developing BMIs is expanding their usability and control, which can be achieved by accurately decoding multiple kinematic and kinetic parameters. To address this, we propose three regression models: partial least squares regressor, multilayered perceptron, and attention based regressor, to decode multiple movement parameters from EEG signals. We evaluated these models on the WAY EEG GAL dataset, focusing on their performance under subject specific and subject independent conditions with two strategies: a single model for all parameters and a baseline with separate models for each parameter. Among all regressors, the attention based regressor achieved the best performance, with an $R^2$ of 0.8 and a latency of 29.2 milliseconds, demonstrating significant improvement in simultaneous multi parameter

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

The Effect of Photorealism Consistency between the Virtual Hands and Environment on the Sense of Body Ownership and Presence in Virtual Reality

arXiv:2607.24047v1 Announce Type: new Abstract: Virtual reality (VR) technology allows users to feel a virtual body as if it was their own (i.e., body ownership illusion). Previous studies have explored how the visual realism of the user's virtual body influences the sense of body ownership in VR, focusing on the aspect of anthropomorphism. However, the effect of photorealism, another element that characterizes the visual realism of a virtual human, has not been systematically examined in the context of body ownership illusion. Therefore, we investigate the effect of the rendering style of virtual hands on the sense of body ownership, hypothesizing that the effect is affected by the rendering style of the virtual environment. In addition, we examine the effect of photorealism on presence (i.e., the sense of being there), as existing studies have offered inconsistent evidence. To this aim, we conducted a 3 x 3 mixed-design remote VR experiment (N=117) that factored in the rendering styl

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

A Design Space for Quantum Circuit Visualizations

arXiv:2607.24042v1 Announce Type: new Abstract: Quantum circuit visualizations play an essential role in supporting sense-making and communication of quantum programs. While several tools exist for rendering quantum circuits, they vary widely in encoding options due to idiosyncrasies among machine and platform providers. We observe an opportunity to coalesce these disparate rendering approaches under a single, unified grammar to enable consistent, cross-platform enhancement of quantum circuit visualizations. However, it is unclear how to design such a grammar to best support the quantum computing community. Towards this end, we contribute a design space of quantum circuit visualizations by analyzing 182 static and 12 interactive cases collected from online tutorials and documentations, research publications, public presentations, and prior systems. Based on our analysis, we discuss how our design space relates to existing visualization principles yet exhibits unique aspects. We conclud

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

SHARE: Towards Head-Mounted AR with User-Centric SLAM in Shared Human-Robot Workspaces

arXiv:2607.23901v1 Announce Type: new Abstract: Human-Robot Collaboration (HRC) in shared physical spaces using Augmented Reality (AR) interfaces is powered by Simultaneous Localization and Mapping (SLAM). Existing multi-agent SLAM systems rely on an edge server to combine visual findings of multiple resource-constrained agents, perform computation, and schedule updates to their local maps. However, the edge treats all agents uniformly and ignores the fundamentally different latency requirements of heterogeneous HRC agents: robots and head-mounted AR users. This uniform resource allocation often results in high lag for user manipulation, as it does not meet the stringent latency requirements of AR. In this work, we design, implement, and evaluate SHARE, a user-centric SLAM system that strategically prioritizes AR user experience while maintaining accurate tracking performance for robots. SHARE builds a first-of-its-kind experience model for HRC agents and adaptively adjusts transmissio

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Evaluating Closed-Loop EEG Feedback for Simulated Prosthetic Vision in Immersive VR: A Sham-Controlled Feasibility Study

arXiv:2607.23889v1 Announce Type: new Abstract: Visual prostheses require users to interpret sparse and distorted artificial percepts through active visual search. We developed an EEG-guided neuroadaptive training platform for simulated prosthetic vision in immersive virtual reality and evaluated its feasibility in a sham-controlled object-localization task. Twenty-two sighted participants searched a virtual desk scene rendered through a low-resolution phosphene simulation while EEG was recorded using a dry-electrode headset integrated with a head-mounted display. During training, participants received post-trial visual feedback based either on a commonly used EEG engagement index, $\beta/(\alpha+\theta)$, or on visually matched non-contingent sham values. Both groups showed comparable within-session improvements in localization performance, consistent with practice, increasing familiarity with the simulated percepts, or refinement of search strategies. EEG-contingent feedback did not

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Meandering Photo Lines: Fluid Multiscale Exploration of Photographic Archives

arXiv:2607.23769v1 Announce Type: new Abstract: We introduce a visualization that interweaves spatial and temporal structures of photo collections into explorable pathways. Existing photo browsing interfaces typically use space and time separately for sorting, filtering, or mapping photos. What remains missing is an integrated visual form that represents the photographic journey itself. This is particularly relevant for the individual photographer's archive, where their unique lifeline is written into the collection, and it is this spatiotemporal trajectory that gives necessary context and meaning to the work. Inspired by the winding nature of rivers, fluid interaction, and un/foldable visualizations, Meandering Photo Lines visualizes adjustable pathways through a photo collection across space, time, and topics. Two complementary views enable visual exploration at different scales: A grid-based arrangement of photo clusters offers a more analytical overview, while meandering spatiotemp

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

arXiv:2607.23670v1 Announce Type: new Abstract: Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to develop a plan before task execution. However, it remains unclear whether the benefits of this feature translate to end-user programming environments such as spreadsheets. Since spreadsheet programmers tend to work iteratively and care less about technical correctness, upfront planning may not fit into their workflows as easily. In this paper, we build a prototype of a Plan Mode for spreadsheet programming and evaluate it against a non-planning baseline through a within-subjects user study (N=24). We found that despite similar task outcomes with both tools, using Plan Mode led to a reduction in refinement and a better perception of the tool across dimensions of creativity support and human-machine collaboration. We discuss the implications of these results for the future design of Plan Modes,

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Multimodal Data Comprehension: Understanding How Visual-Textual Chains of Information Influence Data Interpretation

arXiv:2607.23489v1 Announce Type: new Abstract: Visualizations and text often work together to support effective data communication. Despite this common paradigm, we know little about how the interplay of these modalities affects people's data comprehension. We present a novel experimental paradigm to investigate multimodal data comprehension---the process of people comprehending information from multimodal visual and textual data---across both crowdsourced and think-aloud environments. Our methodology employs two sequential chains for presenting multimodal information---a visualization-first chain and a text-first chain---asking people to describe the data presented iteratively. By comparing how people's data comprehension changes across the chain, we can assess the information contribution of each modality and how they shape subsequent comprehension. We found that the visualization-first chain facilitates exploratory comprehension with hypothesis-driven discovery, whereas the text-fi

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Visualizing in the Mind's Eye: Icon Design Shapes Mental Imagery of Fire Risks

arXiv:2607.23369v1 Announce Type: new Abstract: We introduce mental imagery, or seeing images in the "mind's eye," as a cognitive process that can be shaped by data visualization design and in turn impact decisions. We found in a preregistered study (n = 400) that abstract geometric icons visualizing fire risk data evoked more mental images than concrete house-on-fire icons, and also produced more diverse and personalized mental images. Mediation analysis showed that increased mental imagery subsequently led to risk-averse decisions through evoking negative affect. These findings reveal a nuanced mechanism through which visualization concreteness influences decisions: concrete designs may actually suppress affect-driven behavior by restricting mental imagery.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models

arXiv:2607.23309v1 Announce Type: new Abstract: Recent work in Human-Computer Interaction (HCI) increasingly treats AI models as design materials that have distinctive computational properties to shape design artifacts. Artists learn to work with the model "at play" to explore their emerging properties. The aim of explainability, in this view, is to make visible a crafting and hacking space to enable sustained creative practices with AI. In this chapter, we propose material explainability as a range of activities and artifacts that transform AI models into accessible and inclusive design materials in the workspace of artists, designers, and makers. We present a case study of building a repository of resources to enable artistic explorations of neural audio models in New Interfaces for Musical Expression (NIME) design. Reflecting on our community-building journey and the making of a collection of musical interface designs with a group of artists, we raise three recommendations on enabli

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Cross Sensory Co-Design Tools and Interaction Qualities

arXiv:2607.23298v1 Announce Type: new Abstract: What is the temperature of a loud cry? What is the color of a hand-clap? Cross-sensory and multi-modal interactions offer interesting possibilities in terms of inputs and outputs that might be leveraged for affective and emotional purposes. However, it is hard to design such interactions without the right co-design tools. We developed such tools in the last 10 years: the Loaded Dice and the Wheel of Plush. Both devices incorporate a multitude of sensors and actuators to demonstrate interactions and aid in co-design workshops. The tools incorporate the principle of technical synesthesia: the ability to map any of the included sensors to any of the included actuators. With intuitive simple interaction schemes it is possible to easily and effortlessly create new combinations. We discuss the core principles of the tools and how we realized a meaningful mapping between sensors and actuators. We further discuss how we use the tools for explorat

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

A Taxonomy of Confabulations and the Perception-Reality Gap in LLM-Assisted Immersive Scene Editing

arXiv:2607.23213v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly integrated into immersive environments and design workflows, providing application prospects in areas such as rapid scene prototyping for non-expert users and scene understanding capabilities for accessibility design. While many workflows that incorporate LLMs in immersive spaces are proposed, such systems can exhibit errors, potentially resulting in frustration, loss of user trust, and compromised user safety. This paper studies the underexplored area of LLM confabulations in immersive 3D scene editing contexts. Through an exploratory study with 24 non-expert users, we construct a taxonomy of the different types of confabulation observed in LLM-assisted immersive 3D scene editing. We report their prevalence and disruptiveness, and define the construct perception-reality gap to help understand the gap between the actual and perceived occurrence of confabulations. We highlight the observe

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Beyond Conversations: Spatially-Anchored Previews for Intent Disambiguation in LLM-Assisted Geometry Editing in Virtual Reality

arXiv:2607.23201v1 Announce Type: new Abstract: User intent disambiguation remains a key challenge in intelligent interactive systems. While they have been widely studied in dialogue systems in 2D interfaces, research on how intent disambiguation could be incorporated within Large Language Model (LLM) assisted editing workflows in immersive environments remains limited. Recent advances in LLMs create opportunities to leverage the immersive nature of virtual and augmented reality (VR/AR) environments to provide better disambiguation support. In this paper, we evaluate how traditional dialogue-based disambiguation can be augmented with spatially-anchored graphical previews to resolve ambiguous user commands in LLM-assisted parameter-driven editing workflows. A within-subjects study in which 24 participants completed complex geometry editing tasks in VR simulate scenarios where VR scenes are controlled by numerical parameters. Compared with the condition where disambiguation is not availa

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

From Vibe to Code -- and Back: Lexical Oscillation in the Formation of Design Intent with Generative AI

arXiv:2607.23126v1 Announce Type: new Abstract: Generative AI design tools make natural-language prompts a starting point for design, placing new articulation demands on designers. Rather than treating prompts as the transmission of pre-existing design intent, we ask how design intent is formed through situated interaction with AI. Five expert UI/UX designers (11-20 years' experience, M = 15.4) designed landing-page hero sections with a generative AI tool, recorded through think-aloud and retrospective interviews. Using reflexive thematic analysis, we used lexical granularity (L1 vibe, L2 design-domain, L3 operational language) as a sensitizing lens. Rather than moving from vibe to code unidirectionally, designers showed lexical oscillation, including returns from operational specificity to ambiguity. Mismatches with AI outputs were taken up as occasions for designers to reconsider what they meant, and engagement shifted from instruction to consultation. One non-oscillating trajectory

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Touching or Chatting: The Utility of LLMs and Tactile Charts for Learning about Complex Chart Types by BLV Individuals

arXiv:2607.23065v1 Announce Type: new Abstract: Visualizations are central to communicating data, yet blind and low-vision (BLV) people often lack support for understanding chart types---knowledge that is essential for interpreting new visualizations and collaborating with sighted peers. Prior work found that BLV individuals viewed example tactile charts as more helpful than text-only approaches and preferred them for learning advanced chart types, particularly for understanding spatial layouts and shapes. Meanwhile, large language models (LLMs) are increasingly used by BLV individuals for chart explanation and question answering (QA), but have been studied primarily for dataset exploration rather than chart-type learning. Existing LLM-based chart QA also shows that users frequently ask about layout and structure, yet struggle with spatial concepts and misdirect questions when mental models are weak. We investigate how LLMs influence chart-type learning and whether tactile learning imp

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

The Help Ladder: Skill-Adaptive Peer Scaffolding for Real-Time Collaborative Programming

arXiv:2607.23031v1 Announce Type: new Abstract: Collaborative programming is a widely adopted classroom activity to encourage peer scaffolding, yet real-time collaboration often breaks down into parallel individual work with minimal interaction. Our formative studies reveal that even when students want to collaborate, they are held back by the effort required to understand a teammate's entire problem at once. We present Canary, a system that supports peer scaffolding by breaking down programming obstacles into smaller steps tailored to a student's skill level. Canary alerts potential helpers to specific places where they can start, using AI to turn complex problems into a step-by-step ladder that starts with easy fixes before moving toward harder logic. By providing this gradual ramp-up, Canary enables students to make quick contributions and progressively work toward solving their peers' problems. Our evaluation shows that this staged approach makes helping feel less overwhelming, lea

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Reflections and Recommendations on AI Adoption Practice from a Mixed-Ability Research Group

arXiv:2607.22886v1 Announce Type: new Abstract: Generative AI tools have recently been rapidly adopted by academics in mixed-ability research teams for both personal and professional tasks. While previous work on adoption of AI-based workflows has focused on collaboration and productivity, the perceptions of AI use within research teams remains divided. Through qualitative analysis of interviews of the five members of our mixed-ability research team, we discuss the motivations, challenges, and practices surrounding the use of generative AI in our lab. We reflect on experiences that shaped recommendations for balanced AI use that enable mixed-ability team workflows: (1) managing disability tax & crip time, (2) homogenizing identity, (3) risk disclosure of private information, (4) self-experimentation and miscellaneous tasks, and (5) information seeking. We build upon these themes to present AI practice recommendations we established for our lab to promote AI workflow adoption while pres

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

Validation of a Real-Time Manual Wheelchair Simulator through Biomechanical and Perceptual Measures: A Comparison with Overground Propulsion

arXiv:2607.22851v1 Announce Type: new Abstract: Purpose: Manual wheelchair simulators can provide a safe and controlled environment for propulsion training, biofeedback, and biomechanical assessment. However, their ecological validity must be established to ensure that simulator-based outcomes reflect overground propulsion. This study aimed to evaluate the ecological validity of a real-time manual wheelchair simulator by comparing biomechanical and perceptual outcomes during matched overground and simulator tasks. Materials and Methods: Thirteen participants (10 able-bodied individuals, 3 experienced manual wheelchair users), performed propulsion tasks overground and on the simulator. Task segments included straight-line propulsion, acceleration, turning, ascending, and descending. Biomechanical outcomes were collected using instrumented wheels and compared between environments using temporal, kinetic, and waveform-based measures. Perceived realism and user experience were assessed usi

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.HC

GeoTEAM: A Geospatial Tangible User Interface for Exploration and Visual Analysis of Migration Data

arXiv:2607.22825v1 Announce Type: new Abstract: Migration studies is an interdisciplinary field, requiring collaboration between researchers with varying levels of geospatial and quantitative data literacy. Geospatial tangible user interfaces (GTUIs) offer promising opportunities for embodied and collaborative spatial exploration of migration data. Few GTUIs, however, provide real-time visual feedback of data values to help users collaboratively identify trends. To address this gap, we present GeoTEAM, a novel tangible system for real-time, dynamic regional exploration, co-designed with migration and HCI researchers. GeoTEAM features active tangible dials with built-in touchscreens designed to control time and navigate map layers for net migration and various environmental sub-drivers in a multi-surface environment with tabletop and wall displays. We evaluated this system with nine pairs of researchers possessing multi-expertise in geospatial data literacy. Qualitative findings show th

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

arXiv:2606.13385v2 Announce Type: replace-cross Abstract: LLM-based web agents are increasingly deployed in real-world settings such as e-commerce, where they interact extensively with untrusted web content while executing actions that carry direct financial consequences. This makes them vulnerable to prompt-injection attacks, in which seemingly benign web content conceals adversarial instructions that manipulate the agent's behavior. Existing security benchmarks adopt an \textit{attack-centric} perspective, focusing on the technical feasibility of injections while overlooking the nuanced distribution of resulting harms. In practice, however, prompt-injection risk is victim-dependent: a single exploit can produce asymmetric consequences for different stakeholders, and the same attack pattern may exhibit substantially different effectiveness depending on whom it targets. To capture these properties, we introduce StakeBench, a stakeholder-centric benchmark that systematically categorizes

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

arXiv:2606.09833v2 Announce Type: replace-cross Abstract: AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent collaboration both in preserving human agency and generating economic value, this paradigm remains largely absent from occupational task evaluation, hindered by the difficulty of gathering real human data and accounting for inter-human variability. We introduce CollabSkill, a framework for evaluating human-agent collaboration on real-world occupational tasks. CollabSkill pairs real human workers with AI agents on tasks matched to their occupational background, collecting data that capture the complexity of economically valuable tasks and the usage patterns of real workers. To account for inter-human variability, CollabSkill employs a Bayesian skill rating system to disentangle and quantify the skill contributions of both humans and AI agents. Drawing on over 1,500 prompts from 386 working session

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Beyond Adoption Intention How Trust in Augmented Analytics Relates to Perceived Decision Quality Among Non-Technical BI Users

arXiv:2605.20198v2 Announce Type: replace-cross Abstract: Augmented analytics has transformed how Business Intelligence (BI) systems support decision-making, shifting non-technical managers from manual analysis toward dependence on automated insights. Current BI research often overlooks the cognitive mechanisms and the direct impact of AI-enabled analytics on decision quality. This study employs the theory of cognitive delegation to investigate the association between trust in augmented analytics and perceived decision quality among non-technical BI users. Data were collected from 250 business professionals across various organizational roles in Vietnam between January and March 2025 and analyzed using partial least squares structural equation modeling (PLS-SEM). Findings indicate that augmented analytics capabilities are positively associated with perceived ease of use, usefulness, and trust in BI systems. Trust and usefulness are jointly associated with BI adoption intention and perc

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications

arXiv:2602.18443v2 Announce Type: replace-cross Abstract: Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates eleven large language models generating six-word subject lines for German counselling emails through hierarchical assessment - first categorising outputs, then ranking within categories to enable manageable evaluation. Nine assessors (counselling professionals and AI systems) enable analysis via Krippendorff's $\alpha$, Spearman's $\rho$, Pearson's $r$ and Kendall's $\tau$. Results reveal performance trade-offs between proprietary services and privacy-preserving open-source alternatives, with German fine-tuning consistently improving performance. The study addresses critical ethical considerations for mental health AI deployment including privacy, bias and accountability.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

How College Students Use AI to Navigate Course Readings: Evidence from an Eight-Week Study

arXiv:2602.09907v3 Announce Type: replace-cross Abstract: College students increasingly use AI chatbots to support academic reading, yet we lack granular understanding of how these interactions shape their reading experience and cognitive engagement. We conducted an eight-week longitudinal study with 15 undergraduates who used AI to support assigned readings in a course. We collected 838 prompts across 239 reading sessions and developed a coding schema categorizing prompts into four cognitive themes: Decoding, Comprehension, Reasoning, and Metacognition. Comprehension prompts dominated (59.6%), with Reasoning (29.8%), Metacognition (8.5%), and Decoding (2.1%) less frequent. Most sessions (72%) contained exactly three prompts, the required minimum of the reading assignment. Within sessions, students showed natural cognitive progression from comprehension toward reasoning, but this progression was truncated. Across eight weeks, students' engagement patterns remained stable, with substant

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

arXiv:2507.21134v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their domain-specific safety and compliance becomes critical. While prior work has largely focused on improving LLM performance in these domains, it has often neglected the evaluation of domain-specific safety risks. To bridge this gap, we first define domain-specific safety principles for LLMs based on the AMA Principles of Medical Ethics, the ABA Model Rules of Professional Conduct, and the CFA Institute Code of Ethics. Building on this foundation, we introduce Trident-Bench, a benchmark specifically targeting LLM safety in the legal, financial, and medical domains. We evaluated 19 general-purpose and domain-specialized models on Trident-Bench and show that it effectively reveals key safety gaps -- strong generalist models (e.g., GPT, Gemini) can meet basic expectations, whereas domain-sp

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Robustness and Cybersecurity in the EU Artificial Intelligence Act

arXiv:2502.16184v3 Announce Type: replace-cross Abstract: The EU Artificial Intelligence Act (AIA) establishes different legal principles for different types of AI systems. While prior work has sought to clarify some of these principles, little attention has been paid to robustness and cybersecurity. This paper aims to fill this gap. We identify legal challenges and shortcomings in provisions related to robustness and cybersecurity for high-risk AI systems(Art. 15 AIA) and general-purpose AI models (Art. 55 AIA). We show that robustness and cybersecurity demand resilience against performance disruptions. Furthermore, we assess potential challenges in implementing these provisions in light of recent advancements in the machine learning (ML) literature. Our analysis informs efforts to develop harmonized standards, guidelines by the European Commission, as well as benchmarks and measurement methodologies under Art. 15(2) AIA. With this, we seek to bridge the gap between legal terminology

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes

arXiv:2412.20541v2 Announce Type: replace-cross Abstract: Memes act as cryptic tools for sharing sensitive ideas, often requiring contextual knowledge to interpret them correctly. It makes multimodal meme moderation difficult, as existing work either lacks high-quality datasets for nuanced hate categories or relies on low-quality social media visuals. Here, we curate two novel multimodal hate speech datasets comprising English memes - MHS and MHS-Con, which capture fine-grained hateful abstractions in regular and confounding scenarios, respectively. We benchmark these datasets against several competing baselines. Furthermore, we introduce SAFE-MEME (Structured reAsoning FramEwork) with its two variants: a novel multimodal Chain-of-Thought based framework employing Q&A-style reasoning (SAFE-MEME-QA) and a hierarchical categorization (SAFE-MEME-H) to enable robust hate speech detection in memes. SAFE-MEME-QA outperforms the strongest open-source baseline model, showing an improvement of

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Fairness Interventions in Classification: A Study on AI Explainability

arXiv:2407.14766v4 Announce Type: replace-cross Abstract: This paper presents a philosophical and experimental study of fairness interventions in AI classification, centered on the explainability and transparency of corrective methods, and on the opposition between two fairness criteria, namely Demographic Parity and Equalized Odds. Our main argument is that even as a gap in Demographic Parity is used to diagnose inequality between groups, Equalized Odds constitutes a more reliable fairness criterion to guide bias correction in classification. To establish this, we present FairDream, a fairness package intended for lay users, whose mechanism increases the model's weights of errors on disadvantaged groups. To justify FairDream's results, we analyze its reweighting algorithm, and we present the results of a benchmark experiment in which we compare FairDream with a distinct in-processing correction method that enforces Demographic Parity more drastically, the GridSearch method. We then pr

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

arXiv:2605.02050v2 Announce Type: replace Abstract: This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established practices from disciplines with established RCT traditions, including software engineering, economics, clinical and health sciences, and psychology, we synthesize five principles drawn from established validity frameworks and open-science standards on transparency, repeatability, and verification, which together serve as the conceptual foundation for 33 actionable guidelines adapted for AI evaluation RCT contexts, expressed as requirements with rationales, implementation instructions, and evidence bases. We position the principles and guidelines as serving three key roles for AI evaluation RCTs: a design tool for planning studies, an evaluation rubric for assessing existing work, and a blueprint for standard setting as the field converges on norms. AI evaluation research currently lacks common standard

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments

arXiv:2604.02458v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-world experiments are costly or infeasible. The treatment-effect estimates are often evaluated using statistical realism, the degree to which simulated responses reproduce properties of observed human responses, although whether realism predicts treatment-effect accuracy remains unknown. Here we test this proxy relationship by jointly measuring statistical realism and treatment-effect accuracy on the same simulated responses in a cross-national experiment with 59,508 participants from 62 countries using three LLMs. The correlation between statistical realism and treatment-effect accuracy is weak, and optimizing for statistical realism can even worsen treatment-effect accuracy when selecting models, prompts, and target populations. The pattern replicates in two additional cross-national experiments spa

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Coherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure

arXiv:2603.28371v2 Announce Type: replace Abstract: When an agent can articulate why something works, we typically take this as evidence of genuine understanding. This presupposes that effective action and correct explanation covary, and that coherent explanation reliably signals both. I argue that this assumption fails for contemporary Large Language Models (LLMs). I introduce what I call the Bidirectional Coherence Paradox: competence and grounding not only dissociate but invert across epistemic conditions. In low-observability domains, LLMs often act successfully while misidentifying the mechanisms that produce their success. In high-observability domains, they frequently generate explanations that accurately track observable causal structure yet fail to translate those diagnoses into effective intervention. In both cases, explanatory coherence remains intact, obscuring the underlying dissociation. Drawing on experiments in compiler optimization and hyperparameter tuning, I develop

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Stability of AI Governance Systems: A Coupled Dynamics Model of Public Trust and Social Disruptions

arXiv:2603.20248v2 Announce Type: replace Abstract: AI systems are increasingly entrenched in public governance, yet scholarship lacks formal tools to determine when deviations of public trust in algorithmic institutions dissipate and when they grow into collapse. Stability refers here to asymptotic recovery from finite state perturbations under fixed structural parameters. We address this gap by developing a mathematical framework for institutional trust stability that couples a Friedkin-Johnsen opinion dynamics process with a Hawkes-inspired intensity process for AI controversies. Motivated by the Computers-Are-Social-Actors literature and recent studies of trust in large language models, this bidirectional coupling reveals that governance stability depends on the structural architecture of the information environment rather than absolute trust levels. We derive an exact spectral stability criterion delineating resilience from collapse, demonstrating how event self-excitation and mem

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Scalable and Personalized Oral Assessments Using Voice AI

arXiv:2603.18221v3 Announce Type: replace Abstract: Written work no longer certifies that a student understands it: a polished analysis now says little about who did the thinking. Oral examinations restore that evidentiary link, but they have never scaled, because conducting and grading them is expensive. We report on a system in which voice AI conducts a personalized oral exam and a council of three large language models (LLMs) grades the transcript, each model scoring independently and then revising after reading the others. Across two undergraduate cohorts at NYU Stern (36 students in Fall 2025, 37 in Spring 2026), a voice subscription covered all speaking time and grading stayed under one dollar per exam. The deployments yield five practical engineering lessons that should generalize wherever understanding must be tested under questioning, from job interviews to professional certification.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI Systems in Text-Based Online Counselling: Ethical Considerations Across Three Implementation Approaches

arXiv:2601.08878v2 Announce Type: replace Abstract: Text-based online counselling scales across geographical and stigma barriers, yet faces practitioner shortages, lacks non-verbal cues and suffers inconsistent quality assurance. Whilst artificial intelligence offers promising solutions, its use in mental health counselling raises distinct ethical challenges. This paper analyses three AI implementation approaches - autonomous counsellor bots, AI training simulators and counsellor-facing augmentation tools. Drawing on professional codes, regulatory frameworks and scholarly literature, we identify four ethical principles - privacy, fairness, autonomy and accountability - and demonstrate their distinct manifestations across implementation approaches. Textual constraints may enable AI integration whilst requiring attention to implementation-specific hazards. This conceptual paper sensitises developers, researchers and practitioners to navigate AI-enhanced counselling ethics whilst preservi

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Introducing AI to an Online Petition Platform Changed Outputs but not Outcomes

arXiv:2511.13949v4 Announce Type: replace Abstract: The rapid integration of AI writing tools into online platforms raises critical questions about their impact on content production and outcomes. We leverage a unique natural experiment on Change$.$org, a leading social advocacy platform, to causally investigate the effects of an in-platform ''write with AI'' tool. To understand the impact of the AI integration, we collected 1.5 million petitions and employed a difference-in-differences analysis. Our findings reveal that in-platform AI access significantly altered the lexical features of petitions and increased petition homogeneity, but did not improve petition outcomes. We confirmed the results in a separate analysis of repeat petition writers who wrote petitions before and after introduction of the AI tool. The results suggest that while AI writing tools can profoundly reshape online content, their practical utility for improving desired outcomes may be less beneficial than anticipat

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Framework for Developing University Policies on Generative AI Governance: A Cross-national Comparative Study

arXiv:2504.02636v3 Announce Type: replace Abstract: As generative AI (GAI) becomes increasingly embedded in higher education, universities worldwide are developing policies to govern its ethical, pedagogical, and institutional use. However, these policies vary across national and institutional contexts. We undertake a cross-nationalanalysis of GAI guidelines issued by leading universities in the United States, Japan, and China, identifying key policy orientations and proposing a structured framework to support policy development. Using an extended Technology Acceptance Model as an analytical lens, we examine five domains Perceived Usefulness and Perceived Ease of Use, Perceived Risk, Facilitating Conditions, Social Influence, and Self-Efficacy, and identify 20 key themes through thematic coding. Together, these findings inform the development of the University Policy Development Framework for Generative AI (UPDF-GAI). U.S. universities emphasize faculty autonomy, practical application,

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Designing Within the Lines: Practitioners' Perspectives and Visualisation Tool Evaluation in the Arabic Context

arXiv:2607.24571v1 Announce Type: cross Abstract: Design guidelines and best practices serve as references that support designers throughout the visualisation design process. While considerable effort has identified the elements that contribute to effective data visualisations, little attention has been paid to how language (scripts and reading direction), tool support, and cultural context also shape design decisions. As a result, assumptions of homogeneity persist, with visualisation practices predominantly benefiting users of English and left-to-right (LTR) scripts while overlooking the needs of over two billion Arabic script users. We investigate how Arabic-speaking visualisation practitioners design for right-to-left (RTL) scripts. We report on an analytical evaluation of seven popular GUI-based visualisation authoring tools using an Arabic dataset complemented by interviews with 11 Arabic-speaking practitioners across journalism, design, and data analysis. Our findings reveal tha

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health

arXiv:2607.24275v1 Announce Type: cross Abstract: Ethical governance of AI-driven systems is often expressed through high-level principles and static documentation, creating a gap between regulatory requirements and system-level verification. This challenge is particularly acute in digital phenotyping, where continuous behavioural data raises concerns around consent, privacy, and fairness. In this paper, we propose a computational ethical framework for AI-driven digital phenotyping system in which ethical requirements are formalised as deontic temporal logic constraints, alongside a conceptual ethical agent that oversees the system and ensures that any supervised system satisfies the specified constraints. Using a case study involving financial data and mental health, we model key ethical properties and verify them using the Z3 Satisfiability Modulo Theories (SMT) solver. Our evaluation shows that the framework is logically consistent and that violations of the specified ethical proper

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Mapping the Reddit Bot Ecosystem: Taxonomy and Evolution

arXiv:2607.23941v1 Announce Type: cross Abstract: Automated agents increasingly participate in online communities, yet their population structure and roles remain poorly understood. Using a dataset of 3,389 identified bots and their full activity histories, we construct a taxonomy of bot "species" on the news aggregation and social media platform Reddit based on temporal, community, linguistic, and semantic features. Clustering analysis reveals 18 distinct bot types spanning content-specialized, behavior-driven, and infrastructural roles such as moderation and utility support. In addition, temporal analysis shows that bot numbers and activity expanded rapidly before peaking around the COVID-19 period, then started declining even before Reddit's 2023 API policy changes. However, the overall diversity of bot species has remained remarkably stable. These findings suggest that online bot populations form evolving digital ecosystems.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Sustainable Remote Access Architecture for Digital Inclusion through the Reuse of Discredited TV-BOX Devices

arXiv:2607.23935v1 Announce Type: cross Abstract: Digital inclusion and the volume of electronic waste (e-waste) are major challenges for society, with direct impacts on the educational context. Educational institutions, especially those with limited resources, face difficulties in expanding and maintaining computer laboratories. This paper presents the implementation of a sustainable, low-cost Desktop Virtualization Infrastructure (VDI) developed through the reuse of discarded electronic equipment. The proposed solution employs a refurbished Linux server and repurposes confiscated TV Box devices (originally intended for disposal) into functional ARM-based thin clients by installing a compatible Linux distribution. This approach is aligned with the principles of the circular economy by promoting hardware reuse and seeks to reduce the costs associated with deploying computing environments for education. The paper describes the system architecture, the device repurposing process, and the

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

arXiv:2607.23927v1 Announce Type: cross Abstract: A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut. Feedback exposes two failures: in some models, internal and external judgments swap; in others, accuracy improves while confidence decouples from correctness, dissociations invisible to existing benchmarks. Across models, this pattern implicates active, not aggregate, parameter count. This suggests that as AI systems take on autonomous, multi-turn roles,

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It

arXiv:2607.23893v1 Announce Type: cross Abstract: Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the question one level down, in categories where the buyer picks a person. It issued 2,400 grounded API calls in one two-hour window on 24 July 2026: 120 buyer-intent prompts, four models (GPT-5.6 Sol, Gemini 3.6 Flash, Perplexity Sonar Pro, Grok 4.5), five iterations each, four European markets and five query languages. Every response was coded for whether it named an individual professional, by a rule cascade that never consults a roster and that drops detections resolving to a same-named American city (precision 96.9%, recall 61.7%, so every rate below is a lower bound). All inference corrects for clustering within prompt: intraclass correlation 0.258, effective n 407 against a nominal 2,400. Models named an individual in 25.8% of responses. Category dominates: real estate 35.4% and car dealership

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Visible to the Court: How AI Is (and Isn't) Litigated in U.S. Federal Court Opinions

arXiv:2607.23888v1 Announce Type: cross Abstract: In the United States, artificial intelligence (AI) is rapidly deployed amid limited federal regulation. With courts become a recurring forum in which AI-related practices are scrutinized, it is important to empirically understand the AI litigation landscape to date. We address this gap through a systematic review of 559 U.S. federal court opinions in which AI plays a role in the parties' contentions, taxonomizing (1) common topics of dispute, (2) the AI technologies implicated, and (3) the parties involved, including common plaintiff and defendant types. We identify seven recurring dispute areas, six categories of AI technologies at the center of litigation, and four types of common litigants, alongside legal doctrines used by the litigants. A comparison of this taxonomy to the AI Incident Database revealed substantial gaps in coverage, definitions, and prevalence between documented and litigated harms, suggesting courts capture only pa

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

GEMCo: A Validated, Ethically Releasable Proxy for Inaccessible Counselling Data

arXiv:2607.23621v1 Announce Type: cross Abstract: This paper presents GEMCo, a releasable, human-written proxy for inaccessible counselling data: 86 complete German e-mail counselling conversations (728 messages), expert-authored cases and counsellor sessions with trained role-players. It is validated against a held-out reference of 124 real counselling conversations. The proxy and the real conversations are measured against each other in counsellor strategies and client emotions. The gap is detectable but small. A generative validation supports the analysis. The validation method itself generalises to any domain where real data cannot be shared but a human-made proxy can. Privacy and ethics keep real counselling data closed. GEMCo carries none by design and can be released -- a first step toward language research in this domain.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences

arXiv:2607.23566v1 Announce Type: cross Abstract: Traditional human marking in upper-secondary STEM education creates a structural bottleneck that restricts the frequency of formative mock examinations. This quasi-experimental, mixed-methods longitudinal study (N = 142) investigates the efficacy of deploying a fully automated, handwritten assessment marking platform to remove this bottleneck. Students preparing for STEM A-levels (Mathematics, Further Mathematics, Biology, Chemistry, Physics) were divided into control (N = 70, human marking) and intervention (N = 72, automated marking) groups over a 16-week term. Results indicate that the intervention group achieved a statistically significant improvement in final mock A-level marks (p < .01), scoring on average 16.8% higher than the control group. Furthermore, assessment turnaround time dropped from 11.2 days to < 0.1 days, enabling a fourfold increase in practice volume. The automated system's capacity for granular marginalia, specifi

Source ↗
Showing 3051–3100 of 18402 signals
← Prev Page 62 of 369 Next →