EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Gaming Together on Discord: Teen Gamer's Cross-Platform Practices

arXiv:2608.25942v1 Announce Type: new Abstract: Discord is one of the most popular communication platforms among gamers. While prior research has highlighted its role in community building, relatively little attention has been paid to its original gaming context-how it shapes gameplay and social experiences. To address this gap, we conducted semi-structured interviews with 16 teenage Discord users. Through reflexive thematic analysis, we show how players leverage Discord to create more collaborative and socially enriched experiences that extend beyond the game itself. However, gaming together on Discord also resulted in social and security risks. We conceptualize gaming together on Discord as a cross-platform practice that extends gameplay beyond a game and supports players' social needs. Additionally, cross-platform practice also introduces the 'platform gap,' where fragmented governance between platforms exposed players to risks. To address this tension, we propose design implication

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces

arXiv:2608.25876v1 Announce Type: new Abstract: Generative design interfaces increasingly expose semantic controls that let users steer output with concepts such as "more elegant" or "more minimalist," typically encoded by a vision-language model (VLM). A practical question is whether state-of-the-art VLMs represent objects consistently in terms of the same concept. We audit 6 VLMs by ranking untextured 3D objects along Kansei adjective pairs, where Kansei describes affective impressions of product form, with each axis defined as the difference between the text representations of its two poles. Geometric pairs serve as positive controls, and pairs of unrelated adjectives establish an empirical null. Across 10 categories of ShapeNet database, affective axes converge above the null (mean pairwise rank correlation 0.36 vs. 0.14) but below the geometric ceiling (0.44). The agreement between models is partial and highly uneven: on the three axes shared by all categories, mean convergence ra

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences

arXiv:2608.25771v1 Announce Type: new Abstract: In serious illness, human surrogates often struggle to accurately predict patient preferences (68% accuracy), causing decision conflict. Personalized Patient Preference Predictor (P4) agents offer a potential solution, but prior prototypes treat values as static ratings, ignoring the contextual, situation-dependent nature of medical choices. Grounded in the 'logic of care', we present P4-DT (Dilemma Training), a P4 agent that constructs a patient decision policy by engaging users with varied medical dilemmas, eliciting individual preference reasoning through bi-directional training. In a study with 12 patient-surrogate dyads, P4-DT predicted patient treatment choices with 81.7% accuracy, significantly exceeding chance (OR = 5.61 [2.03, 15.51], p < .001) and outperforming both unassisted surrogates (55.0%; OR = 3.67 [1.59, 8.47], p = .002) and surrogates assisted by P4-DT (61.7%). Comparative prompt analyses showed that incorporating conte

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

AffectSim: A Controllable Interactive 3D Simulation Benchmark for Embodied Affective Perception

arXiv:2608.25664v1 Announce Type: new Abstract: Existing affective benchmarks largely consist of fixed recordings whose observation conditions are determined before inference, making it difficult to systematically study how embodied sensing influences affective perception. We introduce AffectSim, a controllable interactive 3D simulation benchmark for embodied affective perception. Rather than treating affective samples as fixed recordings, AffectSim instantiates emotion-expressive human motions as replayable 3D episodes in which distance, orientation, occlusion, scene geometry, and agent viewpoint can be systematically varied while preserving the underlying behavior and emotion label. AffectSim contains 27{,}647 episodes across five emotion categories and 57 scenes. Its factorized design separates affective behavior from observation conditions, supporting controlled re-observation of the same behavior as well as agent-controlled sensing in an executable 3D environment. To demonstrate t

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Are Concept Bottleneck Models Effective as Decision-Support Systems?

arXiv:2608.25581v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) are interpretable-by-design neural networks that detect human-understandable concepts from the input and use them to generate predictions. By allowing users to inspect the concepts underlying a prediction and explore how predictions change under alternative concept configurations, CBMs have emerged as one of the most prominent approaches to supporting human-AI collaboration. However, user studies investigating their actual effectiveness as decision-support systems remain limited. We present two large-scale user studies (N participants = 705, N observations = 6,959) evaluating how concept-based explanations and user interventions on the model's concepts affect the performance of the human-AI team in two distinct binary classification tasks. Our results show that CBMs, and particularly their interactive component, can improve human-AI team accuracy relative to both unaided human performance and performance w

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Maru: Information Architecture as a Shared Language for Generating Aligned and Persistent User Interfaces

arXiv:2608.25565v1 Announce Type: new Abstract: Generative user interfaces (GenUIs) promise on-demand components tailored to users' needs. As users iterate on information tasks, they construct personal structures over information they encounter---how items are grouped, what gets prioritized, and what terms mean in their context. Yet, current systems leave these structural decisions to the model at each generation, ignoring the structural logic users have established. Without a persistent representational structure shared between user and system, GenUIs have no basis to remain aligned with what users have established. We draw on Information Architecture (IA), a design practice for organizing and structuring information, as a shared language to bridge user-constructed structure and system generation. We present a framework identifying four IA elements---partition, hierarchy, order, and vocabulary---and characterize how each maps to concrete UI generation decisions. We instantiate this fr

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

The Well-Being Palette: An Action-Word Selection Tool Designed for Low-Burden Reflection on Workplace Well-Being

arXiv:2608.25527v1 Announce Type: new Abstract: Background: Workplace well-being interventions need formats that can be used repeatedly with minimal disruption to daily work. We developed the Well-Being Palette, a web-based action-word selection tool designed for brief, low-burden reflection on workplace well-being. Methods: In a three-month exploratory field study at a private-sector corporate research institute in Japan, 88 analyzed participants selected up to three well-being-related action words after reflecting on positive actions or experiences from each workday. We examined application usage, PERMA Profiler scores, selected-word patterns across departments, selected-word diversity using Shannon entropy, and exploratory associations with sharing workshops. Results: During the formal intervention period, the application captured 3,480 input records and 10,104 selected words, and all 72 available action words were selected. Overall PERMA scores increased from baseline to post-inter

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

ScentEcho: Exploring Adsorbent Materials for Accurate Odor Collection and Playback

arXiv:2608.25494v1 Announce Type: new Abstract: Delivering odors that feel realistic and recognizable remains a core challenge for olfactory interaction systems, particularly in applications that demand precise scent delivery. A key limitation lies in the difficulty of capturing, preserving, and playing back real-world scent sources in a reliable and scalable manner. This study explores the potential of adsorbent materials for supporting realistic scent playback. We present ScentEcho, a portable system that enables modular scent collection and release. Through user evaluations, we identify which adsorbent materials tend to perform better for specific odors, and observe that perceived intensity strongly influences similarity ratings. In addition, odor recognition follows a graded pattern, with users moving from broad category identification to more specific source recognition as similarity increases. These findings offer practical insights for designing olfactory interfaces that are bot

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

TailorCoPilot: Enabling Agentic Pattern Making with Version-Controlled State Tracking

arXiv:2608.25462v1 Announce Type: new Abstract: Experience-driven manufacturing, such as garment pattern making, faces a severe generational skills gap because its core expertise relies on undocumented tacit knowledge forged through day-to-day practice. To address this challenge, we present TailorCoPilot, an agentic pattern-making system built upon a specially designed version-control backend TailorTrace. TailorTrace models sewing patterns as structured, discrete states and records their transformations during the pattern-making process as explicit operation sequences defined upon the geometry primitives in the sewing pattern (panels, edges, vertices and stitches). Integrated into a conventional pattern-making GUI, TailorTrace enables seamless documentation of senior experts' tacit pattern-making knowledge without breaking their daily workflow. The documented knowledge further offers interactive, pedagogical scaffolding for novices, while providing a robust foundation to power TailorCo

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Q&A or Document-Based? The Effects of Interface Type on How Screen Reader Users Access Interconnected Documents

arXiv:2608.25382v1 Announce Type: new Abstract: Blind and low-vision (BLV) users are increasingly engaging with large language model (LLM) interfaces to access documents, but it is unclear how such systems support or hinder their ability to build interconnected knowledge. To examine this gap, we compared a Question-Answer Interface (QAI) that supports open-ended conversational inquiry, with a Document Interface (DI) based mostly on traditional structured text document navigation. We recruited 16 BLV screen reader users where they used both interfaces to explore two fictional worlds. Data from interaction logs, concept maps, decision-based tasks, and semi-structured interviews provide comparative insights into how interface design supports knowledge construction. Findings show that participants visited more distinct documents with the DI and formed larger and more correct mental models with the DI than with the QAI. They were also more able to apply knowledge they had gained. Simultaneo

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations

arXiv:2608.25340v1 Announce Type: new Abstract: Agentic AI assistants are increasingly used in everyday life. However, they may also be misused to support harmful manipulation in interpersonal relationships. This problem is role-sensitive. Requests from users who seek to manipulate others should be blocked. Users who seek protection from manipulation should instead receive supportive guidance. We study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents. In multi-turn settings, individually plausible actions may combine into a harmful workflow. We introduce a benchmark of 1,000 five-turn conversations. It covers both attacker-side and victim-side scenarios. It also includes direct and adversarially paraphrased variants. We further propose HRGuard. It includes an online pre-generation gate and a turn-level post-generation gate. The post-generation gate maintains a decayed cumulative risk state and interrupts emerging man

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

AVI-Personality: A Trait-Activated Multimodal Dataset for Personality and Competency Assessment in Asynchronous Video Interviews

arXiv:2608.25316v1 Announce Type: new Abstract: With the rapid development of AI-based personality and job-related competency assessment, Asynchronous Video Interviews (AVIs) are increasingly used in recruitment. However, existing multimodal personality datasets are often based on short, task-free social media videos and crowdsourced apparent personality labels, which limits their construct validity and relevance to structured interview assessment. To address these limitations, we introduce AVI-Personality, a trait-activated multimodal dataset for personality and job-related competency assessment from AVIs. The dataset contains 3,876 interview videos from 646 participants who completed a simulated management traineeship application. Participants answered two generic questions and four personality-targeted questions designed according to Trait Activation Theory. Our dataset provides both self and observer-reported HEXACO personality traits and job-related competency. We validate AVI-Per

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

"Am I Just That Dumb?": Applicability, Action and Verification in Consumer IoT Security Advice

arXiv:2608.25225v1 Announce Type: new Abstract: Public campaigns urge people to change default passwords on Internet of Things (IoT) devices and keep them updated, assuming users can independently determine whether the advice applies. We gave 28 participants in the Netherlands two pieces of government-issued advice reflecting guidance in several countries and asked them to try to apply each to three of six consumer devices selected from bestseller lists, not confirmed feature availability (168 sessions). The protocol asked for each action to be demonstrated rather than completed. Of 84 password sessions, 33 reached no password setting, 50 an account-level setting, and one a device-level setting. Of 84 update sessions, 27 reached no update, 19 a companion-app update, and 38 a verified firmware update. No product had a manufacturer-set credential shared across units as described by the advice; the single device-level credential was unique to its unit. We contribute an account of what gen

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

The Systems Paper is Dead. Long Live the Systems Paper

arXiv:2608.25219v1 Announce Type: new Abstract: The way we structure, conduct, and write up interactive systems research in UIST papers rests on assumptions about constraints that may no longer hold today. What should an impactful UIST paper look like when building working systems is no longer hard? I argue that it should look different, and that we should ask more of our papers once implementation stops being a bottleneck.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Teaching Geometric Proof with Tech: Pitfalls and Possibilities

arXiv:2608.25117v1 Announce Type: new Abstract: Geometric proof is a foundational yet challenging topic in mathematics, requiring students to integrate visual, logical, and notational skills. While technology has enhanced learning in other mathematical domains, its impact on geometric proof remains limited. To investigate this gap, we interviewed 18 geometry teachers to establish the technical requirements of educational proof tools. These requirements inform our review of 33 commercial and research tools. Our findings reveal a critical mismatch: while teachers value certain digital tools for initial planning and exploration activities, they revert to pen-and-paper for formal proof because it supports diagram annotation and provides space for multiple approaches to proof-solving. Annotating the diagram is a key component of the proof-solving workflow that existing tools do not support. We propose four technical and human-centered design guidelines for educational proof tools to meet te

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

What Are We Measuring? Bonding, Trust, and the Evaluation of Human-Robot Relationships

arXiv:2608.24915v1 Announce Type: new Abstract: In human-robot interaction, relationship quality is often quantified using self-report measures, particularly related to "trust", such that a robot's trustworthiness comes to serve as an index of how close or "bonded" a human feels to it. I argue that this is a category error: trust and social bonding are distinct constructs, differing in their antecedents, their timescales, their bodily signatures, the human experience they produce, the robot responses they call for, and the ethical concerns they raise. I propose that we view them as independent dimensions, and describe the resulting two-dimensional space of possible relationship states under this view, with four configurations: avoidance, functional, dependence, and symbiosis. I then draw out some consequences for human-state-aware robotics: (1) social bonding is an explicit estimation target distinct from trust, (2) it should condition online adaptation (3) it reframes what a "failure"

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Blind Edits to Verified Repair: Building Trustworthy User-Side LLM Agents for Web Accessibility

arXiv:2608.24913v1 Announce Type: new Abstract: Assistive agents that adapt web pages on the user's side, at the moment of browsing, could reach the accessibility failures that site authors leave unfixed, and large language models make such agents newly plausible. We contribute three building blocks toward that goal. The first is a complete, privacy-preserving browser agent: a Chrome extension that extracts a page's style sheets, condenses them to fit a local model's context window, asks the model for additive CSS addressing 18 metrics from WCAG and the W3C cognitive accessibility guidance, and injects the result reversibly into the live page. The second is a dual-condition protocol that measures harm as carefully as benefit, applied to six small open-weight models (7B to 14B) on ten violation-rich and ten highly accessible live sites. The diagnosis is sobering but precise: unverified generation improved and regressed pages at similar rates (24 improvements against 20 regressions acros

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Analyzing and Correcting Benevolence Bias in Large Language Models

arXiv:2608.24912v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a "malicious persona" stre

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Plots to Words: Model-Aware Multimodal Explanations as a Foundation for Accessible, Non-Visual Interaction

arXiv:2608.24910v1 Announce Type: new Abstract: Multimodal large language models are increasingly used in interactive systems, yet ensuring consistent, trustworthy reasoning across heterogeneous modalities remains challenging. We present a context-aware, multi-agent framework that integrates textual queries, numerical data, visual representations, and model-derived signals for explainable time-series forecasting. A distinctive feature is that it turns predominantly visual forecasting outputs (e.g., trend plots) into structured, model-aware textual explanations. We argue that this makes the approach a natural foundation for non-visual, accessible interaction of particular relevance to blind and visually impaired users, for whom plot-centric interfaces are largely inaccessible. The framework supports three progressively richer pipelines (baseline, interpretable, explainable), enabling systematic comparison of unimodal, perception-driven, and model-aware responses. In an exploratory evalu

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

arXiv:2608.24909v1 Announce Type: new Abstract: Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are required to generate speech-synchronous gestures online, using only currently available response audio under strict latency constraints. As a result, prior methods are unsuitable for real-time interaction, as they either rely on future speech information or incur substantial inference delay. In this paper, we formulate online co-speech gesture generation for interactive digital humans and propose a real-time interactive framework that couples a streaming speech response module with an online gesture generation module. Specifically, the gesture generator is designed as a causal multimodal autoregressive model that predicts body motion from streaming response speech and motion history, enabling low-latency and speech-aligned

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

PA-CoT: Profile-Adaptive Chain-of-Thought for Personalized Nutritional Consulting

arXiv:2608.24907v1 Announce Type: new Abstract: In health and nutrition consulting, widely used prompting methods pass the user profile as an unstructured block without a dedicated analysis step, leaving personalization as a critical structural gap. We introduce PA-CoT (Profile-Adaptive Chain-of-Thought), a multi-stage prompting method that treats profile interpretation as an explicit, standalone reasoning step prior to response generation. To enable systematic evaluation, we introduce the QPA (Question--Profile--Answer) benchmark -- 200 nutritional consulting samples with structured user profiles scored on four criteria. In a comparative study against 11 comparison methods (CoT, Few-Shot, Role Prompting, DSPy, TextGrad, Self-Refine, and others, plus a Zero-Shot Baseline; 12 total including PA-CoT), PA-CoT achieves the best average score (4.21 on the G-Eval 1--5 scale) and leads on both Personalization (4.71 vs. 4.39) and Safety (4.68 vs. 4.52) with non-overlapping 95\% confidence inte

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

PARAssist: A Framework for Personalized and Adaptive Robotic Assistance from Ambiguous User Requests

arXiv:2608.24905v1 Announce Type: new Abstract: Service robots may encounter ambiguous user requests that require context-aware inference. Users may also have unique preferences with certain tasks when requesting robotic assistance. We introduce PARAssist (Personalized and Adaptive Robotic Assistance), a unique architecture for disambiguating requests in a personalized manner for service robots. PARAssist utilizes vision-language models to determine the physical and cognitive demands of a user's tasks, and passively learns user preferences for assistance by contrasting the demands of tasks the user performs independently with those they request from the robot. When an ambiguous request is received, task candidates are generated from the history of the user's actions, activities, locations, conversations, and requests, as well as the current user and environment state. Task candidates are then evaluated against the learned user preference model to suggest suitable assistance options. Ex

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Evidence-Grounded Mapping of Multimodal Human Sensing Psychological Transdiagnostic Dimensions

arXiv:2608.24903v1 Announce Type: new Abstract: Mobile and wearable sensing enables longitudinal observation of behavior, yet translating these signals into meaningful mental health constructs remains difficult. We introduce a clinician-in-the-loop benchmark for evaluating whether large language models (LLMs) can generate evidence-grounded Brief Hierarchical Taxonomy of Psychopathology (B-HiTOP) item profiles from passive sensing, ecological momentary assessment (EMA), and questionnaire evidence. Using the Generalization of Longitudinal Behavior Modeling (GLOBEM) dataset, we construct 14,592 participant-day instances and align multimodal evidence to 29 B-HiTOP items across five spectra. Since GLOBEM lacks B-HiTOP responses, we evaluate evidence compatibility (C) rather than diagnostic accuracy, separating substantive predictions from abstentions when evidence is insufficient for item-level scoring. Two-stage prediction improves C for EMA and questionnaire evidence, but reduces C under

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Beyond the Chatbot: Co-Learning and Co-Teaching through a Dual-Persona Generative-AI Assistant

arXiv:2608.24902v1 Announce Type: new Abstract: In this paper we present a generative AI application developed to support both teachers and students in secondary education. The system employs two Large Language Models-LLMs, Gemini and DeepSeek, and a Small Language Model-SLM, Gemma, integrated within a Retrieval Augmented Generation - RAG framework, creating a pedagogically grounded, Greek-language assistant capable of adapting its reasoning and communication style to the user role. Unlike conventional chatbots, the assistant introduces pedagogical persona switching, a dual-role mechanism that enables the same AI model to act as both a teaching companion and a learning guide. Utilizing a RAG paradigm tailored to the Greek educational domain, the architecture segments official textbooks into coherent units. Enriched with specific metadata, these units preserve curricular structure and instructional context, demonstrating how generative AI optimizes modern instructional design. The initi

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Stronger Alignment between Brain Activity and LLM Embeddings during Code Writing compared to Prose Writing

arXiv:2608.24900v1 Announce Type: new Abstract: Programming is a critical skill underlying modern software systems, yet the cognitive processes supporting code writing are only beginning to be understood, limiting educational practices and developer tools. At the same time, Large Language Models (LLMs) are increasingly used to assist programming. These models themselves are not well understood and can exhibit undesirable behavior like introducing security vulnerabilities. Given evidence that some cognitive representations may be shared between LLMs and the brain, we seek to improve our understanding on both fronts by relating these two systems to one another. We used Voxelwise Encoding Models (VEMs) to relate LLM embeddings to brain activity measured with functional Magnetic Resonance Imaging (fMRI) during naturalistic writing tasks. Using participants' (n = 23) keystrokes as prompts, we extracted LLM embeddings to predict voxelwise Blood Oxygen Level Dependent (BOLD) signal, quantifyi

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

MCP-Driven Accessibility Tree Standardization for AI-Powered Screen Reader Agents

arXiv:2608.24898v1 Announce Type: new Abstract: Large language model (LLM) agents that interact with graphical user interfaces increasingly rely on either raw screenshots or platform-specific accessibility application programming interfaces (APIs) to perceive interface state. Both approaches have limitations for assistive applications: screenshot-based perception lacks the semantic roles and relationships required by screen readers, while platform-specific APIs such as Windows UI Automation, macOS Accessibility, Android AccessibilityService, and web ARIA require separate integrations for each platform. This paper proposes an architecture that uses the Model Context Protocol (MCP) as a unified transport and schema layer between heterogeneous accessibility frameworks and LLM-based assistive agents. An MCP accessibility server exposes ARIA-aligned roles, labels, states, and focusable-element hierarchies through a platform-independent representation, enabling consistent interaction across

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Agentic World Analysis (AWA) - an alternative way to explore systems and support decision making

arXiv:2608.24896v1 Announce Type: new Abstract: To address increasingly pressing sustainability challenges, various approaches have been developed to foresee possible futures, identify failure modes, detect vulnerabilities, and test potential mitigations. However, environmental systems are highly complex. Especially when coupled with human processes, the scale of uncertainties becomes intractable. To address this challenge, we propose a new approach - Agentic World Analysis (AWA)- combining the strengths of simulation modelling and expert elicitation. The concept of AWA is defined by three properties: 1) AWA uses an agentic AI system to mimic an expert panel that studies the world; 2) AWA projects futures iteratively through analysing scenario trees and learning from this analysis to improve decisions; 3) AWA is auditable. Based on these requirements, we implemented the World Engine by Generative Agents (WEGA) as a possible application of the AWA approach and demonstrated its functiona

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

A Low-Code Approach for the Automatic Personalization of Conversational Agents

arXiv:2605.02384v3 Announce Type: replace-cross Abstract: The rise of Large Language Models (LLMs) has increased the demand for Conversational Agents (CAs) capable of understanding human conversations as part of web applications. While traditional CAs consist of deterministic states, LLMs enhance their capabilities to handle open conversations, handling arbitrary requests. Numerous tools exist that allow non-technical users to create such CAs. Yet, the creation of personalized CAs able to adapt to the profile of end-users to offer an optimal user experience remains in the hands of experienced developers implementing ad-hoc personalizations. In this work, we propose a pipeline that follows a low-code/no-code approach to facilitate the modeling and generation of personalized CAs. A pilot user study was performed to get preliminary results on perceived usability and usefulness and the full pipeline has been implemented on top of an open-source low-code platform.

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Aligning Human-AI-Interaction Trust for Mental Health Support: Survey and Position for Multi-Stakeholders

arXiv:2604.20166v2 Announce Type: replace-cross Abstract: Building trustworthy AI systems for mental health support is a shared priority across stakeholders from multiple disciplines. However, "trustworthy" remains loosely defined and inconsistently operationalized. AI research often focuses on technical criteria (e.g., robustness, explainability, and safety), while therapeutic practitioners emphasize therapeutic fidelity (e.g., appropriateness, empathy, and long-term user outcomes). To bridge the fragmented landscape, we propose a three-layer trust framework, covering human-oriented, AI-oriented, and interaction-oriented trust, integrating the viewpoints of key stakeholders (e.g., practitioners, researchers, regulators). Using this framework, we systematically review existing AI-driven research in mental health domain and examine evaluation practices for ``trustworthy'' ranging from automatic metrics to clinically validated approaches. We highlight critical gaps between what NLP curre

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos

arXiv:2603.25645v2 Announce Type: replace-cross Abstract: Early screening via colonoscopy is critical for colon cancer prevention, yet developing robust AI systems for this domain is hindered by the lack of densely annotated, long-sequence video datasets. Existing datasets predominantly focus on single-class polyp detection and lack the rich spatial, temporal, and linguistic annotations required to evaluate modern Multimodal Large Language Models (MLLMs). To address this critical gap, we introduce Colon-Bench, generated via a novel multi-stage agentic workflow. Our pipeline seamlessly integrates temporal proposals, bounding-box tracking, AI-driven visual confirmation, and human-in-the-loop review to scalably annotate full-procedure videos. The resulting verified benchmark is unprecedented in scope, encompassing 528 videos, 14 distinct lesion categories (including polyps, ulcers, and bleeding), over 300,000 bounding boxes, 213,000 segmentation masks, and 133,000 words of clinical descri

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

arXiv:2601.16529v3 Announce Type: replace-cross Abstract: Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based guidelines. We developed SycoEval-EM, a multi-agent simulation framework to evaluate LLM robustness to adversarial patient persuasion in emergency medicine. Across 19 contemporary LLMs and 1,425 simulated clinical encounters spanning three Choosing Wisely scenarios, acquiescence rates ranged from 0% to 100%, revealing a bimodal distribution. Seven models maintained near-perfect guideline adherence, while six acquiesced in the majority of encounters. Vulnerability varied substantially across clinical scenarios. Acquiescence was highest for CT imaging requests, intermediate for antibiotic prescriptions for sinusitis, and lowest for opioid prescriptions for acute back pain. Model scale, recency, and performance on static medical benchmarks did not consistently predict robustness. All five

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Towards a Bathroom-Centered Human-Building Digital Twin Framework for Indoor Safety Analysis

arXiv:2606.23292v2 Announce Type: replace Abstract: Bathroom use is a critical safety challenge for older adults because wet surfaces, constrained layouts, limited support, and frequent posture transitions are concentrated within a small domestic space. These conditions create risks that cannot be adequately understood by considering either the bathroom environment or human motion in isolation. Existing bathroom safety studies mainly identify hazards, accessibility problems, or design modifications, whereas human-centered sensing studies often focus on activity recognition or fall detection without sufficient semantic understanding of the surrounding environment. This separation limits the interpretation of how older adults interact with fixtures, support surfaces, wet areas, and spatial constraints during daily bathroom activities. To address this gap, this study proposes a bathroom-centered human-building digital twin framework for interaction-aware indoor safety analysis with a spec

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Seeing the Reasoning: How LLM Rationales Influence User Trust and Decision-Making in Factual Verification Tasks

arXiv:2603.07306v2 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly show reasoning rationales alongside their answers, turning "reasoning" into a user-interface element. While step-by-step rationales are typically associated with model performance, how they influence users' trust and decision-making in factual verification tasks remains unclear. We ran an online study (N=68) manipulating three properties of LLM reasoning rationales: presentation format (instant vs. delayed vs. on-demand), correctness (correct vs. incorrect), and certainty framing (none vs. certain vs. uncertain). We found that correct rationales and certainty cues increased trust, decision confidence, and AI advice adoption, whereas uncertainty cues reduced them. Presentation format did not have a significant effect, suggesting users were less sensitive to how reasoning was revealed than to its reliability. Participants indicated they use rationales to primarily audit outputs and calibrate tru

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Tinker Tales: A Tangible Dialogue System for Child-AI Co-Creative Storytelling

arXiv:2602.04109v2 Announce Type: replace Abstract: Conversational AI agents are increasingly explored as creative partners, yet how conversation design shapes child-AI dialogue in co-creative settings remains underexplored. We present Tinker Tales, a tangible dialogue system for child-AI collaborative storytelling, in which educational frameworks (narrative development and social-emotional learning) are instantiated as conversation design, shaping how the agent engages children across four narrative stages. The system combines a physical storytelling board, NFC-embedded toys, and a mobile app mediating multimodal interaction through tangible manipulation and voice-based dialogue. We conducted a home-based user study with 10 children (ages 6-8) across two conversation design conditions varying in how the agent structured elaboration, with and without educational scaffolding. Our findings show that prompt framing shapes the form and consistency of children's narrative contributions, str

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Virtual Reality Alters Perceived Functional Body Size

arXiv:2510.00824v2 Announce Type: replace Abstract: Virtual reality (VR) introduces sensory perturbations that may impact perception and action. The current study was designed to investigate how immersive VR presented through a head-mounted display (HMD) affects perceived functional body size using a passable aperture paradigm. Participants (n=60) performed an action task (sidle through apertures) and a perception task (adjust aperture width until passable without contact) in both physical, unmediated reality (UR) and VR. Results revealed significantly higher action and perceptual thresholds in VR compared to UR. Affordance ratios (perceptual threshold over action threshold) were also higher in VR, indicating that the increase in perceptual thresholds in VR was driven partly by sensorimotor uncertainty, as reflected in the increase in the action thresholds, and partly by perceptual distortions imposed by VR. This perceptual overestimation in VR also persisted as an aftereffect in UR fo

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

The Pin of Shame: Examining Content Creators' Adoption of Pinning Inappropriate Comments as a Moderation Strategy

arXiv:2505.14844v2 Announce Type: replace Abstract: Many social media platforms allow content creators to pin user comments in response to their content. Once pinned, a comment remains fixed at the top of the comments section, regardless of subsequent activity or the selected sorting order. The "Pin of Shame" refers to an innovative re-purposing of this feature, where creators intentionally pin norm-violating comments to spotlight them and prompt shaming responses from their audiences. This study explores how creators adopt this emerging moderation tactic, examining their motivations, its outcomes, and how it compares-procedurally and in effect-to other content moderation strategies. Through interviews with 20 content creators who had pinned negative comments on their posts, we find that the Pin of Shame is used to punish and educate inappropriate commenters, elicit emotional accountability, provoke audience negotiation of community norms, and support creators' impression management go

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Reasonable Motion: A General ASP Foundation for Environment Constrained Movement Trajectory Computation

arXiv:2606.25626v1 Announce Type: cross Abstract: We present a general answer set programming based hybrid quantitative-qualitative method for computing constrained branching trajectory modes for moving objects in real-world settings. The method performs constrained traversal of an environment graph, enumerating geometrically admissible motion behaviours as stable models, each constituting a distinct trajectory mode characterised by both domain-dependent and independent factors such as derived event sequence, map topology, and domain norms. The hybrid trajectory computation method is generally applicable across motion characteristics typically encountered in diverse dynamic domains with moving objects, e.g., autonomous driving. We demonstrate applicability and highlight how computed trajectories are traceable to their underlying stable model, thereby affording verifiable interpretability that purely learned approaches cannot provide. We also perform an empirical evaluation with Argover

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

AI Coaching for Accelerating Human Skill Development with Reinforcement Learning

arXiv:2606.25337v1 Announce Type: cross Abstract: AI copilots can substantially boost human performance through shared control, but excessive assistance can induce over-reliance and skill atrophy. This paper studies how an embodied AI agent can act as a coach that accelerates human motor-skill development. We argue that effective coaching requires strategic scaffolding and stepping back that are aligned with the learner's capability, allowing productive failures that drive learning. We formalize the interactive AI coaching process as a non-cooperative dynamic game in which the learner optimizes task performance while the coach targets the learner's independent competence. Building on this formalism, we develop a reinforcement learning framework combining adaptive shared control with probabilistic models of the coach's causal influence on skill evolution, enabling tractable training of coaching policies. A comprehensive user study (N=33) on first-person-view drone racing shows significa

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

EveLoad: Cognitive Workload Recognition from Event-Based Eye Movements

arXiv:2606.25177v1 Announce Type: cross Abstract: Cognitive workload monitoring is important for adaptive rehabilitation and assistive interfaces, where task difficulty, pacing, and feedback should be adjusted according to the user's cognitive state to avoid overload and under-challenge. Emerging extended reality and robot-assisted rehabilitation environments provide controllable training tasks, but they require unobtrusive sensing methods that can capture rapid ocular dynamics during interaction. Existing eye-movement-based cognitive workload recognition methods mainly rely on frame-based eye trackers, which often suffer from limited temporal resolution and degraded robustness under rapid eye movements. In contrast, event cameras provide microsecond-level temporal resolution, high dynamic range and low latency, making them suitable for capturing fine-grained ocular dynamics. Many previous studies rely on free-viewing or similar paradigms, where gaze locations can vary across tasks. As

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

fARfetch: Enabling Collocated AR-HRC in Large Visually Diverse Environments with VLM-Driven AR Content Adaptation

arXiv:2606.25162v1 Announce Type: cross Abstract: Augmented Reality (AR) can improve collocated human-robot collaboration by making robot state and intent visible and enabling intuitive control, yet large, visually diverse environments like the outdoors challenge both interaction and content legibility, especially at long distances and beyond visual line of sight. We present fARfetch, an AR-HRC system that integrates (i) shared semantic environment mapping across an AR headset and robot that visualizes detected landmarks in AR to support landmark-grounded go-to commands, (ii) a context-aware world-in-miniature representation of the shared environment for fine-grained path authoring, and (iii) vision-language-model driven AR view management that jointly adapts virtual content color, size, and orientation to maintain legibility in large visually diverse environments. We implement fARfetch with a Meta Quest 3 headset and Unitree Go2 quadruped robot, and conduct a within-subjects user stud

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

BCoughBench: Benchmarking Respiratory Acoustic Foundation Models Under Body-Coupled Wearable Sensor Conditions

arXiv:2606.25116v1 Announce Type: cross Abstract: Respiratory acoustic foundation models (FMs) are benchmarked exclusively on smartphone recordings, yet clinical deployment increasingly targets body-coupled (BC) wearables whose sensors attenuate high-frequency content through tissue and bone, leaving FM reliability uncharacterised. We introduce BCoughBench, evaluating five FMs (OPERA-CT/CE/GT, HeAR, M2D+Resp) on nine classification tasks (AUROC, sensitivity at 95% specificity, Expected Calibration Error) and three age regression tasks (MAE vs. a mean-predictor baseline) across five EBEN-simulated BC sensor conditions on five labeled cough datasets. Mean AUROC declines from 0.785 (smartphone) to 0.689-0.723, degrading most under temple vibration pickup ($\Delta$ = -0.096) and least under the soft in-ear ($\Delta$ = -0.062). No FM meets the clinical sensitivity threshold (Se@Sp95 $\geq$ 0.20) on most disease tasks under any BC sensor. Sex classification on the CIDRZ cohort collapses (AUR

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing

arXiv:2606.25108v1 Announce Type: cross Abstract: Autonomous AI systems are transitioning from advisory to autonomous roles for medication prescriptions. Recent United States bill H.R. 238 and Utah's prescription-renewal pilot both authorize AI to prescribe medications in an agentic capacity. While some regulatory guidelines suggest aggregate model performance metrics for clearance, they do not require i) calibrated per-prediction confidence for action-gated thresholds, ii) differentiated communication of uncertainty arising from model ignorance (epistemic) versus genuine clinical ambiguity (aleatoric), and iii) inferential transparency at the moment of decision that allows for liability allocation. Here, we present a regulatory and technical argument (tested with a survey of 136 U.S. prescribing clinicians) positioning these as minimum architectural requirements for safe autonomous prescribing. Our results suggest prescribing clinicians i) would not permit autonomous prescribing witho

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Explainable Control Framework (XCF) based on Fuzzy Model-Agnostic Explanation and LLM Agent-Supported Interface

arXiv:2606.25941v1 Announce Type: new Abstract: Increasing demand for precise and reliable control in complex scenarios has led to the development of increasingly sophisticated controllers, including data-driven approaches employing closed box models and mathematically rigorous yet complex designs. This complexity highlights the needs for explainable control that can provide human-understandable insights into controller behavior. In this paper, an explainable control framework (XCF) along with supporting algorithms and user interface are proposed to explain how controllers determine their control actions and their underlying working mechanism. The novel contributions of this work are threefold: First, the XCF is designed to provide model-agnostic explanations for controllers in closed-loop systems and can optionally refine local explanations by system response dynamics. Second, a novel explanation method, hierarchical fuzzy model-agnostic explanation for control systems (HFMAE-C), is p

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Designing Trustworthy LLM-based Wellbeing Recommendation through Controllable Interaction

arXiv:2606.25809v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate personalized guidance in wellbeing contexts such as physical activity, stress management, and mental health support, enabling fluent and context-aware interaction but relying on largely implicit mechanisms that shape how recommendations are expressed and adapted. We argue that this reliance on implicit adaptation through prompting and alignment limits control over guidance, responsibility framing, and user influence, which is particularly problematic in wellbeing settings where recommendations affect users' actions and long-term outcomes. We propose a system-level perspective in which conversational behavior is structured through explicit interaction constraints, including guidance strategies, explanation styles, degrees of directness, and mechanisms for user control. Building on prior work on tangible recommendations, we show how these constraints address key challenges in we

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Dissociable Spatial and Temporal Effects of Interaction Latency in Virtual Reality

arXiv:2606.25681v1 Announce Type: new Abstract: Motion-to-photon latency is inherent in immersive virtual reality (VR) systems and can arise from multiple sensorimotor loops, including view-contingent latency between head movement and display update and interaction latency between hand movement and the virtual effector. Although prior work shows that interaction latency can impair VR performance, it remains unclear whether common spatial, temporal, and efficiency measures reveal the same latency-related disruption. This study addressed this question by experimentally imposing delays between the physical and virtual hands during manual pointing in VR. Participants pointed to targets on a horizontal surface in VR and in the physical environment as an unmediated baseline. In VR, pointing was performed with a virtual hand avatar controlled by a motion capture pipeline, and additional delays (0-500 ms) were imposed between the participant's hand movement and the rendered movement of the vir

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

When LLM Rationales Become User-Facing: Effects on Trust Perception, Decision-Making, and Gaze Behaviors

arXiv:2606.25489v1 Announce Type: new Abstract: Large language models (LLMs) increasingly show step-by-step reasoning rationales alongside their answers, turning reasoning from an internal model capability into a user-facing interface feature. Yet it is unclear whether such rationales help users judge when trust is warranted or merely persuade through fluent reasoning. We address this gap through the lens of auditable trust calibration: user-facing rationales should help people inspect whether an answer is warranted by evidence. We test this framing in factual verification through two linked studies. Study 1, an online experiment (N=68), manipulated rationale presentation format (instant, delayed, on demand), rationale correctness (correct, incorrect), and certainty framing (none, certain, uncertain). Study 2, a controlled eye-tracking study (N=54), examined how no-, correct-, and incorrect-rationale conditions were associated with users' trust, decision-making, and eye-movement patter

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

The Digital Pirah\~a Condition: Ecological Mismatch and the Reconstruction of Recursive Cognition

arXiv:2606.25287v1 Announce Type: new Abstract: Contemporary digital and AI-mediated environments are reshaping the cognitive ecologies within which human reasoning develops. As everyday activity becomes embedded in datafied infrastructures, cognitive habits adapt to conditions of immediacy, fragmentation, externalisation, and algorithmic filtering. This paper introduces the Digital Pirah\~a Condition, a cultural ecological model explaining how these environments cultivate adaptive but shallow cognitive patterns, epistemic flattening, reduced recursive capacity, and heightened reliance on external scaffolds. While functional within digital systems, these adaptations create an ecological mismatch with the recursive, integrative reasoning required in academic and institutional activity systems. The paper argues that this mismatch is an ecological outcome rather than a psychological deficit, and that addressing it requires intentional cognitive niche construction within educational instit

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

Co-designing a Preliminary Repository of Augmented Reality Concepts for Real-Time Emotion Regulation

arXiv:2606.25271v1 Announce Type: new Abstract: Augmented Reality (AR) can be a positive therapeutic approach to support mental health and emotion regulation. Although AR techniques for therapeutic support exist, there is no user-centered, expert-informed understanding of how real-time AR designs can support people in emotional distress without disengaging them from their ongoing activities. This lack of reusable design resources hinders the adoption of AR for mental health support. This paper addresses this gap by introducing a co-designed collection of AR interventions describing how this technique can support real-time emotion regulation. The repository was created following a two-phase participatory design process. Phase 1 recruited 40 anxiety-prone individuals and used the Nominal Group Technique to list ideas on how AR affordances could support emotion regulation. Phase 2 recruited 10 mental health professionals to organize these ideas into thematic clusters and assess their clin

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

FUTO Swipe: Layout-Agnostic Neural Swipe Decoding

arXiv:2606.25247v1 Announce Type: new Abstract: Neural swipe decoders are typically tied to the keyboard they were trained on, requiring a new corpus and training run for each layout. In this report, we document our approach toward training models that can function on any contiguous mobile keyboard layout. At each point along the swipe, our encoder predicts whether the user is indicating a character and where on the keyboard that character lies. The keyboard layout is supplied at inference time and used to map the spatial and temporal prediction to a logit at each key, rather than being learned during training. Training neural models requires substantial data, but public swipe data is limited, particularly for non-QWERTY layouts. We release swipe.futo.org, the largest MIT-licensed swipe corpus we are aware of, containing over 1M donated swipes from more than 12k donor sessions. To generalize beyond the English QWERTY layout, we apply geometric augmentations to both the swipe trajectory

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.HC

ARTOO-DARTU: Studying AR-HRC With AR Obstruction Mitigation During a Warehouse Task

arXiv:2606.25202v1 Announce Type: new Abstract: Human-robot collaboration (HRC) often requires robot intentions and internal states to be conveyed to users for task efficiency and safety. Recently, augmented reality (AR) situated analytics provide such real-time robot feedback in HRC contexts. However, AR situated analytics can obstruct important real-world elements, posing safety and usability risks, especially when content is dynamically positioned relative to movements of mobile robots in a warehouse HRC scenario. In this paper, we introduce the Augmented Reality Technique Of Obstruction Deterrence while Aiding Robotic Teaming for Users (ARTOO-DARTU), an AR system tailored specifically for warehouse HRC that enables real-time robot situated analytics and control while preserving visibility of the real world through an obstruction detection and mitigation pipeline (ODM) that is uniquely suited for AR-HRC. To evaluate ARTOO-DARTU, we developed Pocket MonstARs, a controlled gamified ab

Source ↗
Showing 901–950 of 1631 signals
← Prev Page 19 of 33 Next →