EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring

arXiv:2606.20138v2 Announce Type: replace-cross Abstract: LLMs can personalize education, although current static-prompt tutoring systems struggle to adapt to diverse academic disciplines. We develop and test a system with subject-aware prompting, based on 14 pedagogical features (e.g., tutor scaffolding, student understanding) extracted from raw transcripts. We first train a prompt routing model in a simulation environment, and then deploy it for online adaptation with actual high-school students. The simulation benchmark shows the router outperforming two static baselines ($0.694$ vs. $0.647$ and $0.64$, $p<0.001$). A/B testing ($N=656$ conversations from 359 students) shows sim-to-real transfer where the model switches from analytical to scaffolding learning strategies. Our adaptive prompt selection mechanism improves instructional efficiency, maintains pedagogical quality and reduces interactions by around 3 turns ($p=0.007$). While a greedy router achieves a comparable exercise co

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Predicting Time Pressure of Powered Two-Wheeler Riders for Proactive Safety Interventions

arXiv:2601.03173v4 Announce Type: replace-cross Abstract: Time pressure critically influences risky maneuvers and crash proneness among powered two-wheeler riders, yet its prediction remains underexplored in intelligent transportation systems. To address this gap, we propose MotoTimePressure (MTPS), a deep learning model combining convolutional preprocessing, dual-stage temporal attention, and Squeeze-and-Excitation feature recalibration, achieving 91.53% accuracy and 98.93% ROC AUC, outperforming six baselines, with only 172K parameters, 0.66 MB model size, and 0.21 ms inference on CPU. To validate and benchmark MTPS, we present a dataset of 129,209 feature windows from 153 simulator sessions by 51 experienced male PTW riders under No, Low, and High Time Pressure conditions. Each sequence captures 63 features spanning vehicle kinematics, control inputs, behavioral violations, and environmental context. Our empirical analysis shows High Time Pressure induces 48% higher speeds, 36.4% gr

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Helping the Helper: LLM-Assisted Problem Articulation for Older Adults Seeking Technology Support

arXiv:2601.10018v2 Announce Type: replace Abstract: Older adults often struggle to articulate technology support needs due to unfamiliar technical terminology and age-related cognitive changes. We explore how large language models (LLMs) can facilitate this problem articulation process. Through a diary study (n = 27), we identified four communication barriers in older adults' queries: verbosity, incompleteness, over-specification, and under-specification. To mitigate these barriers, we developed an LLM pipeline that clarifies context and paraphrases unstructured queries. LLM-rephrased queries significantly improved automated solution accuracy (69% vs. 35%). Furthermore, younger adults (n = 48) acting as technology helpers understood LLM-rephrased queries better (93.7% vs. 65.8%) and reported greater ease in providing support. Older adults (n = 34) also found the resulting solutions highly actionable (94.7%). Finally, we contribute the first synthetic dataset of older adults' technology

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Automated Healthcare Thematic Analysis using Multi-Agent Large Language Model: Algorithm Development and Evaluation

arXiv:2512.16063v2 Announce Type: replace Abstract: Understanding patients experiences is essential for advancing patient-centered care. Qualitative thematic analysis is widely used to explore these experiences, however, the process remains labor-intensive, subjective, and difficult to scale. This study aimed to develop and evaluate Collaborative Theme Identification Agent (CoTI), a multi-agent large language model framework designed to support manual thematic analysis by rapidly generating supporting excerpts, initial codes, and themes. CoTI consists of three agents: Instructor, Thematizer, and CodebookGenerator. The Instructor refines instruction prompts, the Thematizer extracts supporting excerpts and generates initial codes for each transcript, and the CodebookGenerator groups similar codes across all transcripts into a codebook with themes. We evaluated CoTI primarily using 12 heart failure patient transcripts. CoTI-generated outputs were compared against the reference standard de

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Human learning is an understudied but promising lever for boosting human--AI synergy

arXiv:2512.13253v3 Announce Type: replace Abstract: Humans collaborating with artificial intelligence (AI) hold the promise of achieving superior outcomes compared to either acting alone (i.e., human--AI synergy). However, the conditions that facilitate such synergy when humans are advised by AI are not well understood. A recent meta-analysis showed that, on average, human--AI combinations do not outperform the better individual agent. We argue that this pessimistic conclusion arises from insufficient attention to human learning in experimental designs. To substantiate this claim, we re-analyzed all 74 studies included in the original meta-analysis and found that most previous research overlooked design features that foster human learning (e.g., outcome feedback to participants). Our re-analysis further revealed that studies providing outcome feedback show tentatively higher synergy than those without outcome feedback. Crucially, feedback paired with AI explanations was associated with

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

When Do Reactive Notebooks Fail to React?

arXiv:2511.21994v2 Announce Type: replace Abstract: Computational notebooks are convenient for programmers, but can easily become confusing and inconsistent due to the ability to incrementally edit a program that is running. Recent reactive notebook systems, such as Ipyflow, Marimo and Observable, strive to keep notebook state in sync with the current cell code by re-executing a minimal set of cells upon modification. However, each system defines reactivity a different way. Additionally, within any definition, we find simple notebook modifications that can break each system. Overall, these inconsistencies make it difficult for users to construct a mental model of their reactive notebook's implementation. This paper proposes Rex, a fine-grained test suite to discuss and assess reactivity capabilities within reactive notebook systems. We evaluate Rex on three existing reactive notebook systems and classify their failures with the aims of (i) helping programmers understand when reactivity

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

People readily follow personal advice from AI but it does not improve their well-being

arXiv:2511.15352v4 Announce Type: replace Abstract: People increasingly seek personal advice from large language models (LLMs), yet whether humans follow their advice, and its consequences for their well-being, remains unknown. In a longitudinal randomised controlled trial with a representative UK sample (N = 6,474), we found that up to 79% of participants who had a 20-minute discussion with one of three AI chatbots (GPT-4o, LLama-3.3-70B, Gemini 3 Pro) about health, careers or relationships subsequently reported following its advice. Advice-following remained above 65% even for high-stakes recommendations, suggesting that users only weakly calibrate their reliance on AI advice to potential consequences. Based on autograder evaluations of chat transcripts, LLM advice rarely violated safety best practice. However, when queried 2-3 weeks later, participants receiving personal advice from AI showed no sustained well-being benefits compared to a control group who discussed hobbies and inte

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching

arXiv:2608.26094v1 Announce Type: cross Abstract: Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, overlooking physiological dynamics such as muscle mechanics and often modeling actions as monolithic patterns. These limitations hinder fine-grained, biomechanically grounded feedback. We introduce MyoMechanix, a multimodal ecosystem for weight-loaded actions that aligns motion with muscle activity. Expert-annotated, it contains 7,500+ samples of 20 actions from 38 subjects, with synchronized multiview RGB video, 3D pose, sEMG, and additional physiological signals, forming the largest multimodal AQA benchmark to date. We further construct the Fitness Knowledge Graph (FKG), which organizes expert annotations into structured relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment. Building on these representations, we develop CUBIST (Composi

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Simultaneous Digital Communication and Deformation Sensing over a Single Stretchable Interconnect

arXiv:2608.25801v1 Announce Type: cross Abstract: Stretchable hybrid electronics integrate rigid solid-state electronics with stretchable materials and structures to achieve both high deformability and stable electronic performance. However, most existing systems treat stretchability only as a mechanical attribute without exploiting device deformation to encode its own mechanical state. This problem arises from adapting conventional rigid circuit architectures to stretchable substrates, affording a loss in compatibility with the sensors required for strain measurement. This study addresses this issue by proposing a communication-integrated deformation sensing architecture for stretchable hybrid devices. In the proposed approach, standard universal asynchronous receiver-transmitter digital signals transmitted between rigid nodes are amplitude-modulated by strain-induced resistance changes in stretchable liquid metal interconnects. By reading both amplitude changes and digital patterns,

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Development of a Voice-Controlled Tendon-Driven Bionic Hand

arXiv:2608.25222v1 Announce Type: cross Abstract: The impairment of the hands can seriously affect the abilities of every individual to perform the every-day activity, so the design of stable and controllable support devices is a significant field of study. This paper is about the design and implementation of an automated bionic hand which is dedicated to the coordinated finger movement through the simplified and efficient actuation mechanism. The method that the proposed system was designed on is the tendon-based method whereby the servo motors generate the movement of the fingers, with assistance of the angular control which is calibrated. An actuation is controlled by a microcontroller that will be programmed by use of an Arduino-based microcontroller to carry out programmed gestures that include open hand, fist, pinch and half flexion. It has an interface that is voice command enabled to make it easy to interact with a Bluetooth based sender receiver architecture which offers an op

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Longitudinal Robot Learning from Demonstration with Care Providers in a Home Environment

arXiv:2608.25196v1 Announce Type: cross Abstract: Learning from demonstration (LfD) methods enable non-expert end users to teach robots novel skills without explicit programming. However most evaluations of the usability of LfD with non-experts has been conducted in controlled laboratory environments with a robotics experimenter present. In this work we identify non-expert end users' key barriers when teaching robots via demonstration without live robotics expert feedback in a home environment. In our human subjects experiment we support the non-expert end users through two forms of demonstrator guidance developed in prior work: pre-training and adaptive feedback. Towards the ecological validity of the evaluation, we conduct this experimentation over multiple visits, with a population of care providers. Finally, we propose to open source the resulting LfD dataset of care providers teaching a robot assistive tasks over multiple visits to a home environment.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment

arXiv:2608.24920v1 Announce Type: cross Abstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition

arXiv:2608.24904v1 Announce Type: cross Abstract: Inertial sensors at multiple body locations can improve activity recognition, but requiring every sensor at inference increases the deployment burden. We study whether four synchronized IMUs available during training can improve a student that uses only the right-arm IMU during fitting and inference. A frozen four-IMU teacher provides logit and feature targets. Fixed-weight knowledge distillation applies each target with the same strength to every fitting sample, although the student may not benefit equally from them. We introduce dynamic influence weighting (DIW), which tests a one-step candidate update on separate fold-internal training participants. DIW then assigns separate sample-wise gates to the logit and feature losses. On WEAR, we evaluate 19 labels and 68,298 complete windows from 22 participants using subject-disjoint five-fold cross-validation. Pooled out-of-fold macro-F1 is 0.561820 for Supervised and 0.571623 for Fixed-wei

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

arXiv:2608.24901v1 Announce Type: cross Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding the Resonance direction raises the affective score only partially -- in Qwen by +0.29 (approximately 26% of the natural gap). A direct between-direction contrast confirms the shift is facet-specific in Qwen and Llama (not Gemma); we do no

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Producing to Validating: How AI Is Deskilling Freelancers

arXiv:2608.26089v1 Announce Type: new Abstract: Generative AI is promoted as a way to enhance knowledge work, yet its benefits and drawbacks fall unevenly across the workforce. Freelance and gig workers, who commonly lack the upskilling pathways available to traditional employees, face heightened risks to both skill development and job security as AI adoption advances. We review empirical evidence on AI's impact on knowledge-worker workflows and upskilling, then predict the primary and downstream effects of AI adoption among clients and workers in the freelance economy. We anchor this in two cases of the same shift, machine-translation post-editing and software development. We argue that freelancers are the leading edge of a change that also reaches salaried HCI practitioners, and we close with questions for the platforms and clients that mediate this work, and for HCI researchers.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Gaming Together on Discord: Teen Gamer's Cross-Platform Practices

arXiv:2608.25942v1 Announce Type: new Abstract: Discord is one of the most popular communication platforms among gamers. While prior research has highlighted its role in community building, relatively little attention has been paid to its original gaming context-how it shapes gameplay and social experiences. To address this gap, we conducted semi-structured interviews with 16 teenage Discord users. Through reflexive thematic analysis, we show how players leverage Discord to create more collaborative and socially enriched experiences that extend beyond the game itself. However, gaming together on Discord also resulted in social and security risks. We conceptualize gaming together on Discord as a cross-platform practice that extends gameplay beyond a game and supports players' social needs. Additionally, cross-platform practice also introduces the 'platform gap,' where fragmented governance between platforms exposed players to risks. To address this tension, we propose design implication

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces

arXiv:2608.25876v1 Announce Type: new Abstract: Generative design interfaces increasingly expose semantic controls that let users steer output with concepts such as "more elegant" or "more minimalist," typically encoded by a vision-language model (VLM). A practical question is whether state-of-the-art VLMs represent objects consistently in terms of the same concept. We audit 6 VLMs by ranking untextured 3D objects along Kansei adjective pairs, where Kansei describes affective impressions of product form, with each axis defined as the difference between the text representations of its two poles. Geometric pairs serve as positive controls, and pairs of unrelated adjectives establish an empirical null. Across 10 categories of ShapeNet database, affective axes converge above the null (mean pairwise rank correlation 0.36 vs. 0.14) but below the geometric ceiling (0.44). The agreement between models is partial and highly uneven: on the three axes shared by all categories, mean convergence ra

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences

arXiv:2608.25771v1 Announce Type: new Abstract: In serious illness, human surrogates often struggle to accurately predict patient preferences (68% accuracy), causing decision conflict. Personalized Patient Preference Predictor (P4) agents offer a potential solution, but prior prototypes treat values as static ratings, ignoring the contextual, situation-dependent nature of medical choices. Grounded in the 'logic of care', we present P4-DT (Dilemma Training), a P4 agent that constructs a patient decision policy by engaging users with varied medical dilemmas, eliciting individual preference reasoning through bi-directional training. In a study with 12 patient-surrogate dyads, P4-DT predicted patient treatment choices with 81.7% accuracy, significantly exceeding chance (OR = 5.61 [2.03, 15.51], p < .001) and outperforming both unassisted surrogates (55.0%; OR = 3.67 [1.59, 8.47], p = .002) and surrogates assisted by P4-DT (61.7%). Comparative prompt analyses showed that incorporating conte

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

AffectSim: A Controllable Interactive 3D Simulation Benchmark for Embodied Affective Perception

arXiv:2608.25664v1 Announce Type: new Abstract: Existing affective benchmarks largely consist of fixed recordings whose observation conditions are determined before inference, making it difficult to systematically study how embodied sensing influences affective perception. We introduce AffectSim, a controllable interactive 3D simulation benchmark for embodied affective perception. Rather than treating affective samples as fixed recordings, AffectSim instantiates emotion-expressive human motions as replayable 3D episodes in which distance, orientation, occlusion, scene geometry, and agent viewpoint can be systematically varied while preserving the underlying behavior and emotion label. AffectSim contains 27{,}647 episodes across five emotion categories and 57 scenes. Its factorized design separates affective behavior from observation conditions, supporting controlled re-observation of the same behavior as well as agent-controlled sensing in an executable 3D environment. To demonstrate t

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Are Concept Bottleneck Models Effective as Decision-Support Systems?

arXiv:2608.25581v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) are interpretable-by-design neural networks that detect human-understandable concepts from the input and use them to generate predictions. By allowing users to inspect the concepts underlying a prediction and explore how predictions change under alternative concept configurations, CBMs have emerged as one of the most prominent approaches to supporting human-AI collaboration. However, user studies investigating their actual effectiveness as decision-support systems remain limited. We present two large-scale user studies (N participants = 705, N observations = 6,959) evaluating how concept-based explanations and user interventions on the model's concepts affect the performance of the human-AI team in two distinct binary classification tasks. Our results show that CBMs, and particularly their interactive component, can improve human-AI team accuracy relative to both unaided human performance and performance w

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Maru: Information Architecture as a Shared Language for Generating Aligned and Persistent User Interfaces

arXiv:2608.25565v1 Announce Type: new Abstract: Generative user interfaces (GenUIs) promise on-demand components tailored to users' needs. As users iterate on information tasks, they construct personal structures over information they encounter---how items are grouped, what gets prioritized, and what terms mean in their context. Yet, current systems leave these structural decisions to the model at each generation, ignoring the structural logic users have established. Without a persistent representational structure shared between user and system, GenUIs have no basis to remain aligned with what users have established. We draw on Information Architecture (IA), a design practice for organizing and structuring information, as a shared language to bridge user-constructed structure and system generation. We present a framework identifying four IA elements---partition, hierarchy, order, and vocabulary---and characterize how each maps to concrete UI generation decisions. We instantiate this fr

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

The Well-Being Palette: An Action-Word Selection Tool Designed for Low-Burden Reflection on Workplace Well-Being

arXiv:2608.25527v1 Announce Type: new Abstract: Background: Workplace well-being interventions need formats that can be used repeatedly with minimal disruption to daily work. We developed the Well-Being Palette, a web-based action-word selection tool designed for brief, low-burden reflection on workplace well-being. Methods: In a three-month exploratory field study at a private-sector corporate research institute in Japan, 88 analyzed participants selected up to three well-being-related action words after reflecting on positive actions or experiences from each workday. We examined application usage, PERMA Profiler scores, selected-word patterns across departments, selected-word diversity using Shannon entropy, and exploratory associations with sharing workshops. Results: During the formal intervention period, the application captured 3,480 input records and 10,104 selected words, and all 72 available action words were selected. Overall PERMA scores increased from baseline to post-inter

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

ScentEcho: Exploring Adsorbent Materials for Accurate Odor Collection and Playback

arXiv:2608.25494v1 Announce Type: new Abstract: Delivering odors that feel realistic and recognizable remains a core challenge for olfactory interaction systems, particularly in applications that demand precise scent delivery. A key limitation lies in the difficulty of capturing, preserving, and playing back real-world scent sources in a reliable and scalable manner. This study explores the potential of adsorbent materials for supporting realistic scent playback. We present ScentEcho, a portable system that enables modular scent collection and release. Through user evaluations, we identify which adsorbent materials tend to perform better for specific odors, and observe that perceived intensity strongly influences similarity ratings. In addition, odor recognition follows a graded pattern, with users moving from broad category identification to more specific source recognition as similarity increases. These findings offer practical insights for designing olfactory interfaces that are bot

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

TailorCoPilot: Enabling Agentic Pattern Making with Version-Controlled State Tracking

arXiv:2608.25462v1 Announce Type: new Abstract: Experience-driven manufacturing, such as garment pattern making, faces a severe generational skills gap because its core expertise relies on undocumented tacit knowledge forged through day-to-day practice. To address this challenge, we present TailorCoPilot, an agentic pattern-making system built upon a specially designed version-control backend TailorTrace. TailorTrace models sewing patterns as structured, discrete states and records their transformations during the pattern-making process as explicit operation sequences defined upon the geometry primitives in the sewing pattern (panels, edges, vertices and stitches). Integrated into a conventional pattern-making GUI, TailorTrace enables seamless documentation of senior experts' tacit pattern-making knowledge without breaking their daily workflow. The documented knowledge further offers interactive, pedagogical scaffolding for novices, while providing a robust foundation to power TailorCo

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Q&A or Document-Based? The Effects of Interface Type on How Screen Reader Users Access Interconnected Documents

arXiv:2608.25382v1 Announce Type: new Abstract: Blind and low-vision (BLV) users are increasingly engaging with large language model (LLM) interfaces to access documents, but it is unclear how such systems support or hinder their ability to build interconnected knowledge. To examine this gap, we compared a Question-Answer Interface (QAI) that supports open-ended conversational inquiry, with a Document Interface (DI) based mostly on traditional structured text document navigation. We recruited 16 BLV screen reader users where they used both interfaces to explore two fictional worlds. Data from interaction logs, concept maps, decision-based tasks, and semi-structured interviews provide comparative insights into how interface design supports knowledge construction. Findings show that participants visited more distinct documents with the DI and formed larger and more correct mental models with the DI than with the QAI. They were also more able to apply knowledge they had gained. Simultaneo

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations

arXiv:2608.25340v1 Announce Type: new Abstract: Agentic AI assistants are increasingly used in everyday life. However, they may also be misused to support harmful manipulation in interpersonal relationships. This problem is role-sensitive. Requests from users who seek to manipulate others should be blocked. Users who seek protection from manipulation should instead receive supportive guidance. We study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents. In multi-turn settings, individually plausible actions may combine into a harmful workflow. We introduce a benchmark of 1,000 five-turn conversations. It covers both attacker-side and victim-side scenarios. It also includes direct and adversarially paraphrased variants. We further propose HRGuard. It includes an online pre-generation gate and a turn-level post-generation gate. The post-generation gate maintains a decayed cumulative risk state and interrupts emerging man

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

AVI-Personality: A Trait-Activated Multimodal Dataset for Personality and Competency Assessment in Asynchronous Video Interviews

arXiv:2608.25316v1 Announce Type: new Abstract: With the rapid development of AI-based personality and job-related competency assessment, Asynchronous Video Interviews (AVIs) are increasingly used in recruitment. However, existing multimodal personality datasets are often based on short, task-free social media videos and crowdsourced apparent personality labels, which limits their construct validity and relevance to structured interview assessment. To address these limitations, we introduce AVI-Personality, a trait-activated multimodal dataset for personality and job-related competency assessment from AVIs. The dataset contains 3,876 interview videos from 646 participants who completed a simulated management traineeship application. Participants answered two generic questions and four personality-targeted questions designed according to Trait Activation Theory. Our dataset provides both self and observer-reported HEXACO personality traits and job-related competency. We validate AVI-Per

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

"Am I Just That Dumb?": Applicability, Action and Verification in Consumer IoT Security Advice

arXiv:2608.25225v1 Announce Type: new Abstract: Public campaigns urge people to change default passwords on Internet of Things (IoT) devices and keep them updated, assuming users can independently determine whether the advice applies. We gave 28 participants in the Netherlands two pieces of government-issued advice reflecting guidance in several countries and asked them to try to apply each to three of six consumer devices selected from bestseller lists, not confirmed feature availability (168 sessions). The protocol asked for each action to be demonstrated rather than completed. Of 84 password sessions, 33 reached no password setting, 50 an account-level setting, and one a device-level setting. Of 84 update sessions, 27 reached no update, 19 a companion-app update, and 38 a verified firmware update. No product had a manufacturer-set credential shared across units as described by the advice; the single device-level credential was unique to its unit. We contribute an account of what gen

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

The Systems Paper is Dead. Long Live the Systems Paper

arXiv:2608.25219v1 Announce Type: new Abstract: The way we structure, conduct, and write up interactive systems research in UIST papers rests on assumptions about constraints that may no longer hold today. What should an impactful UIST paper look like when building working systems is no longer hard? I argue that it should look different, and that we should ask more of our papers once implementation stops being a bottleneck.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Teaching Geometric Proof with Tech: Pitfalls and Possibilities

arXiv:2608.25117v1 Announce Type: new Abstract: Geometric proof is a foundational yet challenging topic in mathematics, requiring students to integrate visual, logical, and notational skills. While technology has enhanced learning in other mathematical domains, its impact on geometric proof remains limited. To investigate this gap, we interviewed 18 geometry teachers to establish the technical requirements of educational proof tools. These requirements inform our review of 33 commercial and research tools. Our findings reveal a critical mismatch: while teachers value certain digital tools for initial planning and exploration activities, they revert to pen-and-paper for formal proof because it supports diagram annotation and provides space for multiple approaches to proof-solving. Annotating the diagram is a key component of the proof-solving workflow that existing tools do not support. We propose four technical and human-centered design guidelines for educational proof tools to meet te

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

What Are We Measuring? Bonding, Trust, and the Evaluation of Human-Robot Relationships

arXiv:2608.24915v1 Announce Type: new Abstract: In human-robot interaction, relationship quality is often quantified using self-report measures, particularly related to "trust", such that a robot's trustworthiness comes to serve as an index of how close or "bonded" a human feels to it. I argue that this is a category error: trust and social bonding are distinct constructs, differing in their antecedents, their timescales, their bodily signatures, the human experience they produce, the robot responses they call for, and the ethical concerns they raise. I propose that we view them as independent dimensions, and describe the resulting two-dimensional space of possible relationship states under this view, with four configurations: avoidance, functional, dependence, and symbiosis. I then draw out some consequences for human-state-aware robotics: (1) social bonding is an explicit estimation target distinct from trust, (2) it should condition online adaptation (3) it reframes what a "failure"

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Blind Edits to Verified Repair: Building Trustworthy User-Side LLM Agents for Web Accessibility

arXiv:2608.24913v1 Announce Type: new Abstract: Assistive agents that adapt web pages on the user's side, at the moment of browsing, could reach the accessibility failures that site authors leave unfixed, and large language models make such agents newly plausible. We contribute three building blocks toward that goal. The first is a complete, privacy-preserving browser agent: a Chrome extension that extracts a page's style sheets, condenses them to fit a local model's context window, asks the model for additive CSS addressing 18 metrics from WCAG and the W3C cognitive accessibility guidance, and injects the result reversibly into the live page. The second is a dual-condition protocol that measures harm as carefully as benefit, applied to six small open-weight models (7B to 14B) on ten violation-rich and ten highly accessible live sites. The diagnosis is sobering but precise: unverified generation improved and regressed pages at similar rates (24 improvements against 20 regressions acros

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Analyzing and Correcting Benevolence Bias in Large Language Models

arXiv:2608.24912v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a "malicious persona" stre

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Plots to Words: Model-Aware Multimodal Explanations as a Foundation for Accessible, Non-Visual Interaction

arXiv:2608.24910v1 Announce Type: new Abstract: Multimodal large language models are increasingly used in interactive systems, yet ensuring consistent, trustworthy reasoning across heterogeneous modalities remains challenging. We present a context-aware, multi-agent framework that integrates textual queries, numerical data, visual representations, and model-derived signals for explainable time-series forecasting. A distinctive feature is that it turns predominantly visual forecasting outputs (e.g., trend plots) into structured, model-aware textual explanations. We argue that this makes the approach a natural foundation for non-visual, accessible interaction of particular relevance to blind and visually impaired users, for whom plot-centric interfaces are largely inaccessible. The framework supports three progressively richer pipelines (baseline, interpretable, explainable), enabling systematic comparison of unimodal, perception-driven, and model-aware responses. In an exploratory evalu

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

arXiv:2608.24909v1 Announce Type: new Abstract: Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are required to generate speech-synchronous gestures online, using only currently available response audio under strict latency constraints. As a result, prior methods are unsuitable for real-time interaction, as they either rely on future speech information or incur substantial inference delay. In this paper, we formulate online co-speech gesture generation for interactive digital humans and propose a real-time interactive framework that couples a streaming speech response module with an online gesture generation module. Specifically, the gesture generator is designed as a causal multimodal autoregressive model that predicts body motion from streaming response speech and motion history, enabling low-latency and speech-aligned

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

PA-CoT: Profile-Adaptive Chain-of-Thought for Personalized Nutritional Consulting

arXiv:2608.24907v1 Announce Type: new Abstract: In health and nutrition consulting, widely used prompting methods pass the user profile as an unstructured block without a dedicated analysis step, leaving personalization as a critical structural gap. We introduce PA-CoT (Profile-Adaptive Chain-of-Thought), a multi-stage prompting method that treats profile interpretation as an explicit, standalone reasoning step prior to response generation. To enable systematic evaluation, we introduce the QPA (Question--Profile--Answer) benchmark -- 200 nutritional consulting samples with structured user profiles scored on four criteria. In a comparative study against 11 comparison methods (CoT, Few-Shot, Role Prompting, DSPy, TextGrad, Self-Refine, and others, plus a Zero-Shot Baseline; 12 total including PA-CoT), PA-CoT achieves the best average score (4.21 on the G-Eval 1--5 scale) and leads on both Personalization (4.71 vs. 4.39) and Safety (4.68 vs. 4.52) with non-overlapping 95\% confidence inte

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

PARAssist: A Framework for Personalized and Adaptive Robotic Assistance from Ambiguous User Requests

arXiv:2608.24905v1 Announce Type: new Abstract: Service robots may encounter ambiguous user requests that require context-aware inference. Users may also have unique preferences with certain tasks when requesting robotic assistance. We introduce PARAssist (Personalized and Adaptive Robotic Assistance), a unique architecture for disambiguating requests in a personalized manner for service robots. PARAssist utilizes vision-language models to determine the physical and cognitive demands of a user's tasks, and passively learns user preferences for assistance by contrasting the demands of tasks the user performs independently with those they request from the robot. When an ambiguous request is received, task candidates are generated from the history of the user's actions, activities, locations, conversations, and requests, as well as the current user and environment state. Task candidates are then evaluated against the learned user preference model to suggest suitable assistance options. Ex

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Evidence-Grounded Mapping of Multimodal Human Sensing Psychological Transdiagnostic Dimensions

arXiv:2608.24903v1 Announce Type: new Abstract: Mobile and wearable sensing enables longitudinal observation of behavior, yet translating these signals into meaningful mental health constructs remains difficult. We introduce a clinician-in-the-loop benchmark for evaluating whether large language models (LLMs) can generate evidence-grounded Brief Hierarchical Taxonomy of Psychopathology (B-HiTOP) item profiles from passive sensing, ecological momentary assessment (EMA), and questionnaire evidence. Using the Generalization of Longitudinal Behavior Modeling (GLOBEM) dataset, we construct 14,592 participant-day instances and align multimodal evidence to 29 B-HiTOP items across five spectra. Since GLOBEM lacks B-HiTOP responses, we evaluate evidence compatibility (C) rather than diagnostic accuracy, separating substantive predictions from abstentions when evidence is insufficient for item-level scoring. Two-stage prediction improves C for EMA and questionnaire evidence, but reduces C under

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Beyond the Chatbot: Co-Learning and Co-Teaching through a Dual-Persona Generative-AI Assistant

arXiv:2608.24902v1 Announce Type: new Abstract: In this paper we present a generative AI application developed to support both teachers and students in secondary education. The system employs two Large Language Models-LLMs, Gemini and DeepSeek, and a Small Language Model-SLM, Gemma, integrated within a Retrieval Augmented Generation - RAG framework, creating a pedagogically grounded, Greek-language assistant capable of adapting its reasoning and communication style to the user role. Unlike conventional chatbots, the assistant introduces pedagogical persona switching, a dual-role mechanism that enables the same AI model to act as both a teaching companion and a learning guide. Utilizing a RAG paradigm tailored to the Greek educational domain, the architecture segments official textbooks into coherent units. Enriched with specific metadata, these units preserve curricular structure and instructional context, demonstrating how generative AI optimizes modern instructional design. The initi

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Stronger Alignment between Brain Activity and LLM Embeddings during Code Writing compared to Prose Writing

arXiv:2608.24900v1 Announce Type: new Abstract: Programming is a critical skill underlying modern software systems, yet the cognitive processes supporting code writing are only beginning to be understood, limiting educational practices and developer tools. At the same time, Large Language Models (LLMs) are increasingly used to assist programming. These models themselves are not well understood and can exhibit undesirable behavior like introducing security vulnerabilities. Given evidence that some cognitive representations may be shared between LLMs and the brain, we seek to improve our understanding on both fronts by relating these two systems to one another. We used Voxelwise Encoding Models (VEMs) to relate LLM embeddings to brain activity measured with functional Magnetic Resonance Imaging (fMRI) during naturalistic writing tasks. Using participants' (n = 23) keystrokes as prompts, we extracted LLM embeddings to predict voxelwise Blood Oxygen Level Dependent (BOLD) signal, quantifyi

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

MCP-Driven Accessibility Tree Standardization for AI-Powered Screen Reader Agents

arXiv:2608.24898v1 Announce Type: new Abstract: Large language model (LLM) agents that interact with graphical user interfaces increasingly rely on either raw screenshots or platform-specific accessibility application programming interfaces (APIs) to perceive interface state. Both approaches have limitations for assistive applications: screenshot-based perception lacks the semantic roles and relationships required by screen readers, while platform-specific APIs such as Windows UI Automation, macOS Accessibility, Android AccessibilityService, and web ARIA require separate integrations for each platform. This paper proposes an architecture that uses the Model Context Protocol (MCP) as a unified transport and schema layer between heterogeneous accessibility frameworks and LLM-based assistive agents. An MCP accessibility server exposes ARIA-aligned roles, labels, states, and focusable-element hierarchies through a platform-independent representation, enabling consistent interaction across

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Agentic World Analysis (AWA) - an alternative way to explore systems and support decision making

arXiv:2608.24896v1 Announce Type: new Abstract: To address increasingly pressing sustainability challenges, various approaches have been developed to foresee possible futures, identify failure modes, detect vulnerabilities, and test potential mitigations. However, environmental systems are highly complex. Especially when coupled with human processes, the scale of uncertainties becomes intractable. To address this challenge, we propose a new approach - Agentic World Analysis (AWA)- combining the strengths of simulation modelling and expert elicitation. The concept of AWA is defined by three properties: 1) AWA uses an agentic AI system to mimic an expert panel that studies the world; 2) AWA projects futures iteratively through analysing scenario trees and learning from this analysis to improve decisions; 3) AWA is auditable. Based on these requirements, we implemented the World Engine by Generative Agents (WEGA) as a possible application of the AWA approach and demonstrated its functiona

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

IDEAlign: Comparing Ideas of Large Language Models to Domain Expert

arXiv:2509.02855v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to produce open-ended, interpretive annotations, yet there is no validated, scalable measure of idea-level similarity to expert annotations. We (i) introduce the content evaluation of LLM annotations as a core, understudied task, (ii) propose IDEAlign for capturing expert similarity judgments via pick-the-odd-one-out tasks, and (iii) benchmark various similarity methods (text embeddings, topic models, and LLM-as-a-judge) against these human ratings. Applying this approach to two real-world educational datasets (e.g., interpreting math reasoning and feedback generation), we find that most metrics fail to capture the nuanced dimensions of similarity meaningful to experts. LLM-as-a-judge performs best (11~18% improvement over other methods) but still falls short of expert alignment, making it useful as a triage tool rather than a substitute for human review. Our work demonstrates t

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Triadic Novelty: A Structural Typology of Science Innovation

arXiv:2506.17851v3 Announce Type: replace-cross Abstract: Scientific progress depends on novelty, but current evaluation systems often conflate novelty with recognition, favoring work that aligns with existing paradigms over ideas that challenge them. We introduce a theory-driven framework that conceptualizes novelty as a structural process rather than a single scalar outcome. Drawing on network science and theories of scientific discovery, we develop a triadic typology of novelty: Pioneers introduce new topics, Mavericks recombine distant areas of knowledge, and Vanguards reinforce weak but emerging connections. We apply this framework to philanthropic and nonprofit studies, an interdisciplinary and evolving field well suited to examining how novelty is recognized before evaluation norms are fully stabilized. Results show that novelty is not uniformly rewarded: Pioneer contributions are often weakly recognized unless later taken up, Maverick contributions receive consistent recognitio

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia

arXiv:2602.18455v5 Announce Type: replace Abstract: Search engines increasingly display AI-generated answers above organic links, potentially displacing traffic to upstream publishers. We estimate the impact of Google's AI Overviews (AIO) on Wikipedia's search traffic using AIO's staggered geographic rollout and Wikipedia's multilingual structure. Our difference-in-differences design compares monthly external-search referrals to English Wikipedia articles with referrals to the same articles in German and French, and finds that default AIO availability reduced English search traffic by 5.45% and 4.82%, respectively. Our results suggest that answer-producing digital intermediaries can materially reallocate attention away from informational publishers, with implications for content monetization, search platform design, and policy.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Operational Agency: A Permeable Legal Fiction for Tracing Culpability in AI Systems

arXiv:2602.17932v3 Announce Type: replace Abstract: Modern artificial intelligence (AI) systems act with a high degree of independence yet lack legal personhood-a paradox that fractures doctrines grounded in human-centric notions of mens rea and actus reus. This Article introduces Operational Agency (OA)-a permeable legal fiction structured as an ex post evidentiary framework-and Operational Agency Graph (OAG), a tool for mapping causal interactions among human actors, organizations, and AI systems. OA evaluates an AI's observable operational characteristics: its goal-directedness (as a proxy for intent), predictive processing (as a proxy for foresight), and safety architecture (as a proxy for a standard of care). OAG operationalizes that analysis by embedding these characteristics in a causal graph to trace and apportion culpability among developers, fine-tuners, deployers, and users. Drawing on corporate criminal liability, the innocent-agent doctrine, and secondary and vicarious lia

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Copyright Laundering Through the AI Ouroboros: Adapting the 'Fruit of the Poisonous Tree' Doctrine to Recursive AI Training

arXiv:2601.02631v3 Announce Type: replace Abstract: Copyright enforcement rests on an evidentiary bargain: a plaintiff must show both the defendant's access to the work and substantial similarity in the challenged output. That bargain comes under strain when AI systems are trained through multi-generational pipelines with recursive synthetic data. As successive models are tuned on the outputs of its predecessors, any copyrighted material absorbed by an early model is diffused into deeper statistical abstractions. The result is an evidentiary blind spot where overlaps that emerge look coincidental, while the chain of provenance is too attenuated to trace. These conditions are ripe for "copyright laundering"--the use of multi-generational synthetic pipelines, an "AI Ouroboros," to render traditional proof of infringement impracticable. This Article adapts the "fruit of the poisonous tree" (FOPT) principle to propose a AI-FOPT standard: if a foundational AI model's training is adjudged in

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

EduAgentQG: Multi-Agent Personalized Mathematics Question Generation with Explicit Diversity and Objective-Aware Evaluation

arXiv:2511.11635v2 Announce Type: replace Abstract: In intelligent education, personalized mathematics question generation aims to produce mathematics questions that satisfy educational requirements while supporting adaptive assessment and learning. Existing LLM-based single-agent and multi-agent methods improve generation flexibility, but they still tend to rely on aggregated feedback or model randomness, making it difficult to jointly ensure dimension-wise objective alignment and controllable diversity. To address these challenges, we propose EduAgentQG, a multi-agent collaborative framework for personalized mathematics question generation with explicit diversity and objective-aware evaluation. EduAgentQG organizes question generation as a closed-loop process of planning, writing, evaluation, refinement, and checking: structured generation plans and multiple generation directions guide candidate generation, while fine-grained evaluation verifies logical correctness, solvability, and

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Hallucination to Reliability: Generative Modeling and the Structure of Scientific Inference

arXiv:2504.08526v3 Announce Type: replace Abstract: Generative AI is increasingly used in science, but is unavoidably prone to hallucination. I develop a reliabilist account of how generative AI nevertheless gives rise to new scientific knowledge. I analyze hallucinations as non-strategic misrepresentations of the target phenomenon, introduced by a model's generative activity, rather than inherited from training data. Through case studies of AlphaFold and SEEDS, I show how scientific workflows draw on pre-existing knowledge of target phenomena to filter or qualify hallucinatory outputs, thereby preventing their erroneous content from propagating into downstream inference. Finally, I show that workflows are units of epistemic evaluation in their own right.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Learning from waste: Machine Learning for health risk prediction and computer vision-based sorting in Ghana

arXiv:2608.25759v1 Announce Type: cross Abstract: The inappropriate disposal of solid waste remains a significant public health and environmental concern worldwide, including in Ghana. Poor sanitation and improper waste management practices contribute to substantial economic costs and avoidable deaths annually. In 2022, a field study in Atonsu, Kumasi, Ghana, reported a community-perceived relationship between household waste disposal and illness patterns, but only through descriptive analysis without quantitative validation. This study extends that investigation using two data-driven approaches. First, a Random Forest classifier was developed to predict illness categories using waste disposal practices and demographic survey data. On a held-out group of respondents who reported illness (N=69), the model obtained a macro F1 score of 0.63, with disposal method emerging as the most important substantive predictor of illness type. Second, a MobileNetV2 image classification model enabled a

Source ↗
Showing 6001–6050 of 10879 signals
← Prev Page 121 of 218 Next →