Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2608.07504v1 Announce Type: new Abstract: We propose a human bottleneck perspective for understanding how generative AI transforms the innovation process. The central premise is that many constraints traditionally plaguing the innovation process are cognitive and social in origin, rooted in how people generate ideas, evaluate novelty, and communicate through social systems. Generative AI does not act uniformly on these constraints. At each stage, it can deepen some bottlenecks while alleviating others, and predicting these outcomes requires understanding the underlying mechanisms of the constraint itself. We identify bottlenecks in four stages of the innovation process: ideation, screening and testing, preference measurement and consumer insight, diffusion, and market learning. By grounding analysis in human behavior rather than rapidly changing AI capabilities, we offer a framework for assessing whether new developments alleviate or intensify the bottlenecks that matter most at
arXiv:2608.07503v1 Announce Type: new Abstract: Interdisciplinary project-based learning requires students to negotiate differences in language, assumptions, priorities, and working practices. These differences are difficult to surface in text-based team communication, where discussions can become fragmented and AI tools are often used as private side channels rather than shared supports for collective sensemaking. We present Spritz, a Discord-based LLM technology probe that explores how AI might mediate disciplinary boundaries in student project teams. Spritz monitors group chat for signals of semantic or pragmatic boundaries, prompts members to articulate their perspectives through private channels, and returns anonymized syntheses to the shared discussion. We conducted a technology probe study and co-design workshop with 12 university students from technical, business, and design backgrounds. Participants experienced Spritz during a simulated interdisciplinary resource-allocation ta
arXiv:2608.07502v1 Announce Type: new Abstract: This document introduces HAR-IMU-IL, a dataset developed for human activity recognition (HAR) using inertial measurement unit (IMU) sensors within a smart home environment with a focus to support objective functional assessment of older adults' independent living (IL). In particular, HAR-IMU-IL includes recordings of 50 participants performing 17 clinically relevant activities of daily living, spanning 4 functional domains essential for independent living: mobility, hygiene, nutrition and hydration, and medication intake. The dataset was collected using 30 IMU sensors, comprising both wearable and object-mounted devices integrated within a real-world residential setting. The dataset includes multi-sensor inertial data captured under realistic, unconstrained conditions, together with detailed annotations ensuring high temporal accuracy and consistency across sensors. A comprehensive data collection protocol was implemented to preserve ecol
arXiv:2608.07501v1 Announce Type: new Abstract: Navigating an unfamiliar city poses significant challenges, and tourists are among the groups more likely to experience them, particularly when attempting to locate a point of interest (POI). Various factors, such as language barriers or a lack of precise information, further complicate this issue by making it difficult for visitors to explore efficiently. To address these issues, we introduce the tool "Personalized Assistant for Tourist Hints" (PATH), which is based on a methodology for estimating personalized tourist routes. PATH computes optimal paths to specified locations guided by two tailored heuristics: tourist frequency and POI preference. This integration enables efficient and personalized navigation, guiding users toward attractions that best match their preferences. We evaluated our methodology through a case study in Viterbo, Italy, using a dataset comprising over 1000 routes from 265 tourists who visited the city during diff
arXiv:2608.07500v1 Announce Type: new Abstract: Research on human-GenAI collaboration yields conflicting findings: GenAI can enhance creativity yet reduce collective diversity, with uneven benefits across skill levels. Rather than treating these as contradictions, we argue they reflect a core feature of GenAI: abundance. GenAI makes ideas, drafts, and recombinations plentiful, potentially expanding the hypothesis space and surfacing unanticipated possibilities. However, abundance alone doesn't ensure better outcomes. We propose generative fit as a unifying mechanism explaining when abundance yields productive creativity and when it backfires. Drawing on Generativity Theory, generative fit captures how well a system's generative potential complements a community's generative capacities. We develop a conceptual framework for collaborative human-GenAI settings where participants share goals, depend on one another, and must integrate diverse contributions. By mapping abundance to cognitive
arXiv:2608.07499v1 Announce Type: new Abstract: The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based simulated clients. Prior work on simulated clients, however, has not aligned with the specific tasks fundamental to the MI therapy approach. A key task is evoking, in which the counsellor first elicits the client's ambivalence and then strengthens the client's motivation for change. We present Evoke-Sim, a task-aware, multi-stage LLM-based client simulation framework for evaluating MI counsellors in smoking cessation, designed specifically for the evoking MI task. Evoke-Sim employs structured client profiles, an evoking-specific three-stage conversation flow, and a reveal policy that regulates which client profile information might be disclosed at each stage. We show that compared to existing profile-grounded simulated clients, Evoke-Sim is better at differentiating levels of MI quality using task-awa
arXiv:2608.07498v1 Announce Type: new Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing. A central open question is whether persona-prompted LLMs can simulate individual-level social media reactions with sufficient accuracy to support either application, and how accuracy depends on profile completeness, model selection, and the generalization challenge posed by novel post content. This study benchmarks twelve LLM configurations on binary like/dislike prediction across 296 survey-based agent profiles and 26 ground-truth-mapped posts under three profile conditions, with leave-post-out machine learning classifiers as baselines. Across full-profile conditions, accuracy ranges from 75.54% to 96.68%, with a 30-point spread attributable primarily to model selection and confirmed by paired McNemar tests with agent-level bootstrap intervals. GPT-5.5 Pro ac
arXiv:2608.07497v1 Announce Type: new Abstract: Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or powering teachable agents. Recently, large language models (LLM) have enabled richer, more naturalistic interactions with simulated learners; however, no open framework exists for evaluating whether such simulations faithfully reproduce real learner behavior. We introduce EvalConvoLearn, an open-source framework that assesses learner simulations along two axes: learning behavior (skill-conditioned mastery outcomes) and conversational quality (talk moves, error type distributions, question rate, turn length). EvalConvoLearn measures how closely a simulated learner approximates answer distributions observed in data by grounding metrics in authentic tutoring conversation datasets, and anchoring generated tutor responses in existing tutor utterances. The framework is demonstrated on a dataset of tuto
arXiv:2608.07496v1 Announce Type: new Abstract: Agent-based models have historically served as tools for generative explanation, constructing testbeds in which candidate micro-level behavioral rules can be tested for their capacity to produce observed macro-level phenomena. The integration of Large Language Models into agent-based simulation has expanded what these models can represent, but it has also introduced an unexamined shift in how users engage them. We argue that current generative agent-based models (GABMs) inherit the dominant interaction metaphor of conversational LLM interfaces - a question-answer pattern that positions users as consumers of system output rather than explorers of a possibility space. In the context of policy, where problems are wicked and ground truth is unknowable in advance, this metaphor produces a trust deficit that cannot be resolved through improved model accuracy alone. We open a design space we call human-simulation interaction, and argue that warr
arXiv:2608.07495v1 Announce Type: new Abstract: Effective communication during palliative care discussions is a critical clinical skill, yet training clinicians to manage complex patient emotions remains challenging. Large language model (LLM)-based patient simulators provide a scalable approach for communication training, but most existing systems treat patient emotion as static and fail to capture the dynamic emotional shifts observed in clinical interactions. We present EmoPatient, an emotion-directed patient simulator designed to generate evolving emotional responses during palliative care discussions. The system introduces an Emotion Director agent that estimates the patient's emotional state and generates turn-level control signals for emotional intensity, regulatory stability, and interactional guidance. We evaluate EmoPatient through controlled multi-turn physician-patient dialogue simulations and compare it with baseline simulators. Results show improvements across four theory
arXiv:2608.07494v1 Announce Type: new Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as ``plan a three-day Vienna trip'', ``solve the attached mathematical problem'', ``draft an email to inquire review progress'', etc., which are also known as LLM prompts. Crafting clear and well-structured prompts leads to more appropriate LLM feedback, which effectively bridges human-LLM interaction. Although prompting appears accessible to non-expert users, precisely organizing effective prompts is a highly systematic and skillful process, presenting potential challenges even for experienced users. This survey explores the principles, taxonomy, and organization of prompts from a user-centered perspective. Differing from the existing surveys that primarily focus on technical principles and application scenarios of LLMs, this paper provides actionable guidelines for for
arXiv:2608.07493v1 Announce Type: new Abstract: Current AI disclaimers often fail to function as intended due to warning habituation and a transparency paradox. As AI-generated information becomes pervasive in everyday decision-making, effective risk communication is increasingly critical for responsible design. This exploratory study examines how disclaimer placement and persuasive cues shape trust, perceived accuracy, and disclaimer engagement across three high-stakes domains: finance, medicine, and AI-generated content. Using a mixed within-between experimental design with 378 stimulus-level responses from 52 participants, we find that advisory content was generally trusted across conditions, even when disclaimers were present. A significant domain effect showed that medical content received the highest trust ratings. In the AI domain, the findings reveal a transparency paradox: some participants interpreted disclaimers not as warnings, but as signs of system self-awareness and hone
arXiv:2608.07492v1 Announce Type: new Abstract: Religion is an important part of many people's lives, with reading from religious texts being among the most common and important regular practices. Modern technology has in many ways changed the way this religious reading takes place, but the effects of these changes have not yet been studied. While there have been plenty of studies examining the effects of digital media on reading comprehension and similar topics, the study of religious texts is often primarily focused on achieving a religious experience, so the existing research is not sufficient to explore these changes. In our study, we used surveyed students from two universities to ask individuals how they use technology in their personal study of religious texts and religious education courses. Our respondents were predominately Christian. Through qualitative and quantitative analysis of the results, we aim to learn how technology interacts with study of religious texts, in what s
arXiv:2608.07491v1 Announce Type: new Abstract: Ontologies are widely used biomedical science and clinical practice. However, no recent works have analyzed the usability of ontology development software. We survey ontology researchers to assess the usability of 15 front-end ontology tools using the System Usability Scale (SUS). Among 38 respondents, Protege and WebProtege were most used but showed only moderate usability (SUS ~60). Familiarity significantly predicted usability scores (p=0.016). Results highlight a usability gap in ontology tooling critical for advancing biomedical data integration.
arXiv:2608.07490v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players learn from repeated interaction. We study experience-sensitive game learning: how gameplay experience changes the decision-making behavior of humans and language agents. We formulate experience-sensitive game learning as a framework for analyzing behavioral change across repeated gameplay, rather than only final score or win rate. We introduce a suite of interactive games with reusable strategic structure, together with cross-game greedy-to-global metrics and game-specific behavioral diagnostics that make experience-driven change observable from action traces. We also collect repeated-game trajectories from human players and evaluate recent self-evolving language agents in the same behavioral metric space. Our results show that human players exhibit interpretable and relatively stable shifts from local
arXiv:2608.07489v1 Announce Type: new Abstract: Human-environment interactions, a classic topic in geography, suggest that individuals and their environments might shape each other. Yet the specific mechanisms underlying these interactions regarding human personality traits have not been explored. This study examines the associations between human Big Five personality traits and built environment characteristics derived from street view imagery across four cities in Texas, United States, providing a descriptive foundation for understanding these complex human-environment dynamics. By integrating fine-resolution self-reported personality assessments with computer vision analysis of urban environments, we identified significant spatial clustering of personality traits at the ZIP code level. Our regression analyses reveal that built environment features and socioeconomic characteristics explain substantial variance in personality distributions, with Openness showing the strongest model fi
arXiv:2608.07486v1 Announce Type: new Abstract: Virtual reality (VR) and head-mounted displays are constantly gaining popularity in various fields such as education, military, entertainment, and bio/medical informatics. Although such technologies provide a high sense of immersion, they can also trigger symptoms of discomfort. This condition is called cybersickness (CS) and is quite popular in recent publications in the virtual reality context. This work proposes a novel experimental analysis using symbolic machine learning that ranks potential causes for CS. We estimate the CS causes and rank them according to their impact on the classification capabilities of CS. The experiments are performed using two distinct virtual reality games. We were able to identify that acceleration triggered cybersickness more frequently in a race game in contrast to a flight game. Furthermore, participants less experienced with VR are more prone to feel discomfort and this variable has a greater impact in
arXiv:2608.07484v1 Announce Type: new Abstract: Virtual reality (VR) is an imminent trend in games, education, entertainment, military, and health applications, as the use of head-mounted displays is accessible to everyone. While VR provides immersive experiences, it still does not offer an entirely perfect situation, mainly due to cybersickness (CS) issues. In this work, we propose a novel approach for predicting upcoming CS symptoms. Our solution is able to suggest whether the user of VR is entering into an illness situation. We adopted random forest classifiers and validated our solution using 16 different machine-learning techniques, which presented the best results. For training purposes, we built our own dataset through a CS profile questionnaire that we also propose in the present work. The questionnaire is focused on registering and identifying the user's susceptibility to CS, considering their historical conditions and also their response to the immersive environment developed
arXiv:2608.07483v1 Announce Type: new Abstract: Traditional human-centered design is grounded in human needs, behaviors, and experiences, especially throughout the development of products and services. Recent advances in artificial intelligence and machine learning, accelerated by the public launch of ChatGPT in November 2022, have created a wave of AI adoption across research, industry, and design practice. While some of this rapid adoption reflects the current enthusiasm and overuse surrounding AI tools, it has also introduced lasting changes to how designers and researchers understand users, generate ideas, evaluate systems, and make design decisions. This position paper examines how AI is reshaping human-centered design and argues that some of these changes are likely to extend beyond short-term hype. It also considers the risks and limitations introduced by this shift, including over-reliance on AI, reduced human judgment, bias, data security concerns, and unclear accountability.
arXiv:2608.07482v1 Announce Type: new Abstract: Generative AI is reshaping HCI work by accelerating brainstorming, writing, prototyping, and interface generation. However, faster production does not automatically produce human-centered design. This position paper argues that GenAI should be treated not as a replacement for human-centered methods, but as a co-thinking partner driven by human goals, ethics, cognition, and accountability. After reflecting on HCI coursework, design activities, and AI-assisted workflows, I argue that traditional HCD principles become more important as AI systems become more capable. Users still bring limited attention, mental models, trust issues, and cognitive biases into interaction, while GenAI introduces new risks around over-reliance, de-skilling, shallow reasoning, hallucination, unclear authorship, and reduced accountability. The future of HCI should therefore focus on human-AI collaboration that preserves human agency, critical thinking, and respons
arXiv:2608.07481v1 Announce Type: new Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style task. Two models - GPT-4o as Czar and Claude Opus-4.5 as Player - are evaluated on a binary humor-selection task constructed so that success cannot follow from self-preference. A reflected-cell stability procedure isolates 244 hands on which the two models hold deterministic but opposite preferences, partitioned into a 97-hand context pool and a 147-hand held-out test pool. The Player is then evaluated across five graded conditions: default self-preference, generic Czar-modeling instruction, model-identified Czar, prior Czar selections, and prior Czar selections with rationales. This gradient is designed to separate two sources of improvement: framing effects, in which the Player is told to attend to a Czar without seeing any of the Czar's behavior, and direct behavioral evidence, in which
arXiv:2608.07477v1 Announce Type: new Abstract: This thesis examines the fairness of Automated Machine Learning (AutoML) tools in human resource hiring systems through the combined lenses of regulation, business strategy, and Human-Computer Interaction (HCI). It argues that fairness is no longer merely an ethical concern but a critical determinant of usability, trust, legal compliance, and organizational adoption. While AutoML platforms improve efficiency by simplifying model selection and deployment, they also risk perpetuating discriminatory outcomes when trained on biased historical hiring data. Existing platforms prioritize technical performance over fairness, leaving non-expert business users unable to detect or mitigate bias effectively. The study investigates fairness gaps in AutoML tools through four research questions focused on fairness mechanisms, interface transparency, human oversight, and product design priorities. Drawing on frameworks such as the Technology Acceptance M
arXiv:2606.16337v3 Announce Type: replace-cross Abstract: Predictive modeling for clinical decision support requires not only strong predictive performance but also transparent decision logic. Although deep learning and tree-based ensemble methods can achieve high accuracy, their black-box nature remains a major obstacle to clinical deployment. This challenge is further compounded by common characteristics of medical data, including limited sample sizes, severe class imbalance, and feature evolution arising from changes in diagnostic criteria and clinical documentation. To address these issues, we propose Medical Heuristic Learning (MHL), an instantiation of the learning beyond gradients paradigm for clinical prediction from structured medical data. Instead of relying on neural network weight updates, MHL uses a large language model (LLM) driven workflow that integrates statistical probes, medical knowledge probes, rule synthesis, and code-level iterative refinement to optimize a deter
arXiv:2606.15198v2 Announce Type: replace-cross Abstract: City landscapes viewed through home windows influence quality of life, yet perceptions of actual window views at the urban scale remain understudied. This study presents an approach for large-scale mapping of perceptions using 12,334 window view images (WVIs) collected from actual residential properties listed on real estate platforms in Wuhan, China, representing a rarely explored form of urban view imagery that offers advantages over the rendered or simulated window views commonly examined in previous studies. Through a non-immersive virtual reality platform, we collected 27,477 pairwise comparisons across six perceptual dimensions (e.g. preference) from 304 participants based on 499 WVIs. A hybrid neural network model was trained to predict human perceptions of all crowdsourced WVIs and map their spatial distribution. Results reveal significant spatial autocorrelation with distinct hot and cold spots across the whole city. Fl
arXiv:2606.08281v2 Announce Type: replace-cross Abstract: Safe physical human-robot interaction (pHRI) is fundamentally a problem of interaction dynamics: the robot must track a commanded motion, yield under human forces, respect actuator and joint limits, and stay predictable under persistent contact. Classical impedance control shapes this through a virtual spring-damper, but a sustained force produces the bias $e_\infty=-K_d^{-1}F_h$, trading accuracy for safety. We propose a predictive framework that makes interaction dynamics explicit through a linear double-integrator backbone: an operational-space feedforward cancels gravity and Coriolis terms and normalizes the task inertia, leaving a configuration-independent state-transition matrix with robot dependence isolated in the input matrix. This converts nonlinear torque-controlled pHRI into a linear constrained-control problem, so offset-free tracking, actuator feasibility, sampled-data joint-limit safety, and passivity filtering fo
arXiv:2605.22774v3 Announce Type: replace-cross Abstract: Assessing cognitive load continuously and at low latency would help adaptive human-computer interaction, but it remains hard because labeled data are scarce and models generalize poorly across subjects. Recent ECG foundation models, pre-trained on millions of clinical diagnostic ECG recordings, yet they do not apply directly to wearable devices when the sensor configuration and the task both differ. We present CogAdapt, a framework that adapts a clinical ECG foundation model to wearable cognitive load assessment. CogAdapt has two parts. LeadBridge is a learnable adapter that maps 3-lead wearable signals to a 12-lead-compatible representation. ProFine is a progressive fine-tuning strategy that unfreezes encoder layers in stages while limiting representational drift in the pre-trained model. On two public datasets (CLARE and CL-Drive) under leave-one-subject-out cross-validation, CogAdapt reaches macro-F1 of 0.626 and 0.768, impro
arXiv:2604.25596v2 Announce Type: replace-cross Abstract: Formal models for concurrent and distributed systems describe machines; the people who operate them are either ignored or treated as external environment. Yet, key distributed systems -- notably grassroots platforms -- include people operating their personal machines (smartphones), and their faithful description must include the states of both people and machines and how they jointly effect system behaviour. Here, we propose volition-guarded multiagent atomic transactions -- executed atomically by machines and guarded by their people's volitions -- as a novel mathematical foundation for specifying systems consisting of people operating machines. Each agent's state consists of a volitional state and machine state; a transaction is enabled when the machine precondition holds and the guarding persons are willing. For example, befriending two people is guarded by both; unfriending, by either; voluntary swap of coins and bonds is gua
arXiv:2604.11730v4 Announce Type: replace-cross Abstract: Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits that improve medical outcomes. In-person interventions are costly and difficult to scale, especially in resource-limited regions. Digital health interventions offer a cost-effective approach, potentially supporting independent living and self-management. Automating such interventions, especially through machine learning, has recently gained considerable attention. Ambivalence and hesitancy (A/H) play a primary role for individuals to delay, avoid, or abandon health interventions. A/H are subtle and conflicting emotions that place a person in a state between positive and negative evaluations of a behaviour, or between acceptance and refusal to engage in it. They manifest as affective inconsistency across modalities or within a modality, such as language, facial, vocal expressions,
arXiv:2603.02070v3 Announce Type: replace-cross Abstract: When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facilitate an iterative reasoning and elicitation process, where the human's role is to guide the AI planner according to their preferences and expertise. In this context, explanations that respond to users' questions are crucial to improve their understanding of potential solutions and increase their trust in the system. To enable natural interaction with such a system, we present a multi-agent Large Language Model (LLM) architecture that is agnostic to the explanation framework and enables user- and context-dependent interactive explanations. We also describe an instantiation of this framework for goal-conflict explanations, which we use to conduct a user study comparing the LLM-powered interaction with a baseline template-based explanation interface.
arXiv:2602.19107v3 Announce Type: replace-cross Abstract: Robotaxis are emerging as a promising form of urban mobility, but removing human drivers fundamentally reshapes passenger-vehicle interaction and raises new design challenges. To inform robotaxi design based on real-world experience, we conducted 18 semi-structured interviews and autoethnographic ride experiences to examine users' perceptions, experiences, and expectations for robotaxi design. We found that users valued benefits such as increased agency and consistent driving. However, they also encountered challenges such as limited flexibility, insufficient transparency, and emergency handling concerns. Notably, users perceived robotaxis not merely as a mode of transportation, but as autonomous, semi-private transitional spaces, which made users feel less socially intrusive to engage in personal activities. Safety perceptions were polarized: some felt anxiety about reduced control, while others viewed robotaxis as safer than h
arXiv:2512.10785v3 Announce Type: replace-cross Abstract: Generative AI offers new opportunities for individualized and adaptive learning, e.g., through large language model (LLM)-based feedback systems. While LLMs can produce factually correct feedback for relatively straightforward conceptual tasks, delivering high-quality feedback for tasks that require advanced domain expertise, such as physics problem solving, remains a substantial challenge. This study presents the design and implementation of an LLM-based feedback system for physics problem solving grounded in evidence-centered design and reports a first evaluation within the German Physics Olympiad. Participants rated the usefulness and correctness of the generated feedback for each implemented problem. The collected ratings indicate that the feedback was generally perceived as useful and highly correct. However, an in-depth analysis revealed that the feedback contained errors in 20% of cases; errors that often went unnoticed b
arXiv:2508.21010v3 Announce Type: replace-cross Abstract: Existing Causal-Why Video Question Answering (VideoQA) models often struggle with higher-order reasoning, relying on opaque, monolithic pipelines that entangle video understanding, causal inference, and answer generation. These black-box approaches offer limited interpretability and tend to depend on shallow heuristics. We propose a novel, modular paradigm that explicitly decouples causal reasoning from answer generation, introducing natural language causal chains as interpretable intermediate representations. Inspired by human cognitive models, these structured cause-effect sequences bridge low-level video content with high-level causal reasoning, enabling transparent and logically coherent inference. Our two-stage architecture comprises a Causal Chain Extractor (CCE) that generates causal chains from video-question pairs, and a Causal Chain-Driven Answerer (CCDA) that derives answers grounded in these chains. To address the la
arXiv:2508.16771v3 Announce Type: replace-cross Abstract: Code Language Models (CodeLLMs) learn token importance from data correlations, whereas human developers attend selectively to semantically salient code. We present EyeMulator, a model-agnostic method that injects human visual-attention priors into CodeLLM fine-tuning without architectural changes. EyeMulator distills eye-tracking data into semantic salience and gaze-transition priors, then uses them to reweight token-level training losses. Across six backbones, two data regimes, and three CodeXGLUE tasks, the reported configurations yield positive matched-metric deltas in all 36 model-task-setting cells. Effects are largest for structure-preserving completion and translation, while summarization shows smaller but positive METEOR deltas. Session-mode and component-ablation analyses further show that reading, writing, semantic, and transition-derived priors provide complementary signal. Human-attention artifacts are available at h
arXiv:2507.02950v3 Announce Type: replace-cross Abstract: Large language models (LLMs) may support counseling training, yet evidence from Japanese-language interactions and automated quality ratings remains limited. We examined 18 fixed Japanese-language counseling transcripts generated through artificial intelligence (AI)-to-AI interactions under three counselor conditions: GPT-minimal (GPT-4-turbo with a minimal role instruction), GPT-SMDP (GPT-4-turbo with the Structured Multi-step Dialogue Prompt [SMDP]), and Claude-SMDP (Claude-3-Opus with SMDP). Fifteen counseling experts rated transcripts on four adapted global scales from the Motivational Interviewing Treatment Integrity coding manual and an overall-quality item; three newer LLMs independently rated the same transcripts in three iterations. In this fixed stimulus set, SMDP-condition dialogues received higher expert ratings for cultivating change talk, partnership, empathy, and overall quality than GPT-minimal dialogues; the two
arXiv:2504.09662v4 Announce Type: replace-cross Abstract: Multi-agent large language model simulations have the potential to model complex human behaviors and interactions. If the mechanics are set up properly, unanticipated and valuable social dynamics can surface. However, it is challenging to consistently enforce simulation mechanics while still allowing for rich and emergent dynamics. We present AgentDynEx, an AI system that helps set up, track, and repair simulations. Specifically, AgentDynEx introduces milestones that act as checkpoints and failure conditions that act as guardrails to ensure dynamics are relevant and mechanics are respected as the simulation progresses. It also introduces a method called nudging, where the system dynamically reflects on simulation progress and gently intervenes if it begins to deviate from intended outcomes. A technical evaluation found that nudging enables simulations to progress further without reducing the presence notable dynamics compared to
arXiv:2410.20696v3 Announce Type: replace-cross Abstract: Case studies have shown that software disasters snowball from technical issues to catastrophes through humans covering up problems rather than addressing them and empirical research has found the psychological safety of software engineers to discuss and address problems to be foundational to improving project success. However, the failure to do so can be attributed to psychological factors like loss aversion. We conduct a large-scale study of the experiences of 600 software engineers in the UK and USA on project success experiences. Empirical evaluation finds that approaches like ensuring clear requirements before the start of development, when loss aversion is at its lowest, correlated to 97% higher project success. The freedom of software engineers to discuss and address problems correlates with 87% higher success rates. The findings support the development of software development methodologies with a greater focus on human fa
arXiv:2202.14019v3 Announce Type: replace-cross Abstract: Maintaining proper form while exercising is important for preventing injuries and maximizing muscle mass gains. Detecting errors in workout form naturally requires estimating human's body pose. However, off-the-shelf pose estimators struggle to perform well on the videos recorded in gym scenarios due to factors such as camera angles, occlusion from gym equipment, illumination, and clothing. To aggravate the problem, the errors to be detected in the workouts are very subtle. To that end, we propose to learn exercise-oriented image and video representations from unlabeled samples such that a small dataset annotated by experts suffices for supervised error detection. In particular, our domain knowledge-informed self-supervised approaches (pose contrastive learning and motion disentangling) exploit the harmonic motion of the exercise actions, and capitalize on the large variances in camera angles, clothes, and illumination to learn
arXiv:2606.22484v2 Announce Type: replace Abstract: The adoption of agentic AI coding systems -- where autonomous agents generate, review, test, and deploy code with minimal human intervention -- creates a governance challenge in regulated industries. Existing frameworks address AI-assisted development maturity or the productivity-reliability tension but offer no mechanism for calibrating human oversight intensity to regulatory impact. We present the Governed AI-Assisted Engineering (GAIE) framework, a three-tier graduated human oversight model for agentic code generation in regulated domains. GAIE introduces the Oversight Classification Model (OCM), a deterministic decision function that classifies code generation tasks by regulatory impact, customer proximity, reversibility, and data sensitivity to route them through one of three oversight tiers: human-in-the-loop (strategic functions), human-over-the-loop (customer-impacting), or automated-with-monitoring (internal). Each tier defin
arXiv:2605.15932v2 Announce Type: replace Abstract: Designing safe and sustainable chemicals is critical to combat chemical pollution in our environment. Computational and AI-assisted methods have been developed to aid de novo molecule design. However, data on the environmental impacts of chemical compounds are sparse, resulting in low-fidelity machine learning (ML) oracles and unreliable candidate proposals. Furthermore, many automated molecular design approaches rely on numerical scoring functions that cannot fully capture the nuanced chemical intuition of expert scientists required for real-world molecular design. Instead, we present GEMS - an interactive visual analytics tool for human-in-the-loop molecular optimization that lets domain experts directly collaborate with an evolutionary genetic algorithm. Users continuously guide the search using domain knowledge through high-level, parametric modification of the scoring function alongside direct, granular control over molecule popu
arXiv:2605.12613v3 Announce Type: replace Abstract: WhatsApp is one of the most widely used messaging platforms globally, with billions of users sharing information in private groups. Yet, it offers little infrastructure to support moderation and group governance. In the absence of platform-level oversight, group admins bear the responsibility of governing group behavior. In this paper, we explore how WhatsApp group admins collaborate with AI tools to create, enforce, and maintain group rules. Drawing on a two-phase speculative design study with 20 admins in India, we examine how participants interacted with an AI assistant (Meta AI) to co-create rules and responded to a series of probes illustrating AI-assisted moderation features. Our findings show that while admins appreciated the AI's ability to surface overlooked rules and reduce their moderation burden, they were highly sensitive to issues of relational trust, data privacy, tone, and social context. We identify how group type and
arXiv:2603.28944v2 Announce Type: replace Abstract: Artificial intelligence (AI) predictions are increasingly used to inform human decisions. Here, using a behavioral implementation of the classic Newcomb's paradox in 1,305 participants, we show that AI predictions can also shape the reasoning people use to make a decision. In this paradigm, perceived predictive authority can alter how people reason about their future actions, leading them to forgo a guaranteed reward. Over 40% of participants treated AI as such a predictive authority about their own behavior, significantly increasing the odds of forgoing the guaranteed reward by a factor of 3.39 (95% CI: 2.45-4.70) and reducing earnings by 10.7-42.9%. The effect appeared across AI presentations and decision contexts and remained detectable even when predictions repeatedly failed. When people perceive AI as capable of predicting their personal behavior, the mere presence of AI predictions may shape their decision-making, narrowing the
arXiv:2601.20749v2 Announce Type: replace Abstract: How do students develop AI literacy through everyday practice rather than formal instruction? While normative AI literacy frameworks proliferate, empirical understanding of how students actually learn to work with generative AI remains limited. This study analyzes 10,536 ChatGPT messages from 36 undergraduates over one academic year, revealing five use genres -- academic workhorse, emotional companion, metacognitive partner, repair and negotiation, and trust calibration -- that constitute distinct configurations of student-AI learning. Drawing on domestication theory and emerging frameworks for AI literacy, we demonstrate that functional AI competence emerges through ongoing relational negotiation rather than one-time adoption. Students develop sophisticated genre portfolios, strategically matching interaction patterns to learning needs while exercising critical judgment about AI limitations. Notably, repair work during AI breakdowns
arXiv:2601.11049v2 Announce Type: replace Abstract: We examine whether large language models (LLMs) can predict biased decision-making in conversational settings, and whether their predictions capture not only human cognitive biases but also how those effects change under cognitive load. In a pre-registered study (N = 1,648), participants completed six classic decision-making tasks via a chatbot with dialogues of varying complexity. Participants exhibited two well-documented cognitive biases: the Framing Effect and the Status Quo Bias. Increased dialogue complexity resulted in participants reporting higher mental demand. This increase in cognitive load selectively, but significantly, increased the effect of the biases, demonstrating the load-bias interaction. We then evaluated whether LLMs (GPT-4, GPT-5, and open-source models) could predict individual decisions given demographic information and prior dialogue. While results were mixed across choice problems, LLM predictions that incor
arXiv:2512.09802v4 Announce Type: replace Abstract: This paper presents the initial stages of a design study aimed at developing a dashboard to visualize gameplay data of the Commander format from Magic: The Gathering. We conducted a user-task analysis to identify requirements for a data visualization dashboard tailored to the Commander format. Afterward, we proposed a design for the dashboard, leveraging visualizations to address players' needs and pain points for typical data analysis tasks in the context domain. Then, we followed-up with a structured user test to evaluate players' comprehension and preferences of data visualizations. Results show that players prioritize contextually relevant, outcome-driven metrics over peripheral ones, and that canonical charts like heatmaps and line charts support higher comprehension than complex ones such as scatterplots or icicle plots. Our findings also highlight the importance of localized views, user customization, and progressive disclosure
arXiv:2509.13742v4 Announce Type: replace Abstract: Science communication revision requires writers to dynamically balance scientific exposition and narrative engagement - a process where writers often struggle with competing directions. Existing LLM-assisted tools help with co-writing, but offer limited support for navigating this iterative, multi-directional revision process. To address this gap, we designed Spatial Balancing, an exploratory revision environment that maps rhetorical goals and revision strategies onto a two-dimensional spatial canvas for experienced science communication creators with domain expertise but lacking formal professional training. By building a design space of communication strategies and embedding them into a spatial exploratory canvas, our system treats feedback as navigational cues rather than prescriptive judgments. Our findings show that this integrated revision environment helps writers stay focused on writing goals, reason about revision as trajecto
arXiv:2508.08242v2 Announce Type: replace Abstract: Group decision-making often suffers from uneven information sharing, hindering decision quality. While large language models (LLMs) have been widely studied as aids for individuals, their potential to support groups of users, potentially as facilitators, is relatively underexplored. We present a pre-registered randomized experiment with 1,475 participants assigned to 281 live groups completing a hidden profile task--selecting an optimal city for a hypothetical sporting event--under one of four facilitation conditions: no facilitation, a one-time message prompting information sharing, a human facilitator, or an LLM (GPT-4o) facilitator. We find that LLM facilitation increased information shared within a discussion by raising the minimum level of engagement with the task among group members, and that these gains came at limited cost in terms of participants' attitudes towards the task, their group, or their facilitator. Whether by human
arXiv:2507.12721v3 Announce Type: replace Abstract: Human-AI interfaces play a pivotal role in integrating clinicians' expertise with artificial intelligence to enhance both healthcare practice and research. However, designing effective interfaces in this domain remains a significant challenge. The inherent complexity of medical data, the influence of domain-specific conventions, and the diverse needs of clinical users compound the challenge of developing practical and usable solutions. In this study, we review existing solutions and synthesize a set of design patterns - recurring approaches that support the design of human-AI interfaces in clinical settings. We conducted a comprehensive literature review of human-AI interaction designs in clinical contexts, through which we identified 15 information entities commonly presented to users and 12 design patterns used to organize and communicate this information effectively. For each design pattern, we summarize the underlying design probl
arXiv:2507.03670v2 Announce Type: replace Abstract: Writing longer prompts for an AI assistant to generate a story increases psychological ownership, a user's feeling that the writing belongs to them. To encourage users to write longer prompts, we evaluated two interaction techniques that modify the prompt entry interface of chat-based generative AI assistants: pressing and holding the prompt submission button, and continuously moving a slider up and down when submitting a short prompt. A within-subjects experiment investigated the effects of such techniques on prompt length and psychological ownership, and results showed that these techniques increased prompt length and led to higher psychological ownership than baseline techniques. A second experiment further augmented these techniques by showing AI-generated suggestions for how the prompts could be expanded. This further increased prompt length, but did not lead to improvements in psychological ownership. Our results show that simpl
arXiv:2504.17331v3 Announce Type: replace Abstract: Locomotion plays a crucial role in shaping the user experience within virtual reality environments. In particular, hands-free locomotion offers a valuable alternative by supporting accessibility and freeing users from reliance on handheld controllers. To this end, traditional speech-based methods often depend on rigid command sets, limiting the naturalness and flexibility of interaction. In this study, we propose a novel locomotion technique powered by large language models (LLMs), which allows users to navigate virtual environments using natural language with contextual awareness. We evaluate three locomotion methods: controller-based teleportation, voice-based steering, and our language model-driven approach. Our evaluation combines eye-tracking data analysis, including exploratory explainable machine learning analysis with SHAP, and standardized questionnaires (SUS, IPQ, CSQ-VR, NASA-TLX) to examine user experience through both obj
arXiv:2503.20666v2 Announce Type: replace Abstract: Thematic analysis (TA) is a widely used qualitative approach for uncovering latent meanings in unstructured text data. TA provides valuable insights in healthcare but is resource-intensive. Large Language Models (LLMs) have been introduced to perform TA, yet their applications in high-stakes healthcare settings, particularly for qualitative clinical interview analysis, remain limited. Here, we propose TAMA: A Human-AI Collaborative Thematic Analysis framework using Multi-Agent LLMs for clinical interviews. We leverage the scalability and coherence of multi-agent systems through structured conversations between agents and coordinate the expertise of cardiac experts in TA. Using interview transcripts from parents of children with Anomalous Aortic Origin of a Coronary Artery (AAOCA), a rare congenital heart disease, we demonstrate that TAMA outperforms single-agent LLM TA approaches, achieving higher thematic hit rate, coverage, and dist