EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CL

Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support

arXiv:2608.05151v1 Announce Type: new Abstract: Wastewater operators need answers grounded in how their plant's variables interact and how fast effects propagate, not in generic pretraining text, when asking causal questions such as "why is N2O rising?" or "what happens if I cut aeration by 20%?". We compare three concrete ways to ground a frozen Qwen2.5-32B-Instruct model in an architecturally interpretable wastewater simulator (CCSS-IX): a live simulator oracle (Method 1), structured parameter injection (Method 2), and a Decoupled Recall-Reasoning (DRR) retriever (Method 3). On a 198-question causal benchmark the three reach 99.5%, 79%, and 75.8%, forming a deployment ladder above the strongest retrieval-augmented baseline at 48%. The DRR retriever has 110M parameters and trains per plant in ~17 seconds; after cross-plant transfer to a biologically distinct plant it still reaches 88%, while Method 2's static table cannot transfer. On a 60-question counterfactual benchmark only Method

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

MoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis

arXiv:2605.00025v3 Announce Type: replace-cross Abstract: Speech neuroprosthesis systems decode intended speech from neural activity in the absence of audible output, offering a path to restoring communication for individuals with speech-impairing conditions. Current approaches decode predominantly from motor cortical areas, discarding others -- such as area 44, part of Broca's area -- that may encode complementary linguistic information. We introduce MoDAl (Modality Decorrelation and Alignment), a framework that discovers complementary neural modalities through the interplay of two objectives in a shared projection space. A contrastive loss aligns each of several parallel brain encoders with the text embeddings of a pretrained large language model (LLM), while a decorrelation loss prevents the encoders from coalescing to duplicative representations. We prove that these objectives are in productive tension: Contrastive alignment induces transitive modality coalescence, which decorrelat

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

'OpenBloom': A Stigma-Sensitive LLM Design Probe for Navigating Reproductive Well-being Conversations with Young Adults

arXiv:2606.15536v2 Announce Type: replace Abstract: The growing use of large language models (LLMs) by young adults seeking sensitive health information has raised important questions in Human-AI Interaction about how these systems can support understanding and navigation of reproductive well-being. In response to Feminist HCI principles, we introduce OpenBloom, a web application and an exploratory design probe that uses LLMs to generate question-based prompts from reproductive health articles. Through a user study with 34 young adults across 136 interactions with OpenBloom, we provide an initial assessment of the system while exploring how participants' reflections engage with culture and value sensitivities. We found that while OpenBloom outputs meet expectations of "safe" and non-offensive, they tend to paraphrase or rely on factual recall, which may lead to value dilution. We discuss implications under contestability and value-sensitive frameworks for future LLM-mediated reproducti

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Large Language Model Counterarguments in Older Adults: Cognitive Offloading or Susceptibility to Moral Persuasion?

arXiv:2604.22356v2 Announce Type: replace Abstract: This study examined whether counterarguments generated by large language models (LLMs) influence the moral judgments of younger and older adults, and whether these effects vary by dilemma type, cognitive functioning, trust in AI, and prior LLM experience. Using the switch and footbridge trolley dilemmas, 130 participants (56 younger adults and 74 older adults) were presented with ChatGPT-generated counterarguments that opposed their initial judgments. More than 30% of participants reversed their judgments in both dilemmas (32.31% in the switch dilemma and 36.92% in the footbridge dilemma). Older adults tended to be more likely than younger adults to reverse their judgments and showed a significantly greater degree of judgment change in the switch dilemma. In the emotionally aversive footbridge dilemma, older adults with lower cognitive functioning were significantly more likely to align with the LLM-generated counterargument. General

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Language Scent: Exploring Cross-Language Information Navigation

arXiv:2604.03604v2 Announce Type: replace Abstract: While multilingual users often switch between languages when seeking information, this process remains undersupported by current systems where information is typically siloed by language. Our formative study reveals that users select their search language based on its perceived value for their current information need, a concept we formalize as language scent. Language scent extends Pirolli and Card's information foraging theory - which explains how users navigate among already-encountered sources - to the multilingual case, where users' choice of query language determines which sources they can encounter in the first place. Building on this insight, we designed Niffler, a multilingual information seeking system that provides proximal cues for gauging the language scent of different languages. Finally, we conducted a lab study with 16 multilingual speakers to understand Niffler's utility, usage patterns and application contexts.

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Clinician input steers AI toward accurate and harmful recommendations

arXiv:2603.14158v2 Announce Type: replace Abstract: Large language models (LLMs) are entering clinical workflows, yet evaluations rarely assess how clinician reasoning shapes model behavior during clinical interactions. Using 61 curated NEJM Case Records, we tested how expert or misleading clinician reasoning influenced AI-generated differential diagnoses and next step recommendations across 21 reasoning variants from 8 proprietary and open-source models. After clinician exposure, LLM-clinician concordance increased: simulations with >=3 overlapping differential diagnoses rose from 65.8% to 93.5%, and those with >=3 overlapping next step recommendations from 20.3% to 53.8%. Expert context significantly improved correct final-diagnosis inclusion in all 21 models (mean +20.4 pp), reflecting both improved reasoning and passive content echoing, while adversarial context significantly degraded performance in 14 models (mean -5.4 pp). Expert context also significantly increased leading-diagn

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

A Lexical Analysis of online Reviews on Human-AI Interactions

arXiv:2511.13480v2 Announce Type: replace Abstract: This study focuses on understanding the complex dynamics between humans and AI systems by analyzing user reviews. While previous research has explored various aspects of human-AI interaction, such as user perceptions and ethical considerations, there remains a gap in understanding the specific concerns and challenges users face. By using a lexical approach to analyze 55,968 online reviews from G2.com, Producthunt.com, and Trustpilot.com, this preliminary research aims to analyze human-AI interaction. Initial results from factor analysis reveal key factors influencing these interactions. The study aims to provide deeper insights into these factors through content analysis, contributing to the development of more user-centric AI systems. The findings are expected to enhance our understanding of human-AI interaction and inform future AI technology and user experience improvements.

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

MASS: Multiplayer World Models with Authoritative Shared State

arXiv:2608.06257v1 Announce Type: cross Abstract: Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redundant compute, view inconsistencies, and poor scalability. We propose MAS (Multiplayer world models with Authoritative Shared State) to resolve this limitation. Inspired by multiplayer game architectures, MAS disentangles world dynamics and view rendering. A learned Logic Engine advances a global, authoritative typed state from joint actions without any hand-written transition function, acting as the sole recurrent memory and synchronization reference. From this shared state, a learned Rendering Engine generates independent and consistent views for any requested camera on demand. This explicit disentangling allows MAS to achieve superior state accuracy and lower cross-view inconsistency compared to state-of-the-art multi-view baselines on a matched multiplayer Snake benchmark. It advances p

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation

arXiv:2608.06221v1 Announce Type: cross Abstract: Learning from demonstration (LfD) provides a developmental framework through which robots can develop motor skills by observing and imitating human dynamics, reducing reliance on explicit programming to teach a skill to a robot. The resulting human-like robot motion is recognised as a key factor in building trust and enabling natural collaboration in human-robot interaction. This paper presents a framework for learning human-like robot motion from demonstration, including data collection, probabilistic trajectory learning, and perceptual user evaluation. A dataset of 3,142 handwriting demonstrations was collected from 22 participants across all 52 Latin alphabet character-case combinations via a touchscreen teleoperation interface, capturing planar position, contact force, and timing. Building on the widely used Gaussian Mixture Model and Gaussian Mixture Regression approach for learning from demonstration, the framework is extended in

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators

arXiv:2608.06219v1 Announce Type: cross Abstract: Intuitive teleoperation interfaces are crucial for the safe and effective operation of robotic manipulators in challenging environments. In the nuclear industry, surface contact tasks such as swab sampling require precise path and force tracking, obstacle avoidance, and sustained operator attention, which conventional joystick interfaces struggle to support effectively. This study designs and evaluates a novel touchscreen teleoperation interface that maps continuous finger movements directly to robotic manipulator motions, provides finer velocity control, and integrates control with visualization, enabling more natural, precise, and intuitive surface interaction than conventional controllers. A comparative user study with 20 participants evaluated task performance and workload using the proposed touchscreen, a conventional joystick, and a single-click autonomous mode. Tasks simulated realistic surface manipulation using a Franka Emika P

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India

arXiv:2608.06027v1 Announce Type: cross Abstract: In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write. Reaching them requires a spoken conversation. Today that work falls to frontline health workers who enroll beneficiaries one at a time, a poor use of stretched capacity. We built FormBharo ("fill the form" in Hindi), a voice agent that fills a structured form over a phone call under tight latency and cost budgets by pairing Large Language Models (LLMs) with deterministic, rule-based validation and flow control. It is being piloted with ARMMAN, an NGO running large-scale maternal and child mobile-health programs in India, to enroll low-income, Hindi-speaking mothers in antenatal and postnatal care. To our knowledge, it is the first voice agent piloted to fill an enrollment form for this population. We openly release FormVoiceAgentBench, a benchmark pairing human-recorded Hindi audio with 3,760 multi-tur

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Unified Agent: Managing Interactions across Devices

arXiv:2608.05729v1 Announce Type: cross Abstract: As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time. Yet existing agent systems still fall short in this scenario. This is because observations are scattered across devices and moments, but mainstream systems are not designed around this fact: a single agent that treats devices as tools lacks effective state management for all devices across time, and multi-agent systems coordinate across agents but do not maintain the compact carried state a cross-device, cross-time request needs. We argue that the agent should maintain an effectively designed state that organizes engagement evidence, stated facts, and the standing request in a compact, action-ready form for deciding its action given the current observation. To compare state designs, we construct a benchmark of user-agent interaction across devices and time. We instantiate this principle in Unified Agent, a statef

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows

arXiv:2608.05602v1 Announce Type: cross Abstract: Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or trusted by users. These questions remain necessary, and each can contribute to warranted reliance. However, they do not directly specify warranted reliance as a distinct evaluative target: the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. We argue that this requires an account of epistemic trustworthiness: what makes a system epistemically worthy of reliance. Drawing on philosophical accounts of trustworthiness as competence and audience-orientati

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

PromptShield Home: Ambient Multimodal Prompt Injection Defense for Smart-Home Agents

arXiv:2608.05495v1 Announce Type: cross Abstract: Smart-home assistants increasingly use multimodal large language models (MLLMs) that perceive video and audio directly. This raises a safety question specific to the home: can the agent tell a genuine user command from ambient or externally-sourced content, television speech, on-screen text, or an overheard conversation, that merely looks like a command? We introduce PromptShield-Home, a pilot benchmark of realistic smart-home scenarios spanning addressee ambiguity, screen/audio injection, health-monitor false triggers, mixed occupancy, and a legitimate-command floor, and use it to compare three abstraction layers: traditional detectors (L0), a single MLLM agent (L1; vision, vision+ASR, and audio-visual), and multi-agent mediation (L2; voting, role specialists, cross-model arbitration). Because the label distribution is skewed toward inaction, aggregate accuracy is misleading, a constant always-block predictor scores 82%, so we report u

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers

arXiv:2608.05478v1 Announce Type: cross Abstract: Graphical Abstracts (GAs) visually summarize the key findings of academic papers, playing a crucial role in facilitating the understanding of research content. Recently, advancements in vision-language models and image generation models have enabled the automatic generation of scientific figures based on paper content. However, most conventional methods output the generated results as raster graphics, making post-editing (e.g., text modification and layout changes) highly difficult. This poses a significant challenge, as they are unsuitable for the iterative figure revision process inherent in paper writing and peer review. To tackle these challenges, we define the novel task of generating editable GAs from paper content and propose GenGA, a new GA generation framework that directly produces figures in vector format. By generating figures as a collection of vector elements with a hierarchical structure, GenGA produces outputs that can b

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Failing Gracefully: Mitigating Impact of Inevitable Robot Failures

arXiv:2608.05313v1 Announce Type: cross Abstract: Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failures such as software crashes, hardware degradation, or unpredictable interactions. While roboticists strive to minimize failures, some remain inevitable, making it critical to mitigate their potential consequences for safe and reliable deployment. This paper introduces a novel safety formulation that evaluates both the probability of impactful interactions between robots and surrounding entities during failures, and the severity of their outcomes. By quantifying the impact of failures on different entities, our approach enables robots to make informed planning decisions that balance safety with task efficiency. To support systematic evaluation, we also present FailBench, a MuJoCo-based simulation framework for studying robot-environment interactions under diverse failure modes, including sensing issu

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

arXiv:2608.06202v1 Announce Type: new Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and report accuracy as the primary outcome metric, without accounting for conditions such as web search that may have effects on model behavior in deployment. We audit these assumptions for one of the most widely-used LLMs, comparing two modalities, ChatGPT's chat UI and OpenAI's API, with and without web search enabled. We use a stratified total sample of 401 prompts from two popular benchmarks, BBQ and SafetyBench, collecting 4,812 total responses across three repeated runs per prompt. Beyond standard performance measures, we evaluate model output dimensions including response consistency, response text similarity, citation grounding, and abstention behavior. For instance, chat UI responses were

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Reducing belief in conspiracy theories as they unfold using large language models

arXiv:2608.06151v1 Announce Type: new Abstract: The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational dialogues with a large language model (LLM) can reduce belief in immediately unfolding conspiracies. In experiments conducted in the days following the July 2024 assassination attempt on Donald Trump and the September 2025 assassination of Charlie Kirk, U.S. adults (Experiment 1: N = 472; Experiment 2: N = 1035) holding conspiratorial views about the crisis event engaged in a multi-turn conversation with an LLM prompted to reduce their conspiracy belief. Compared to control participants who either discussed an irrelevant topic with an LLM or viewed a static fact sheet, participants in the LLM treatment showed significantly reduced conspiracy beliefs in both experiments. We also found evidence of downstream effects of the LLM treatment, observing reduced belief in different conspiracies one to two months

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Divergent Perceptuomotor Recalibration in Virtual Reality and Video-Passthrough Mixed Reality on the Same Head-Mounted Display

arXiv:2608.06132v1 Announce Type: new Abstract: Virtual reality (VR) and video-passthrough mixed reality (MR-VPT) can be delivered on the same headset, but it is unclear whether these two interaction modalities produce comparable perceptuomotor behavior. Although delivery through the same headset controls many display-level characteristics, VR and MR-VPT differ in both their visual-feedback pipelines and the action-relevant visual information available to guide movement, such as whether the surrounding environment and the user's body are synthetically rendered or preserved through passthrough. This study compared visually guided manual pointing in VR and MR-VPT using the same headset. Forty adults were assigned to either a VR or MR-VPT group and completed a pointing task in the physical, unmediated reality (UR) before and after performing the same task in their assigned XR modality. The analyses revealed numerous important insights. First, groups had comparable baseline performance in

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

"I don't know anything about laptops!" - User Perception of Digital Product Advisors Adapting to Their Knowledge Levels

arXiv:2608.06091v1 Announce Type: new Abstract: Conversational commerce uses digital assistants to support the search process and decision-making in e-commerce. Effective communication in these interactions can be facilitated by assistants adapting their communication style to users and supporting shared understanding. An open challenge in this context is adapting the presentation of complex product information to users with varying levels of domain knowledge. To investigate strategies for such knowledge-level adaptation, we set up a chatbot-assisted laptop search scenario. In a between-subjects experiment (n = 251), we examined novice and expert perceptions of product attribute recommendations presented as technical information only (T), or augmented with performance categories (TC), attribute explanations (TE), or both (TCE). For novices, approaches with explanations (TE, TCE) were perceived as more helpful and led to higher perceived learning than those without. Novices also rated t

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Cleo: A Transparent and Controllable Chatbot for Conversational Commerce

arXiv:2608.06068v1 Announce Type: new Abstract: We demonstrate Cleo, a transparent and controllable conversational product advisor that addresses the challenges of opacity, unpredictability of LLMs, and the complexity of comparisons in conversational commerce. With our chatbot system, we make four contributions: First, we introduce transparency by prompting the LLM to reflect on interpreted user needs, while an auditable ranking mechanism reveals loss values per attribute, explaining ranking decisions. Second, we propose controllability through a hybrid architecture separating deterministic ranking from language generation. A ranker applies categorical filters and numeric loss functions over 3,638 product specifications. Meanwhile, a constrained LLM generates grounded descriptions constrained to catalog evidence, thus mitigating the risk of hallucinated or persuasive content. Third, we provide decision support in the form of natural-language comparisons and a highlights feature. These

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

arXiv:2608.06013v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonst

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

PoseForge: Editable Pose Analytics for AI-Assisted Sports Coaching

arXiv:2608.05971v1 Announce Type: new Abstract: Athletic coaching increasingly relies on video analysis, yet raw footage lacks tools to quantify motion or simulate valid technique corrections. Drawing on formative interviews with eleven cricket experts (coaches, performance analysts, captains, and players), we introduce PoseForge, a visual analytics system that extracts 3D skeletal poses from single-camera sports videos for interactive movement analysis. In a cricket batting case study, PoseForge computes interpretable kinematic metrics such as feet gap and elbow angle, compares them against scientifically derived norms, and uses an AI coach to suggest targeted adjustments, presented visually and through natural-language feedback (e.g., "increase feet gap by 10 cm"). Users can directly modify poses via mouse interaction or natural-language instructions, with inverse kinematics maintaining anatomical plausibility and real-time updates of metrics and comparisons. An evaluation with the s

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

A Modular Workflow for Multimodal Reading Experiments

arXiv:2608.05966v1 Announce Type: new Abstract: We introduce a web-based modular workflow for real-time multimodal experiments in naturalistic online reading. The workflow integrates eye tracking, EEG, and interaction data from mouse and keyboard, synchronizes them via Lab Streaming Layer, and links gaze to browser-based text at the word, sentence, and AOI levels. It is designed as a reusable experimental procedure that can be adapted to different sensors, tasks, and analysis goals. As a use case, we apply the workflow to a study of selective exposure in online news search and reading. During the experiment, gaze-derived measures are computed online, while EEG and other synchronized streams are processed immediately after task sessions based on fixation-triggered segmentation. The resulting behavioral, neural, and linguistic metrics support selecting text passages for targeted post-task rating or labelling within the same lab session. The workflow thus provides a general basis for mult

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Topic Matters: How Linguistic Properties can Shape Reading Behaviour in Selective Exposure Studies

arXiv:2608.05942v1 Announce Type: new Abstract: Research on selective exposure frequently relies on eye tracking to study reading behaviour, often assuming that texts across different controversial topics are comparable once basic controls are applied. This assumption is problematic if topic-dependent linguistic properties systematically shape how users read and allocate attention. We therefore examine whether such properties relate to differences in reading behaviour in selective exposure contexts. We analyse linguistic features and eye-tracking data from a laboratory study in which 68 participants searched for and read news articles on climate change and migration policy. Our results reveal systematic differences in both textual characteristics and reading behaviour across topics. These findings identify an important methodological confound in selective exposure research and highlight the need to account for topic-specific linguistic properties when interpreting eye-tracking measures

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Temporal Tracking of Reeb-Space Sheets

arXiv:2608.05837v1 Announce Type: new Abstract: Time-varying bivariate fields arise in many scientific applications, where the relationship between two scalar quantities evolves over time. While topological methods such as merge trees provide an effective framework for identifying and tracking features in univariate data, analogous approaches for bivariate fields remain comparatively underexplored. Reeb spaces extend topological analysis to multivariate data by representing fiber connectivity through a collection of interconnected sheets, making these sheets natural candidates for describing bivariate structures. However, establishing temporal correspondences between sheets is challenging due to the structural complexity of Reeb spaces, sensitivity to noise, and the difficulty of defining meaningful similarity measures across timesteps. We present a framework for tracking Reeb space sheets in time-varying bivariate fields. The method establishes correspondences between sheets in consec

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

SpaceVLA: Spatially Grounded VLA for Robotic Manipulation with User-Authored Grasp and Place Anchors

arXiv:2608.05730v1 Announce Type: new Abstract: Vision-language-action (VLA) models follow language commands but often lack explicit spatial intent for manipulation. We present Visual Intent Anchors, an XR pipeline that lets users specify grasp and placement regions and renders them as image-space overlays for VLA control. We collect 200 Unity pick-and-place demonstrations and fine-tune OpenVLA-7B with LoRA on temporally subsampled annotated observations. The policy predicts tokenized 7-DoF incremental actions from marked RGB observations and language. We evaluate the policy in closed-loop Unity trials, achieving a grasp success rate of 91.25% and mean grasp and placement errors of 0.5 cm and 0.7 cm, respectively.

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

ASIDE: From Conflict Participants to Co-Observers Through Dyadic Spectator Reflection

arXiv:2608.05690v1 Announce Type: new Abstract: When people argue over text, they share a record of what was said but may hold different accounts of what it meant. Existing AI reflection tools typically work from one person's account, while dyadic tools support co-expression without making interpretation gaps inspectable. We present ASIDE, a system for Dyadic Spectator Reflection (DSR). DSR follows the sequence externalize independently, then encounter together. From a past chat conflict, ASIDE creates a pixel-art theatrical replay with revisable AI-generated inner-state hypotheses. Partners first confirm hypotheses about themselves and separately edit their interpretations of the other. After both finish, they view the co-annotated scene together and review Divergence Cards that pair a confirmed account with the partner's reading at a specific conversational beat. In an exploratory study with 10 couples who revisited real text-based conflicts, participants described the theatrical rep

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

CaRing: Preventing Carpal Tunnel Syndrome based on Daily Activities from Always-Available Input Device

arXiv:2608.05619v1 Announce Type: new Abstract: We present CaRing, a ring worn on the base knuckle of the index finger, a wearable system for detecting the start and end of mouse use to help prevent Carpal Tunnel Syndrome, in which the damage to the median nerve is permanent. CaRing senses finger movement, which neither a software timer nor a wrist-worn device detects. The displacement reported by an optical flow sensor is accumulated into a running value, then a zero point is measured while the hand rests on the desk at the start of each session. With this formulation, the start and end thresholds are expressed relative to the session's zero point. CaRing does not introduce any per-user parameter. We empirically demonstrate that approximately $90\%$ of start and end events are detected within two seconds of the researcher's label, using 35 recordings and a lab study with ten users.

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Toward Resilient Human-AI Collaboration: A Lifecycle Taxonomy of Sociotechnical Risks and Cascading Failures

arXiv:2608.05614v1 Announce Type: new Abstract: As AI systems become increasingly integrated into consequential domains such as healthcare, journalism, education, scientific research, organizational decision-making, and defense, effective human-AI collaboration has emerged as a critical challenge. However, the sociotechnical risks that undermine collaboration are often studied in isolation, obscuring the recurring failure mechanisms that cut across domains. This paper presents a lifecycle-oriented synthesis of human-AI collaboration risks spanning four stages: task allocation, interaction, feedback, and adoption. Drawing on evidence from diverse application domains, we identify six recurring cross-domain risk clusters: Trust Miscalibration, Cognitive Burden, Accountability Gap, Capability Erosion, Goal Misalignment, and AI Anxiety and Technostress. We further propose a conceptual interaction model that illustrates how these risks emerge from sociotechnical drivers, interact through cas

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

A Multi-Layer System for Ultra-High-Resolution Static 360-Degree Telepresence

arXiv:2608.05570v1 Announce Type: new Abstract: 360-degree video telepresence offers strong immersive potential but remains constrained by the limited resolution of current capture and display hardware. Many telepresence installations feature fixed viewpoints and largely static scenes, yet optimization strategies tailored to such setups have received limited attention. We present a multi-layer, ultra-high-resolution system for static 360-degree telepresence that combines an 8K panoramic camera with a rotatable 4K pan-tilt-zoom (PTZ) camera. Our approach builds a three-layer representation: (1) a tile-based ultra-high-resolution panoramic background, generated by offline stitching high-detail 4K PTZ scans onto the base 8K panorama to achieve effective resolution beyond native capture, and represented as a set of spatial tiles; (2) a dynamic update layer that composites foreground motions from the 8K stream via real-time high-resolution background matting; and (3) a region-of-interest 4K

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Turing's Frist Imitation Game: Design Concepts and a Human-Approximates-Machine Reading

arXiv:2608.05558v1 Announce Type: new Abstract: This paper examines Turing's 1948 report, "Intelligent Machinery", as an important conceptual source for the later imitation games. Its first contribution is to identify and integrate the design concepts underlying the 1948 chess-based imitation game: the possibility that intelligent machines may make mistakes, the exclusion of irrelevant physical features, the role of the human judge, and Turing's claim that intellectual activity consists mainly of search. The paper's second contribution is to argue that restricting the human contestant to a rather poor chess player increases the role of intellectual search and makes human behaviour more comparable to machine behaviour. This interpretation presents the 1948 game as a human-approximates-machine game and suggests that the imitation game framework can be used not only to ask whether machines imitate humans, but also to examine when human intelligence becomes machine-like under specific task

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.HC

Mixed Uncertainty in One View: Co-Visualizing Statistical Variability and Qualitative Confidence

arXiv:2608.05487v1 Announce Type: new Abstract: Forecasting involves multiple forms of uncertainty, including both uncertainties that can be quantified directly (quantitative uncertainty) and those that must be expressed through experts' subjective judgments about the forecast and its context (qualitative confidence). Past work has established that conveying both quantitative uncertainty and qualitative confidence in forecasts can alter readers' decision making, but little research investigates the impact of how these forms of uncertainty are presented. In this work, we present three preregistered human-subjects studies (total n = 923) on how different methods of visualizing qualitative uncertainty alongside line charts' confidence intervals affects non-experts' decision making. In particular, we investigate representing qualitative uncertainty separately via text and icons, and integrated into quantitative confidence intervals via color, transparency, and a blurred stroke design. In E

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Gender-Based Heterogeneity in Youth Privacy-Protective Behavior for Smart Voice Assistants: Evidence from Multigroup PLS-SEM

arXiv:2603.27117v2 Announce Type: replace-cross Abstract: This paper investigates how gender shapes privacy decision-making in youth smart voice assistant (SVA) ecosystems. Using survey data from 469 Canadian youths aged 16-24, we apply multigroup Partial Least Squares Structural Equation Modeling to compare males (N=241) and females (N=174) (total N = 415) across five privacy constructs: Perceived Privacy Risks (PPR), Perceived Privacy Benefits (PPBf), Algorithmic Transparency and Trust (ATT), Privacy Self-Efficacy (PSE), and Privacy Protective Behavior (PPB). Results provide exploratory evidence of gender heterogeneity in selected pathways. The direct effect of PPR on PPB is stronger for males (Male: \b{eta} = 0.424; Female: \b{eta} = 0.233; p < 0.1), while the indirect effect of ATT on PPB via PSE is stronger for females (Female: \b{eta} = 0.229; Male: \b{eta} = 0.132; p < 0.1). Descriptive analysis of non-binary (N=15) and prefer-not-to-say participants (N=39) shows lower trust and

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations

arXiv:2604.17359v2 Announce Type: replace Abstract: Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no real one. We gave GPT-4o-mini, Gemini-3-Flash, DeepSeek-V3 and GLM-4.7 each of 120 demographic cohorts under two framings, one written as a clinician enters a patient and one as a person describes themselves, and scored all 28,800 responses against survey-weighted PHQ-8 anchors derived from NHANES microdata. Case by case the output holds up: 97.3% of elevated presentations satisfy the DSM-5 gateway rule, violating it at 2.68% against a chance null of 10.4%. As populations, four things fail at once. Every benchmarkable group returns inflated by 2.8 to 5.5 PHQ-8 points, and 18.2% of simulated patients screen at the treatment threshold against 7.5% of adults. Population Black-White and Hispanic-White disparities do not survive the simulation, with two models attenuating each gap and two flattening or in

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data

arXiv:2603.00059v3 Announce Type: replace Abstract: How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article presents a comparison between a human-respondent survey of 420 Silicon Valley coders and developers and synthetic survey data designed to simulate real survey takers generated by five leading Generative AI Large Language Models: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2. Our findings reveal that while AI agents produced technically plausible results that lean more towards replicability and harmonization than assumed, none were able to capture the counterintuitive insights that made the human survey valuable. Moreover, deviations grouped together for all models, leaving the real data as the outlier. Our key findi

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Geo-Standardizing 3D Modeling of Surface/Subsurface Objects and Related Logical Spaces on Celestial Bodies: Case Studies for Moon and Mars

arXiv:2601.06182v2 Announce Type: replace Abstract: Establishing frameworks for promoting the realization of various activities on celestial bodies sustainably is of great significance for different contexts, such as preserving the scientific evidence and space heritage. Therefore, this research first proposes a conceptual model that covers the different types of features, attributes, and relationships between them to comprehensively delineate the surface/subsurface objects and related logical spaces on celestial bodies. It then implements this conceptual model as a CityJSON extension in such a way that allows for creating the three-dimensional (3D) geodatasets that represent these objects and spaces in a standardized manner. Moreover, the usefulness of this study is demonstrated through creating CityJSON datasets that include 3D models of exemplary surface/subsurface objects from the Moon and Mars, such as a historical landing site and related logical spaces, such as exclusion zones f

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Auditing Sex/Gender Disparities in Emergency Triage with LLM-based Paired Comparisons

arXiv:2511.17124v2 Announce Type: replace Abstract: We present a domain-agnostic paired-comparison approach that uses Large Language Models (LLMs) to quantify sex/gender-related asymmetries in documented clinical decision-making. The method trains an LLM to emulate observed decisions, then evaluates sex-swapped pairs in which only sex is flipped, holding documented clinical content constant. We apply it to emergency triage, analyzing more than 140,000 Bordeaux University Hospital (France) admissions and testing methodological portability on MIMIC-IV, spanning a different language, population, and healthcare system. Fine-tuning Mistral NeMo 12B for triage prediction and using Mistral Small 24B for pair generation, we find otherwise identical presentations were more likely to receive a lower-severity predicted score as female than male: 1.1% (95% CI 0.9-1.3) in the French cohort, 2.2% (1.7-2.7) in MIMIC-IV. Predictions are sensitive to both tabular and textual sex markers, with the asymm

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Scientific Discovery in the Age of AI and Supercomputing

arXiv:2511.12686v2 Announce Type: replace Abstract: Artificial intelligence (AI) and high-performance computing (HPC) are transforming scientific capabilities and the way science is conducted. Yet their combined impact on scientific discovery remains poorly understood, as do inequalities in access to these capabilities across countries and institutions. Drawing on metadata from more than five million scientific publications (2000-2024) across 27 fields, we examine how the convergence of AI and HPC correlates with scientific breakthroughs. Our results show that this computational synergy is most pronounced at the scientific frontier: research combining AI and HPC is more likely to introduce novel ideas and achieve top-cited status than either conventional work or research using AI or HPC in isolation. We also document growing disparities in access to supercomputing resources and AI expertise, which are increasingly concentrated in a small number of regions (dominated by the United State

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

arXiv:2608.06141v1 Announce Type: cross Abstract: This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and non-standard language varieties are not only technical errors, but also implicit linguistic policies that reproduce colonial language hierarchies. Drawing on linguistic capital, raciolinguistic ideology, language policy research, and decolonial computing, we show how data, metrics, and model priors determine whose voices become machine-legible. We introduce the Three Harms (3M) taxonomy---Misrecognition, Misalignment, and Mistrust---and a seven-layer situatedness model for linguistic diversity in ASR and ASR-mediated voice interfaces. We then propose a participatory framework and minimum audit protocol for culturally competent ASR, positioning affected communities as co-designers, evaluators, and governance partners.

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

CourseGraph: Finding overlaps and differences in Computer Science courses across universities

arXiv:2608.05910v1 Announce Type: cross Abstract: Student mobility programs such as Erasmus+ enable students to take courses at other universities, broadening their academic and cultural horizons. However, this flexibility also leads to a practical challenge: ensuring that students do not take courses elsewhere that substantially overlap with courses in their home curriculum. In this work, we propose CourseGraph, a methodology that automates the evaluation of external courses based on insights obtained from the process followed by curriculum administrators when assessing courses for inclusion in a degree program. Course- Graph extracts information such as course titles, descriptions, and learning outcomes from the course webpage. Then, this information is represented semantically using a BERT-based language model, after which the pair-wise similarity between courses can be computed. This information is then used by a Random Forest classifier to determine whether a candidate course abro

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Mapping the Emerging Curriculum for AI-Assisted Software Engineering via Syllabus Analysis

arXiv:2608.05898v1 Announce Type: cross Abstract: As Generative AI coding tools reshape professional software development, universities have begun designing courses to prepare students for AI-assisted development workflows. By analyzing the syllabi of these courses, we can gather empirical evidence about these courses, reveal how this emerging curricular area is being defined, and gain guidance for future curriculum design. We analyzed 23 publicly available syllabi and course materials of upper-division, credit-bearing courses that meet specific criteria, including explicitly addressing Generative AI in software engineering. Through iterative qualitative coding, we characterized courses' learning objectives, assessments, topics, and documented AI tools. Our analysis reveals commonalities and differences among these courses that allow researchers and educators to study and develop future courses.

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025

arXiv:2608.05889v1 Announce Type: cross Abstract: Large language models (LLMs) can leave small stylistic traces in text written with their help. The most discussed is the em-dash (U+2014), especially the unspaced form word---word, which is normal in typeset English prose but unusual in U.S. press writing, where AP style calls for spaced dashes. This study asks whether that trace is measurable in congressional press releases. In a preregistered design (OSF: 10.17605/OSF.IO/U5NEY), 146,239 scraper-sourced releases from 480 House and Senate offices (2021-2025, the open congress-press dataset) were analyzed: density of unspaced prose-form em-dashes per 1,000 characters of cleaned text, Poisson/negative-binomial models with a length offset, clustering by office. Density stayed within 0.10-0.12 per 1,000 characters through 2021-2024, then rose to 0.217 in 2025, more than twice the four-year baseline; the share of releases with such an em-dash rose from ~13% to 24.8%. The primary frequency ra

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

arXiv:2608.05576v1 Announce Type: cross Abstract: When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines. This gap between LLM and human writing has been noted across a variety of domains. LLMs tend to produce "average" writing, while human writing contains more diverse content that covers a broader distribution. Existing work has shown the existence of this distributional "gap", but no work has proposed a systematic way to measure it. Our paper proposes a human-grounded framework that uses the empirical distribution of human writing on a topic to measure the distributional breadth of LLM-generated content on that same topic. We propose two metrics, LLM Coverage (LLM-Cov) and In-Boundar

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Small Foundation Models of Human Cognition and Behaviour

arXiv:2608.05224v1 Announce Type: cross Abstract: Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. In-distribution, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choic

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning

arXiv:2608.05166v1 Announce Type: cross Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles the effect of exposure to a biased user turn from the effect of the turn's semantic content, alongside a benchmark of 24,300 jury-validated user prompts spanning all 81 cells of a 9x9 target-human bias interaction matrix. Across eight frontier LLMs, we find that biased conversational context systematically increases bias expression relative to zero-shot baselines in 6 of 8 models. We identify two competing behavioral dynamics underlying this effect: conversational exposure to biased reasoning generally amplifies downstream bias tendencies, while explicitly stated bias cues often trigger alignment-related suppression behaviors that reduce overt bias expression. We release our framework, codebase, and dataset to

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria

arXiv:2608.06364v1 Announce Type: new Abstract: The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including fraud and reduced user control over digital technologies, raising concerns about digital sovereignty. This research examines how Artificial Intelligence (AI) in Nigerian mobile applications affects digital sovereignty, examined through platform transparency as a key indicator of user awareness and control. Using an interpretive approach, the research combines the forensic analysis of selected Android applications with contextual document analysis to identify AI features and evaluate disclosure practices. The findings show that AI is widely implemented in the applications, yet transparency about its use remains limited. A socio-economic analysis of Nigeria further shows an increasing dependence on consumer digital platforms, moderate AI awareness, and uneven patterns of interaction. By providing empirical evidence on AI trans

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Precision Medicine to Precision Education: A Vision for AI-Powered Student Digital Twins, Preventive Student Success, and Career-Aligned Academic Pathways

arXiv:2608.06322v1 Announce Type: new Abstract: Higher education remains largely reactive in its approach to student success. Institutions frequently identify academic problems only after students have failed courses, fallen behind in degree progression, accumulated excessive debt, or departed without a credential. Healthcare faced a similar challenge decades ago. It responded by shifting from reactive treatment to preventive care powered by predictive models, risk stratification, electronic health records, and artificial intelligence (AI). This paper argues that higher education stands at an analogous inflection point. Drawing on advances in learning analytics, educational data mining, machine learning, workforce analytics, and digital twin technologies, we propose a paradigm we call Precision Education. Under this framework, AI continuously analyzes academic, behavioral, financial, and career data to identify emerging risks, recommend personalized interventions, optimize educational

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries

arXiv:2608.06166v1 Announce Type: new Abstract: The article reports on a blind Turing Test experiment, assessing the performance of out-of-the-box leading LLMs on three Italian legal professional exams: the Bar, Judges and Notary exams. Leading LLMs were asked to generate full written exam papers, which were made indistinguishable from human submissions and anonymously evaluated by expert examiners, using the same criteria applied in real examinations. Results reveal marked differences across both models and tasks. While some LLMs match or exceed top human performance in adversarial legal argumentation and doctrinal analysis, all models fail in the notary exam, which requires goal-directed legal planning under strict formal and substantive constraints. Beyond ranking models, the study identifies task-specific strengths, limitations and recurring legal failure patterns. Although limited to out-of-the-box systems, the findings provide qualitative evidence on the current scope and boundar

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Algorithmic Flattening of Sound: Computational Evidence and Justice Implications of AI Music Homogenization

arXiv:2608.06106v1 Announce Type: new Abstract: This paper audits whether large-scale generative music systems exhibit measurable musical homogenization relative to human-produced music, and develops a justice-centered account of why this matters. We audit two commercially deployed systems (Suno and Lyria 3) across four genres (Afrobeats, K-pop, Dance Pop, and Heavy Metal). For each system and genre, we generate 100 tracks and compare them against human corpora of equal size, using 72 music information retrieval (MIR) features and multiple diagnostics of dispersion, redundancy, and separability. We define homogenization as reduced acoustic variation in standard computational audio features including rhythm and timing, timbre/spectral shape, and dynamics, both within genres and across genre boundaries. We also generate tracks using only a genre name as the prompt, with no additional instructions, to reveal each system's default musical tendencies. The results show two structurally disti

Source ↗
Showing 10501–10550 of 11035 signals
← Prev Page 211 of 221 Next →