EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation

arXiv:2604.20468v3 Announce Type: replace-cross Abstract: Industrial robot applications require increasingly flexible systems that non-expert users can easily adapt for varying tasks and environments. However, different adaptations benefit from different interaction modalities. We present an interactive framework that enables robot skill adaptation through three complementary modalities: kinesthetic touch for precise spatial corrections, natural language for high-level semantic modifications, and a graphical web interface for visualizing geometric relations and trajectories, inspecting and adjusting parameters, and editing via-points by drag-and-drop. The framework integrates five components: energy-based human-intention detection, a tool-based LLM architecture (where the LLM selects and parameterizes predefined functions rather than generating code) for safe natural language adaptation, Kernelized Movement Primitives (KMPs) for motion encoding, probabilistic Virtual Fixtures for guide

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling

arXiv:2507.02950v4 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly evaluate generated dialogue, but repeatable scores do not necessarily align with professional judgment. This observational fixed-benchmark study compared four configured LLM evaluator systems (GPT-5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Fable 5) with aggregated ratings from 15 counseling experts on 18 complete simulated AI-to-AI counseling sessions conducted in Japanese. The sessions represented three counselor conditions across six prespecified client profiles. Each system scored every transcript three times on four motivational interviewing-informed dimensions and overall quality. All four systems assigned higher scores than the expert panel for softening sustain talk and overall quality, although differences varied across systems and constructs. Single-run intraclass correlation coefficients ranged from .33 to .96, showing that high run-to-run reliability did not ensure closer exp

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

The BS-meter: Detecting Politics and Labour through ChatGPT's Language

arXiv:2411.15129v3 Announce Type: replace-cross Abstract: What can we learn about language from studying how it is used by ChatGPT and other large language model (LLM)-based chatbots? In this paper, we analyse the distinctive character of language generated by ChatGPT, in relation to questions raised by natural language processing pioneer, and student of Wittgenstein, Margaret Masterman. Following frequent complaints that LLM-based chatbots produce "bullshit," in the sense of Frankfurt's popular monograph On Bullshit, we conduct an empirical study to contrast the language of 1,000 scientific publications with typical text generated by ChatGPT. We then explore whether the same language features can be detected in two well-known contexts of social dysfunction: George Orwell's critique of political speech, and David Graeber's characterisation of bullshit jobs. Using simple hypothesis-testing methods, we demonstrate that a statistical model of bullshit can reliably relate the Frankfurtian

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Algorithmic Fairness Perceptions in the Global South: Evidence from Bangladesh on Ride-Sharing, Beauty Filters, and Large

arXiv:2508.05281v2 Announce Type: replace Abstract: Algorithmic fairness research comes almost entirely out of North America and Western Europe, so we know little about how people elsewhere judge the algorithms they already rely on every day. We asked people in Bangladesh directly: a bilingual (Bangla and English) survey of 199 participants rated fairness across three everyday scenarios -- ride-sharing prices that shift with context, AI beauty filters that reshape appearance, and large language models that handle cultural values differently than a human would. Four patterns stood out. Context changes the verdict even when the outcome doesn't: a 20% price surge during a medical emergency feels less fair than the identical surge on a casual trip (2.00 vs. 2.17 on a 5-point scale, Wilcoxon p = .006), a small effect uneven across income groups (largest among middle-income participants). People already view surge pricing critically in general; context sharpens the judgment rather than creat

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Toward a New Science of AI as Cognitive Infrastructure

arXiv:2507.22893v3 Announce Type: replace Abstract: Contemporary human-AI interaction research overlooks how AI systems fundamentally reshape human cognition pre-consciously, a critical blind spot for understanding distributed cognition. This paper introduces "Cognitive Infrastructure Studies" (CIS) as a new interdisciplinary domain to reconceptualize AI as "cognitive infrastructures": foundational, often invisible systems conditioning what is knowable and actionable in digital societies. These semantic infrastructures transport meaning, operate through anticipatory personalization, and exhibit adaptive invisibility, making their influence difficult to detect. Critically, they automate "relevance judgment," shifting the "locus of epistemic agency" to non-human systems. Through narrative scenarios spanning individual (cognitive dependency), collective (democratic deliberation), and societal (governance) scales, we describe how cognitive infrastructures reshape human cognition, public re

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Communication styles and reader preferences of LLM- and human-authored COVID-19 information explanations: a case study

arXiv:2505.08143v2 Announce Type: replace Abstract: With the wide adoption of large language models (LLMs) in information assistance, it is essential to examine their alignment with human communication styles and values. We situate this study within health fact-checking, where effective communication is critical for correcting misconceptions and building trust. Although recent studies have explored LLMs for fact-checking and health communication, differences between LLM and human communication styles and associated reader perceptions remain under-explored. We compiled a dataset of 1,498 health misinformation claims and explanations from authoritative fact-checking organizations and generated LLM responses to inaccurate health information. Drawing on health communication theories, we evaluated communication styles across three dimensions: information linguistic features, sender persuasive strategies, and receiver value alignments. We further assessed reader perceptions through a blinded

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Losing One's Story: How Vulnerable Users Experience Harm in Online Support Seeking

arXiv:2311.15427v2 Announce Type: replace Abstract: Online support communities are a critical resource for individuals facing distress, stigma, and limited offline support. However, these spaces are not uniformly supportive, particularly for users with little margin for error. In this paper, we examine how vulnerable users experience harm within online support-seeking interactions. Through 25 semi-structured interviews with Reddit users, we show that support seeking is often a constrained practice shaped by structural vulnerability. We introduce the concept of \emph{narrative harm} to describe how participants experience harm through losses of narrative authority, coherence, and space. Their personal disclosures are questioned, reframed, or displaced by others, reflecting asymmetries in voice, credibility, and interpretive power embedded in platform dynamics and moderation regimes. In response, users engage in defensive strategies, including selective self-disclosure, self-silencing, i

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Relative Attention-based One-Class Adversarial Autoencoder for Continuous Authentication of Smartphone Users

arXiv:2210.16819v4 Announce Type: replace Abstract: Behavioral biometrics-based continuous authentication is a promising authentication scheme, which uses behavioral biometrics recorded by built-in sensors to authenticate smartphone users throughout the session. However, current continuous authentication methods suffer some limitations: 1) behavioral biometrics from impostors are needed to train continuous authentication models. Since the distribution of negative samples from diverse attackers are unknown, it is a difficult problem to solve in real-world scenarios; 2) most deep learning-based continuous authentication methods need to train two models to improve authentication performance. A deep learning model for deep feature extraction, and a machine learning-based classifier for classification; 3) weak capability of capturing users' behavioral patterns leads to poor authentication performance. To solve these issues, we propose a relative attention-based one-class adversarial autoenc

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Toward an Actionable Socioeconomic-Aware HCI

arXiv:2108.13477v3 Announce Type: replace Abstract: Although inequities for individuals in different socioeconomic situations are starting to capture widespread attention, less attention has been given to the socioeconomic inequities that saturate socioeconomic-diverse individuals' user experiences. To enable HCI practitioners to attend to such inequities and avoid unwittingly introducing them, in this paper we consider a wide body of research relevant to how an individual's socioeconomic status (SES) can affect their user experiences with technology. We synthesize this foundational research to produce a core set of 6 evidence-based SES "facets" (attribute types and value ranges) that directly relate to user experiences for individuals in different SES strata. We then harness these SES facets to produce actionable paths forward -- including a new structured method we call SocioeconomicMag -- by which HCI researchers and practitioners can bring new socioeconomic-aware practices into the

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Comparative Evaluation of 3D Reconstruction Methods for Immersive Visualization of Laboratory Objects

arXiv:2608.27301v1 Announce Type: cross Abstract: In this study, we examined whether current 3D reconstruction methods can support the creation of realistic holographic representations of laboratory objects for educational use. In this regard, we compared four approaches: photogrammetry, a neural radiance field (NeRF)-based method, Gaussian splatting, and LiDAR. These methods were used to generate holographic models of common laboratory items and their fidelity was evaluated by graduate students. Participants assessed the models for shape, color, texture, and visual defects using a repeated-measures design. Across objects, the NeRF-based method produced the most consistently high-fidelity representations, particularly for transparent, reflective, or low-texture items that were difficult to capture with other approaches. Shape and color were generally reproduced more successfully than texture, suggesting that some visual properties remain more challenging to represent accurately in educ

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

BrailleBench: Investigating Multi-Criteria Braille Comprehension in Large Language Models

arXiv:2608.27268v1 Announce Type: cross Abstract: Although Large language models (LLMs) mediate access to knowledge and computational assistance, their capabilities should benefit vulnerable groups in the same way. However, it is unclear whether existing AI systems are inclusive enough for blind and deafblind users to access the same functionality through Braille, whose indicators, contractions, and digital representations introduce distinct requirements for model comprehension. To this end, we introduce BrailleBench, a benchmark for evaluating LLMs in Braille comprehension from different Criteria. BrailleBench aligns 5,570 instances from five datasets, including mathematics, commonsense, and multi-hop question answering across English and Braille Grades 1 and 2. Different configurations are designed to understand whether the systems can comprehend Braille-authored content, express answers in Braille, and complete end-to-end Braille interaction. To ensure the quality and prevent evalua

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Soft Active Electromyography Interface for Machine Learning-Enabled Silent Speech Recognition

arXiv:2608.27048v1 Announce Type: cross Abstract: Silent speech recognition (SSR) provides an alternative communication pathway in the absence of audible speech. However, conventional approaches are limited by the need for constant facial attachment, privacy concerns, and unstable signal acquisition. Here, we propose a soft, active electromyography (EMG) interface that enables word-level SSR using machine learning. Worn on the hand, the device uses a fingertip electrode that can be positioned near the lips to acquire EMG signals only when needed. The interface integrates liquid metal (LM) interconnects, transparent flexible printed circuit (FPC) electrodes, and elastomer encapsulation to ensure high mechanical stability during finger motion. A deep neural network trained on these stable signals achieved a mean accuracy of 97.2 $\pm$ 1.3% across three subjects in classifying a 30-word vocabulary, demonstrating robust linguistic discrimination. Furthermore, real-time drone control valida

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency

arXiv:2608.26885v1 Announce Type: cross Abstract: Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review. Methods. After a conservative title-only screen, 1,131 records were screened by one review lead, four trained assistants screening non-overlapping subsets, and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, agreement, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational. Results. No workflow recovered all verified eligible records. The human

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Processing/p5 Defined through Practice and Learning

arXiv:2608.26614v1 Announce Type: cross Abstract: Processing/p5 libraries across different programming languages enact consistent priorities for creative coding as a designed experience. While different programming language ecosystems, like Java and JavaScript, are each associated with their own affordances, community norms, and patterns of use, Processing/p5 sketches across these languages share similarities. Based on case studies of building an implementation of Processing/p5 in two host languages, JavaScript and Lua, we propose a list of software decision-making guiding aspects that constitute Processing/p5, regardless of host language. We discuss this framework in the context of decisions in other exploratory and creative tools that demonstrate how each of the guiding aspects can be operationalized differently than in the case studies. The proposed list highlights opportunities for learning, research, and artistic practice through creation of new Processing/p5 libraries for creativ

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Direct Manipulation and Natural Language Programming, Together at Last?

arXiv:2608.26359v1 Announce Type: cross Abstract: Decades of programming languages research has contributed novel approaches to program editing that go beyond modifying text, including direct manipulation programming, structure editing, and automated refactoring tools. However, the rapid growth of natural language programming largely reinforces a view of programs as text and program editing as (unstructured) text transformation. How can we develop unified programming systems that bridge the gap between these approaches, supporting multiple editing paradigms in concert? We take a first step toward answering these questions by introducing a framework that enables program editing via both direct manipulation and natural language, and instantiate this framework in a variant of the $\texttt{cartokit}$ direct manipulation programming system. Our key insight is to treat programs as sequences of structured edits and to use an edit language as a shared interface for both direct manipulation and

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

"A Second Set of Eyes": The Process and Challenges of Software Documentation Review

arXiv:2608.26232v1 Announce Type: cross Abstract: Organizations assign documentation work to technical writers, yet the knowledge required to produce it is distributed across developers, managers, and other practitioners. Prior work has established quality criteria for judging "good" documentation, but it has not examined how practitioners bring that expertise to improve documentation quality or the challenges they face in doing so. Through semi-structured interviews with experienced technical writers ($n=31$) from different organizations, our work reveals the individual and collaborative effort required to maintain documentation quality. We identify five distinct stages of the documentation review process: self review, technical review, editorial review, play testing, and post-publication feedback. Each stage draws on practitioners with distinct expertise to address quality across content, presentation, and user experience. Our findings surface organizational and technical challenges

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion

arXiv:2608.26185v1 Announce Type: cross Abstract: Equal participation in co-located discussion is important for effective collaboration, yet people often hold back when they anticipate negative interpersonal or professional consequences, especially when raising a point requires voicing it themselves. We present SecondVoice, a mixed-reality system that enables people to speak up through an embodied virtual proxy. By separating what is said from who says it, SecondVoice brings hesitant points into the live spoken discussion without putting the speaker on the spot. Using a private overlay, users specify their intent through a structured specification process rather than composing a full utterance. The system reformulates the input and voices it into the conversation through the proxy. We characterize a design space of participation channels under social risk. In a preliminary within-subject study (N = 16), we compare the complete SecondVoice system with an anonymous text-board channel acr

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education

arXiv:2608.26184v1 Announce Type: cross Abstract: AI programming tutors provide scalable support, yet lack the behavioral context human tutors rely on to adapt support to learners' needs. We present TutorTrace, a dataset and behavioral abstraction pipeline that makes learners' behavioral context visible and computable in real time from low-level IDE telemetry. Across four deployments in two introductory Python courses (N=480), TutorTrace captures approximately 180K telemetry events, 13,633 behavioral segments, and 27 continuously computed metrics. From this foundation, we derive a taxonomy of learner activity before the first AI query, between consecutive queries, and across the full session, enabling systems to respond not just to what learners say, but to what they have done leading up to the help-seeking moment. In a preliminary classroom evaluation, behavior-aware prompts were associated with a decrease in intervals between queries with no independent work from 50.0% to 20.7%. As a

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Persona in LLM-Based HRI

arXiv:2608.26182v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for verbal interaction in social robots, yet prompt design in human-robot interaction (HRI) remains underspecified. As a result, robots may present hallucinated capabilities, unclear behavioural boundaries, and misleading personas. This paper develops a framework for prompt design in LLM-based robots and introduces a structured prompt template comprising eight functional components through which robot behaviour can be specified, bounded, and adapted. The framework is grounded in a review of prior LLM-based HRI work and complemented by survey and discussion data from HRI experts gathered at the Robo-Identity workshop at IEEE RO-MAN 2025 (N=27). The qualitative findings highlight limited legibility of robot personality, the need for user adaptation, and strong ethical concerns about safety, deception, and governance. Based on these findings, we present prompting guidelines accompanied by

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

arXiv:2608.26163v1 Announce Type: cross Abstract: Cough events during live spoken conversations carry clinically valuable respiratory signals, yet existing dialogue systems treat them as acoustic noise to be discarded. We present HealthCUES (Clinical Understanding from Embodied Sounds), a streaming pipeline for paralinguistic respiratory monitoring in real-time conversational agents, a capability that, to the best of our knowledge, is absent from all prior systems. HealthCUES processes audio through a rolling buffer aligned with dialogue turn boundaries, enabling sub-second event detection without interrupting conversational flow. Beyond binary cough detection, the system provides fine-grained analytics: (i) differentiation between coughing and throat clearing, (ii) cough subtype classification (dry, wet, barking, whooping) with confidence scores, and (iii) temporal duration estimation with start-end boundaries. To prevent alert fatigue, HealthCUES introduces dialogue-aware gating mech

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Search

arXiv:2608.26152v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models have reduced, or even eliminated, the need for human input. But rather than replacing human cognitive effort, LLMs may instead serve as cognitive tools to extend human abilities, particularly when they are engaged in a task requiring open-ended conceptual exploration and creative ideation. However, we are yet to understand how these models may enhance such generative human cognitive abilities in human--AI interactions. In this study, we explore and evaluate the ability of LLMs to follow and enhance human mental trajectories during semantic memory search. To test this, we use the semantic fluency task (SFT), a classic cognitive paradigm requiring generative semantic memory retrieval that has long served to cha

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

arXiv:2608.26145v1 Announce Type: cross Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in supporting literature review writing. Twenty AI-generated literature reviews based on research sources from Semantic Scholar and Arxiv were evaluated by two researchers across 15 dimensions. Our findings reveal that AI-generated literature reviews require human oversight to meet academic publishing standards. As context windows increase, LLMs can incorporate broader information and maintain coherence across longer inputs, but they also exacerbate issues such as content repetition, omission of critical work, and a tendency towards descriptiveness over synthesis. Our work shows that AI-generated reviews can provide foundational overviews, but their output must be critically evaluated and refined

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach?

arXiv:2608.27443v1 Announce Type: new Abstract: AI agents are poised to become a primary interface to digital products, acting across email, files, payments, and personal data. People without professional software backgrounds need understandable, reusable ways to control actions across services. We examine a mechanism in which a language model maps actions to plain-language consequence categories with user-authored "allow", "ask", or "never" rules. We ask what is gained and lost when decisions are made in advance as reusable rules rather than separately for each action. We analyzed 113 participants without professional software backgrounds across three conditions: per-action human-in-the-loop approval (HITL), automated per-action model review (AUTO), or user-authored consequence policy (POLICY). Participants judged 2 examples in each of 4 consequence categories; POLICY participants then set one rule per category. All supervised an 18-action simulated day, including 7 overreach actions.

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Beyond Harassment: Exploring the Harm Experienced by People with Disabilities in Social Virtual Reality

arXiv:2608.27390v1 Announce Type: new Abstract: People with disabilities (PWD) are increasingly engaging in social virtual reality (VR) platforms, where immersive and embodied interactions can intensify negative experiences. While prior work has examined harassment in VR, little is known about the harms experienced by PWD and the perceived severity associated with different harassment and disability types. Unlike harassment, which represents behaviors, harm is more critical to designing effective protections, as it reflects the consequences and impact; the realism of VR and the vulnerability resulting from disability identity can further amplify such impact. To characterize and model harms for PWD, we conducted a literature review, followed by an online survey with 67 PWD to understand participants' harassment experiences and resulting harms in social VR. We identified 19 types of harm in 5 categories, and reported the severity perception of each type of harm. Finally, we analyzed our

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

A Point-of-Prescription Safety-Check System for Adverse Drug Reactions in Rural Bangladeshi Hospitals: A Feasibility Study

arXiv:2608.27239v1 Announce Type: new Abstract: Adverse drug reactions (ADRs) are a major, largely preventable source of patient harm. In high-income settings, electronic health records store a patient's allergy history and warn prescribers when a contraindicated drug is ordered; in rural Bangladeshi public hospitals no such record exists for outgoing patients, a single physician may see on the order of one patient per minute, and a patient's history of severe reactions does not survive between visits. This paper proposes and outlines the evaluation of a lightweight, smartphone-based safety-check system for this setting. At registration a soft identifier (a phone number) is recorded; after the physician writes a prescription, its image is captured, the brand names are resolved to active ingredients using national drug references, and the ingredients are matched against the patient's recorded severe reaction history. The system is retrieval-based rather than predictive, and is silent by

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Surrounded by Friends: Design and Evaluation of Immersive Layouts of Egocentric Network for Visual Analytics

arXiv:2608.27194v1 Announce Type: new Abstract: This paper explores design considerations for egocentric network layouts in immersive environments, providing fresh empirical insights that enhance egocentric network analysis. An egocentric network focuses on the topological and semantic relationships around a focal node (ego) and its neighboring nodes (alters), targeting local sub-networks rather than the whole network. Traditional desktop environments, limited by display constraints, often face visual clutter as node numbers grow. Building on recent findings that immersive environments enhance network analysis, we explore layouts tailored for these spaces. We begin by identifying essential design properties and dimensions for egocentric network layouts, taking into account the unique features of immersive environments. Based on these, we design four layouts-Cube, Cylindrical, Radial, and Spherical-that vary across design dimensions. We evaluate these layouts in a user study with 24 par

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Exploring Normativity in Stable Diffusion: Insights for XAI in the Arts

arXiv:2608.26980v1 Announce Type: new Abstract: Generative text-to-image (T2I) systems are increasingly adopted in creative practice, yet their normative behaviors remain underexplored from the perspective of creative practitioners. In this workshop paper, we present a within-subject study with 14 creative practitioners using Stable Diffusion to create illustration from two tasks of differing specificity. We investigate whether and how practitioners perceive normative behavior in a T2I system and how it impacts their creative process depending on task specificity. Our findings show that participants perceived normative behavior through invariant patterns, stereotypical output, and unsolicited omission or addition of details. These experiences led to feelings of disempowerment and creative compromise. We discuss implications for XAIxArts, including prompt transparency and artist empowerment in creative contexts.

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Dynamic Tree Colors: Adaptive Discriminable Hierarchies with Minimum Instability

arXiv:2608.26734v1 Announce Type: new Abstract: Hierarchical color maps can support users in the analysis of hierarchical data. For large hierarchies, dynamic color maps can improve discriminability upon user interactions, but the incremental color changes may cause users to lose their orientation in the data set. To address this challenge, we present Dynamic Tree Colors, a dynamic hierarchical color map that can be configured to a suitable tradeoff between discriminability and color stability. We also define quality metrics for both criteria and investigate our algorithm's performance with respect to these metrics as well as a user study with 18 participants. Our results indicate that Dynamic Tree Colors yields good results in a wide range of application scenarios, but it does not achieve the performance of the state-of-the-art algorithm Cuttlefish in the specific scenario that algorithm was designed for.

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

RegulAR: Graph-Grounded Error Recognition and Assistance for Procedural Tasks in AR

arXiv:2608.26715v1 Announce Type: new Abstract: Errors are inevitable in procedural tasks, yet most AR guidance systems focus on step-by-step instruction delivery rather than helping users recognize and recover from mistakes. We present RegulAR, an AR task assistant for procedural error recognition and recovery. RegulAR models task instructions as a hierarchical dependency graph and combines this structure with a Multimodal Large Language Model (MLLM) to interpret egocentric observations during execution. This enables RegulAR to track progress, identify deviations by error type, estimate their impact on later steps, and deliver appropriately salient interventions through an in-situ head-up display that visualizes task state and recovery guidance. By making procedural structure explicit, RegulAR supports not only next-step guidance, but also reasoning about what went wrong, why it matters, and how users can get back on track. In a within-subject study (N=12), participants reported bette

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

EmoSay: Artificial Intelligence-Driven Text-to-Emotional-Speech System for Affective Communication in Extended Reality

arXiv:2608.26566v1 Announce Type: new Abstract: While contemporary neural text-to-speech (TTS) systems have achieved high levels of intelligibility, they frequently lack the emotional nuance required for authentic affective communication. This limitation is particularly critical in Extended Reality (XR), where the absence of emotionally expressive audio can diminish user presence and spatial immersion. We present EmoSay, an Artificial Intelligence-driven Text-to-Emotional-Speech (TTES) system designed to bridge the semantic-affective gap in immersive environments. EmoSay modulates a neural synthesis pipeline using discrete emotional prompts, delivering the output through a Unity-based interface featuring high-fidelity spatialized audio. The system was evaluated through a comprehensive user study focusing on perception, engagement, and the subjective sense of empathy. Our results demonstrate that EmoSay significantly enhances the immersive experience, achieving a System Usability Scale

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Kale: A Transformation-Safe Spreadsheet System

arXiv:2608.26345v1 Announce Type: new Abstract: Spreadsheet formulas can refer to rectangular ranges of arbitrary size. When a user changes the structure of a referenced table, the spreadsheet system updates the references to refer to a new range. Unfortunately, this new range may differ from the user's expectations, introducing bugs in spreadsheets. We describe a user study showing that standard reference semantics are error-prone, resulting in significant risk to users. We introduce Kale, a prototype system that eliminates the risk of inserting these kinds of bugs by restricting the kinds of references that can be expressed. We show that Kale can be used effectively by users to complete tasks that are error-prone in traditional spreadsheet systems. Finally, we describe a corpus study that evaluates the extent to which the reference restrictions in Kale might have implications on users.

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Calibration-Free Cuffless Blood Pressure Estimation Using Multimodal ECG-PPG Fusion on a Google Pixel Watch

arXiv:2608.26325v1 Announce Type: new Abstract: Inadequate blood pressure (BP) monitoring and management outside of clinical settings can worsen major cardiovascular risk factors such as hypertension. While cuff-based devices are commonly used for at-home monitoring, these devices can be inconvenient for daily use due to their sensitivity to body positions, upper-arm constrictions, and limited portability. A promising alternative is emerging in the form of consumer-grade smartwatches, where physiological signals related to cardiac activity can be used to estimate BP non-invasively and continuously across daily living conditions. In this work, we use data collected from a Google Pixel Watch in 40 participants to develop and compare several algorithm approaches for BP estimation. We found that our proposed deep learning model achieved the strongest overall performance, and that fusing smartwatch signals with demographic information improved model generalizability to unseen individuals. H

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.HC

Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning

arXiv:2608.26166v1 Announce Type: new Abstract: Advancing reasoning capabilities allow large language models (LLMs) to tackle increasingly complex problems, while reasoning traces - intermediate steps toward solutions - open up high-stakes applications by enabling human inspection of AI decision-making. However, current approaches prioritize model performance over human interpretability, limiting effective human-AI collaboration. In this study, we design and evaluate a human-centered approach that structures reasoning traces based on self-contained, verifiable steps, enabling users to independently assess and correct AI reasoning. Our approach uses XML-like tags to encode reasoning content and metadata, facilitating targeted feedback. Evaluation on mathematical reasoning tasks shows our approach maintains equivalent performance to standard Chain-of-Thought reasoning while enhancing interpretability. User studies demonstrate significant improvements in perceived usefulness and ease of u

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training

arXiv:2608.20574v2 Announce Type: replace-cross Abstract: Open-ended language-model evaluation often substitutes another model or a small preference panel for a missing answer key. We introduce FlavourBench, which instead compiles dense answer maps from a versioned culinary environment. Each task asks for a three-ingredient portfolio from eight candidates; before inference, Epicure scores all 56 portfolios. We evaluate 27 frontier endpoints on the same 534 substitution, pairing, and constraint tasks, yielding 14,418 complete model-task observations. Anchor-cluster bootstraps and multiplicity-controlled paired tests resolve 101 of 351 model contrasts. Grok 4.6 has the largest point estimate at 65.1, but the corrected evidence does not identify a unique best endpoint. The ranking replicates across independently compiled panels and remains similar under alternative metrics, task filters, family weights, and three public Epicure checkpoints. We then run a preregistered, three-seed post-tra

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade

arXiv:2509.07274v4 Announce Type: replace-cross Abstract: Migration has been a core topic in German political debate, from postwar expellee displacement to labor migration and recent refugee movements. Large-scale analysis of such political discourse has traditionally required extensive manual annotation, limiting coverage. Large language models (LLMs) offer a scalable alternative. Using a theory-driven annotation scheme, we examine how well LLMs annotate subtypes of solidarity and anti-solidarity in German parliamentary debates and whether the resulting labels support valid downstream inference. We first evaluate multiple LLMs across model size, prompting strategies, fine-tuning, historical versus contemporary data, and systematic errors. The strongest models, especially GPT-5 and gpt-oss-120B, achieve macro- F1 scores comparable to human agreement, although their systematic errors can bias downstream results. We therefore combine soft-label model outputs with Design-based Supervised

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Small Changes, Big Impact: Demographic Bias in LLM-Based Hiring Through Subtle Sociocultural Markers in Anonymised Resumes

arXiv:2603.05189v3 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed in resume screening pipelines. Although explicit PII (e.g., names) is commonly redacted, resumes typically retain subtle sociocultural markers (languages, co-curricular activities, volunteering, hobbies) that can act as demographic proxies. We introduce a generalisable stress-test framework for hiring fairness instantiated in the Singapore context: 100 neutral job-aligned resumes are augmented into 4100 variants spanning four ethnicities and two genders, differing only in job-irrelevant markers. We evaluate 18 LLMs in two settings: (i) Direct Comparison (1v1) and (ii) Score & Shortlist (Top-Score Rates), each with and without rationale prompting. We find that even without explicit identifiers, models recover demographic attributes with high F1 and exhibit systematic disparities, with models favouring markers associated with Chinese and Caucasian males. Ablations show language mark

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Exploring the Role of Automated Feedback in Programming Education: A Systematic Literature Review

arXiv:2602.00089v2 Announce Type: replace Abstract: Automated feedback systems have become increasingly integral to programming education, where learners engage in iterative cycles of code construction, testing, and refinement. Despite its wider integration in practices and technical advancements into AI, research in this area remains fragmented, lacking synthesis across technological and instructional dimensions. This systematic literature review synthesizes 61 empirical studies published by September 2024, offering a conceptually grounded analysis of automated feedback systems across five dimensions: system architecture, pedagogical function, interaction mechanism, contextual deployment, and evaluation approach. Findings reveal that most systems are fully automated, embedded within online platforms, and primarily focused on error detection and code correctness. While recent developments incorporate adaptive features and large language models to enable more personalized and interactiv

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Testing Fairness with Utility Tradeoffs: A Wasserstein Projection Approach

arXiv:2505.11678v4 Announce Type: replace Abstract: Ensuring fairness in data driven decision making has become a central concern across domains such as marketing, lending, and healthcare, but fairness constraints often come at the cost of utility. We propose a statistical hypothesis testing framework that jointly evaluates approximate fairness and utility, relaxing strict fairness requirements while ensuring that overall utility remains above a specified threshold. Our framework builds on the strong demographic parity (SDP) criterion and incorporates a utility measure motivated by the potential outcomes framework. The test statistic is constructed via Wasserstein projections, enabling auditors to assess whether observed fairness-utility tradeoffs are intrinsic to the algorithm or attributable to randomness in the data. We show that the test is computationally tractable, interpretable, broadly applicable across machine learning models, and extendable to more general settings. We apply

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

arXiv:2608.27309v1 Announce Type: cross Abstract: Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them. We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Research Design Tracking and Assessment for the Social Sciences

arXiv:2608.27049v1 Announce Type: cross Abstract: Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis. We introduce Automated Research Design Tracking and Assessment (ARDTrA), a task that involves detecting the research design used in a paper and assessing the quality of its application. We create an expert-annotated dataset of papers covering six families of counterfactual research designs and evaluate the task using a multi-turn RAG-based conversational pipeline. Across four retrieval strategies, four LLMs and six embedding models, we find that passage length is the main driver of performance, explaining 52-66% of the variance. A per-research-design analysis also shows that human and machine difficulty do not align: the designs that prove hardest for the system are not those on which expert annotators disagree most, pointing to two independent sources of task diffi

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems

arXiv:2608.26849v1 Announce Type: cross Abstract: User behavior simulation with large language models~(LLMs) is increasingly used to support multi-agent ecosystem simulation. Existing simulators typically rely on static user profiles inferred from historical observations, which become inadequate in socially intensive environments such as live streaming where interaction dynamics continuously reshape user behavior. We propose \textbf{LiveSim}, an LLM-based framework for live-stream ecosystem simulation. It represents users as editable behavioral hypotheses and progressively refines them through trajectory-grounded interactions, where discrepancies between simulated and observed trajectories reveal missing environmental shaping effects. These signals are further extracted as transferable environment-behavior patterns and accumulated in a collective behavioral memory to improve user-level behavioral fidelity and support ecosystem-level simulation. Experiments on real-world live-stream ris

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

DIRECT: Decomposing Audience Preference and Creative Effect in Visual Content Analytics

arXiv:2608.26584v1 Announce Type: cross Abstract: Which visual choices make a post perform better? A growing literature answers this question with pooled coefficients estimated across many creators, which platforms translate into creative recommendations. We show that these coefficients blend two distinct patterns that can point in opposite directions for the same attribute. The first, audience preference, arises because creators who favor a style attract differently composed audiences, so their posts perform differently because of who is watching, not what any single post does. The second, creative effect, captures how a creator's audience responds when she departs from her usual look. Pooled estimation averages the two, and audience preference can be large enough to reverse the signal that creative direction requires. We propose DIRECT (Decomposed Identification of Response Effects via Causal Tools), a panel-based causal-inference framework that separates them, combining the Mundlak

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Decolonial Discourse in Postcolonial Contexts: How YouTubers Negotiate Audience Tensions, Platform Governance, and State Influence

arXiv:2608.26351v1 Announce Type: cross Abstract: Decolonial discourse on online platforms is often framed in terms of creator motivations and expressive possibilities. In this paper, we examine what it takes to sustain such discourse under layered sociotechnical constraints. Drawing on semi-structured interviews with YouTubers engaging in Bengali decolonial discourse, we analyze how audience publics, platform governance, and state influences shape what becomes sayable, visible, and viable. We show how fragmented postcolonial identities among audiences produce legitimacy policing, harassment, and coordinated backlash, requiring ongoing relational labor from creators. At the platform level, differential monetization, opaque moderation, and copyright regimes reorganize which publics are economically viable and reinforce existing hierarchies. Further, intermediaries such as multi-channel networks mediate regulatory pressure, introducing political risks and constraints on participation. In

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Assessing Company Contributions to Societal Resilience: Extending the Societal Capacity Assessment Framework to Agentic AI

arXiv:2608.27238v1 Announce Type: new Abstract: Companies that deploy AI agents and make them available to others are creating the sociotechnical circumstances under which this technology integrates into existing social and economic structures. AI-deploying companies are institutional actors that actively shape society's capacity to withstand and govern the consequences of agentic AI. In view of these societal impacts, companies can build societal resilience by designing and promoting safer implementations of AI agents. To operationalize this goal, this paper adapts the indicator-based Societal Capacity Assessment Framework (SCAF) to measure how a company's deployment decisions contribute to societal resilience, inverting its original measurement of societal resilience as a backdrop for deployment decisions (Gandhi et al., 2025). Our procedure has two steps: a conceptual step in which we design a suite of indicators that define what SCAF's vulnerability, coping, and adaptive capacities

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Animarium: an open, reproducible pipeline for synthetic populations of Italian cities, from ISTAT sources to open data (Tech Report v1)

arXiv:2608.27111v1 Announce Type: new Abstract: Synthetic populations of eleven Italian municipalities (1,814,317 individuals in 887,937 households) generated from published aggregates alone: ISTAT census and register tables, census-section counts, the national civic-address register, public-use survey microdata, and six municipal open-data portals, every source certified in a registry with licence, fingerprint and declared affordances. Four rings give every attribute a declared place: a maximum-entropy joint model of up to nine demographic attributes; whole-vector donation of twenty-three attitudinal and health variables from survey respondents; placement to census section, single year of age and address; and households constrained by the census size distribution per section. Every downstream layer (detailed titles, work, names, biographies) is a declared derivation adding no information. The pipeline is deterministic to the byte: regenerating all eleven municipalities from the tagged

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Reclaiming Epistemic Agency: A Critical Framework for Human-Generative AI Co-Agency in Education

arXiv:2608.26937v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) has been primarily framed as an impartial educational tool. However, this framing overlooks an even larger shift: the reassignment of epistemological authority from teachers to students to machines. This paper presents a conceptual evaluation of the extent to which GenAI redistributes students' and teachers' ability to act in classrooms to produce knowledge, validate each other's claims, and create evidence of student learning while collaborating with and competing against humans. This evaluation draws on various theoretical paradigms, including Distributed Agency, Self-Determination Theory, Society 5.0, and Technology Integration Paradigms, including TPACK and SAMR. While all of the theoretical paradigms evaluated are relevant to the role of agency within education mediated by AI, none of them address the ongoing disparity regarding equitable distribution of power, ownership of the data used to

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

How Does Science Education Research Respond to Sociopolitical Change? A BERTopic Analysis of Korean Research

arXiv:2608.26675v1 Announce Type: new Abstract: Research fields do not evolve in isolation: their questions and priorities shift with policy, curriculum reform, and broader social change. Analyzing published literature can reveal not only how a field matures but also how it responds to these conditions. Prior work in science education has focused on identifying research topics and their trends, but paid less attention to the external conditions in which research is produced. We examine Korean science education research from 2008 to 2025, a case in which centralized curriculum revision, government education initiatives, and demographic decline are prominent. Using BERTopic, an embedding-based topic modeling technique, we identify major topics and temporal trends, and analyze their associations with selected sociopolitical factors. We interpret each topic and distinguish three groups: sociopolitical, subject-specific, and student-related topics. Within the first group, science teacher pr

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

ClassVision: AI-Powered Classroom Attendance System

arXiv:2608.26173v1 Announce Type: new Abstract: Students and working professionals have to go through the attendance process every day. Traditional methods of marking attendance using pen and paper or online platforms are human-intensive and time-consuming. To address the challenges in manual attendance processes, this research explores the use of face detection (FD) and face recognition (FR) technology to automate the attendance process, particularly in educational settings, and build a ClassVision course attendance system. We also propose an automated attendance system featuring a human-computer interaction (HCI) and user-friendly web interface that utilizes real-time image processing to identify and recognize students in classrooms and automatically record their attendance. We identified RetinaFace as the best face detection model, and when combined with Face Recognition for verification, it provided the most promising results with a cropped embedding of 50x50 pixels.

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints

arXiv:2608.26171v1 Announce Type: new Abstract: Multi-stage LLM hiring pipelines (resume improvement, interview question generation, answer feedback) can fabricate credentials, inflate qualifiers, and invent experience. We evaluate two mitigations, prompt guardrails and human-in-the-loop (HITL) checkpoints, against a fully automated baseline. In a controlled experiment (10 synthetic resumes x 2 job descriptions x 3 repetitions x 3 conditions; 180 runs), the baseline (C1) produced at least one unsupported claim in 96.7% of outputs (mean 6.80 findings/output). Prompt guardrails (C2) reduced finding density by 86% (6.80 to 0.92/output), but 50.0% of outputs still contained a fabrication, showing prompt-level mitigation alone is insufficient. A human checkpoint after resume improvement (C3) eliminated all identity fabrications, reduced finding density by 59% (6.88 to 2.82/output), reduced item-level fabrication from 96.7% to 75.0% (p=.022), and cut capture of JD-embedded trap requirements

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Revision-Aware Success Prediction from Multi-Attempt Programming Trajectories

arXiv:2608.26169v1 Announce Type: new Abstract: Programming outcome prediction plays a central role in data-driven programming education, supporting learner modeling, timely intervention, and adaptive assistance. Yet predicting submission success is difficult due to heterogeneous error states, short-term revisions, and uneven future-horizon availability in programming trajectories. This study examines three prediction tasks under a unified formulation: whether the current attempt is accepted (Task~1), whether the next attempt is accepted (Task~2), and whether acceptance is reached within a three-attempt recovery window (Task~3). Each task is evaluated across current-only, pairwise, and multi-step input regimes using ML, DL, and transformer-based pretrained models (PTM), represented by LinearSVM, XGBoost, BiGRU, BiLSTM, GraphCodeBERT, and CodeT5+. Results show a consistent pattern: the current-only regime is the most reliable, while pairwise and multi-step history provide no consistent

Source ↗
Showing 9301–9350 of 11035 signals
← Prev Page 187 of 221 Next →