EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

LearnAI: Just-in-Time AI Co-Creation Across Disciplines at a University

arXiv:2608.19164v1 Announce Type: new Abstract: As generative AI reshapes professional and educational practice, institutions face a challenge: how to support diverse learners, from non-coders to advanced students, in building confidence and practice with AI-supported problem solving. Most institutional responses bifurcate into conceptual workshops for general audiences or technical courses for computer science majors, leaving few spaces where mixed-ability learners can engage common AI tasks at levels matched to their prior experience. This experience report presents the LearnAI Framework, a two-layer model for just-in-time AI co-creation piloted at a comprehensive teaching university. The Wide-Exposure Layer embeds short presentations in existing courses to build AI awareness at scale, reaching students and faculty across 18 courses in five disciplines. The Customized Co-Creation Layer provides opt-in, one-on-one sessions where clients work with trained undergraduate tutors through a

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hot Games: Towards a Holistic Assessment of the Planet Warming Emissions of Video Games based on 2024-2025 Data

arXiv:2608.19040v1 Announce Type: new Abstract: Following on recent reports on specific platforms or companies, this paper provides an assessment of the global impact of the production and use of video games. It draws together publicly available data on game development, hardware, games sold, download sizes, time spent playing games on different platforms, and subscriptions to multiplayer and cloud game services. It provides an update to figures published 2020 and 2022. Crucially, our account of emissions related to video games considers a wide range of categories, yet contains enough detail to be critiqued and improved in the future.

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Open at the Edge, Captured at the Center: llama.cpp and the Political Economy of Local AI Inference

arXiv:2608.19001v1 Announce Type: new Abstract: Open AI scholarship has focused on model releases and cloud ecosystems, leaving the local inference infrastructure that makes open-weight models runnable on user-owned devices largely unexamined. We address this gap through a mixed-methods analysis of llama.cpp, combining 7,681 merged pull requests from March 2023 through March 2026 with repository discussions, corporate statements, and contributor blogs. We show that local inference broadens participation at execution while relocating capture into the infrastructure that makes execution possible. Through hardware backends, model integration labor, and Hugging Face's February 2026 absorption of the project, we document how control shifts to hardware vendors, model distributors, and core maintainers while model owners and individual contributors bear the cost of making models runnable. These dynamics suggest that preserving openness outside the cloud requires attention to the infrastructur

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Epistemic Subordination: Generative AI and the Infrastructure of Knowledge

arXiv:2608.18758v1 Announce Type: new Abstract: Generative AI does not merely produce biased outputs. It encodes the majority's way of knowing as the default infrastructure of knowledge itself. We call this epistemic subordination. The training process compresses the full breadth of human expression into a single probabilistic model whose statistical baseline reflects the languages, assumptions, and cultural frameworks of the dominant culture. Minority epistemologies are not excluded but absorbed: present in the training data, yet structurally subordinated in the output. The result is not a collection of discrete biases that can be audited and corrected. It is an epistemic condition embedded in the architecture from which all outputs emerge. This unified harm cuts across three legal domains -- anti-discrimination law, cultural and linguistic rights, and democratic viewpoint pluralism -- and each fails to address it for the same structural reason: existing law regulates downstream, at t

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Turning interest into institutional change: teaching advocacy for sustainable research

arXiv:2608.18601v1 Announce Type: new Abstract: Systemic change across the digital research landscape is required to reduce the environmental impact of digital research, but while many researchers and technical professionals are motivated to act, they often lack the skills required to translate motivation into lasting organisational change. We present an open-access course that teaches the foundations of advocacy and organisational change to researchers, research software engineers, and research technical professionals. Structured around the UNICEF five-step advocacy cycle, the course covers stakeholder analysis, power mapping, coalition building, storytelling, framing and messaging, and evaluation. It is grounded in the UK policy landscape, including the Concordat for Environmental Sustainability and the UKRI Environmental Sustainability Strategy, and uses a fictional case study to make concepts concrete. The course is designed as a dual-layer resource in which workshop slides and ext

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks

arXiv:2608.18554v1 Announce Type: new Abstract: Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent's performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are scored through blind pairwise comparisons by an LLM judge panel with task-specific rubrics, replicated across ten runs. Rankings across the two regimes are only modestly correlated, and the automation winner loses augmentation on five of seven tasks. A

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Fabricated Front: Generative AI and the Opacity of Workplace Performance

arXiv:2608.18369v1 Announce Type: new Abstract: Generative AI (GenAI) has become a fixture of workplace life. Current research asks chiefly what this implies for jobs and outputs, measured in productivity, displacement, or bias. What remains underexamined are the interactional reconfigurations that GenAI produces at work. The emerging concept of effort opacity has begun to fill this gap by highlighting the systematic decoupling of observable output from human engagement. When GenAI makes interactional cues less diagnostic, it weakens the reciprocal exchange that sustains collaborative trust. Extending this account of effort opacity, we examine the interactional mechanics that produce opacity in everyday workplace encounters. Drawing on Erving Goffman's dramaturgical framework and 1,250 interview transcripts from Anthropic's AI Interviewer dataset, we identify five opacity mechanisms through which workplace fronts are reorganized: voice (whose stance the words index), provenance (who ca

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Capability-Based Planning for AI Crisis Preparedness

arXiv:2608.18357v1 Announce Type: new Abstract: Capability-based planning drives preparedness in defense and homeland security, but has yet to be applied seriously to AI. Government AI preparations follow a predict-then-act paradigm: rank risks by likelihood and impact, then prepare for the highest expected harm. AI resists prediction: expert timelines disagree by orders of magnitude, and official reviews concede that likelihood-based risk assessment fails for exactly this class of risk. Drawing on principles of decision making under deep uncertainty, we propose a methodological framework in three parts: a scenario library sampled systematically across declared axes; a rating procedure that assesses each government capability against each scenario on coarse, gated criteria; and a prioritization step that maps the resulting matrix onto decision rules a government might adopt. Through a pilot across the four most severe AI-enabled threat classes, we illustrate the kind of insight the ins

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

arXiv:2608.18296v1 Announce Type: new Abstract: As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows t

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Global Index on Responsible AI 2026 : Conceptual Framework and Methodology

arXiv:2608.18122v1 Announce Type: new Abstract: This report presents the methodology of the Global Index on Responsible AI (GIRAI), 2nd Edition. This edition refines the 1st Edition by strengthening the distinction between framework existence and implementation, restructuring dimensions from three to five thematic areas, introducing more granular variables for framework quality, and applying a multi-stage review and validation process. An independent statistical pre-audit was conducted to assess the coherence and robustness of the framework. GIRAI assesses responsible AI governance across five dimensions: Inclusion and Diversity, Ethics and Sustainability, Labour and Skills, Trust and Safety, and Use of AI in Public Service. Each dimension has a number of indicators (38 in total), organised into three pillars, namely AI Policy (17 indicators on government frameworks and implementation, assessed through primary data), CSO Engagement (5 indicators, primary data), and Enabling Conditions

Source ↗
technology Thu, 18 Jun 2026 15:03:36 -0400
EdTech Mag (Higher)

The New Campus Reality: Building Cyber Resilience Against Ongoing Threats

It’s not a matter of if, but when. This cybersecurity maxim is true for almost any organization, but it is especially true for higher education institutions. They are continuing to experience a significant uptick in attacks, numbering about 4,200 per week in 2026 across higher education institutions, according to Randy Rose, vice president of security operations and intelligence at the Center for Internet Security (CIS). “We’re holding steady for 2026, but that’s not necessarily a good thing,” says Rose. “Depending on who’s measuring it, higher education saw anywhere from a 20% to 40%…

Source ↗
technology Thu, 18 Jun 2026 15:00:15 -0400
EdTech Mag (Higher)

Why Data Readiness Is the Foundation for AI Readiness in Higher Education

Every board wants to know the AI plan, but AI readiness starts with a question most institutions haven't answered: is your data ready? Simply put, AI readiness starts with data readiness. You don’t build a house without a solid foundation. The stronger your data as your foundation is, the greater opportunity that you have to build, and we are all building right. Our goal is not to be static. Our goal is to help our organizations grow, be more effective for our students and achieve the outcomes that higher ed is there to provide. Click the below banner to explore building data governance…

Source ↗
technology Thu, 18 Jun 2026 14:13:38 -0400
EdTech Mag (Higher)

Data Governance Is Just the Beginning: Why University IT Leaders Must Also Master These Data Disciplines

In addition to CIOs establishing themselves as leaders when it comes to a unified data strategy and university leadership understanding that data governance is the foundation of AI readiness, there is a growing understanding that data governance is a required discipline, essential to data-centric transformation on campuses. However, there are other data considerations to be mindful of, as well. Click the banner below to explore how to build a foundation for scalable AI at your higher ed institution.

Source ↗
technology Thu, 18 Jun 2026 13:10:52 -0400
EdTech Mag (K-12)

AI Phishing Gains Inside Access to Vulnerable K–12 Data

Artificial intelligence has rapidly transformed K–12’s cyberthreat landscape, turning phishing scams into more sophisticated, multichannel attacks that exploit trust, familiarity and the platforms educators and students use every day. Phishing is no longer just an inbox problem — it’s an “everywhere” problem. For many years now, we’ve taught K–12 staff and teams to check an email sender’s address as one way to stay safe. In today’s threat landscape, the advent of AI-powered vishing, deepfake impersonations and automated social engineering, that advice is now obsolete. Cyber fraud is now…

Source ↗
technology Thu, 18 Jun 2026 09:00:00 +0000
Tech & Learning

Creating 5 Pillars To Guide AI Use In Your District

Innovative Leader Award - Director of Information Technology Kadion Phillips discusses implementing AI in a school district as well as how to bolster cybersecurity.

Source ↗
technology Thu, 16 Jul 2026 14:10:27 -0400
EdTech Mag (K-12)

Five Keys to Cyber Resilience in K–12 Schools

Although cybersecurity has been on K–12 districts’ radar for several years, many have yet to achieve maturity. The goal is resilience — the ability not only to detect an attack that is already underway, but also to anticipate threats, manage risks, minimize the impact of attacks and restore secure operations quickly. Here are five pillars of a resilient cybersecurity posture. Click the banner below for a cybersecurity roadmap built for K–12 schools.

Source ↗
technology Thu, 16 Jul 2026 13:13:33 -0400
EdTech Mag (K-12)

5 Questions to Ask When Choosing EdTech Solutions

Technology can accelerate learning when chosen with care, but it can also become a distraction if it doesn’t meet the everyday needs of students and educators. Education program budgets nationwide are in flux, and demands on educators are high, so it’s crucial to evaluate edtech tools strategically. Here are five questions every school leader should ask before making their next technology investment. Click the banner below to find the guidance you need to manage your device ecosystem.

Source ↗
technology Thu, 16 Jul 2026 13:02:53 -0400
EdTech Mag (Higher)

Why Higher Ed’s Pandemic Playbook Is the Blueprint for AI Infrastructure

Higher ed IT leaders grappling with integrating rapidly accelerating artificial intelligence throughout their campuses need only flip back a few pages in their playbooks to the pandemic. Institutions that invested in hybrid cloud flexibility and business continuity before 2020 pivoted to remote work and hybrid learning far more quickly than those still in legacy environments. Universities that leaned into early modernization in the age of AI find themselves equally well positioned now. The mindset shift that serves college and university IT teams — and one that Nutanix is built around — is…

Source ↗
technology Thu, 16 Jul 2026 09:00:00 +0000
Tech & Learning

Summer Prep To Protect Your School From AI-Enabled Explicit Content

3 actionable steps to guide your response when students use AI to generate explicit content

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence

arXiv:2606.16944v2 Announce Type: replace-cross Abstract: Theory of mind (ToM), the capacity to ascribe mental states to others and use those ascriptions for prediction and inference, is widely assumed to be essential for effective human-machine integration. Existing AI-ToM models address \emph{how} to mentalize, but leave the question of when largely unaddressed. The central question is: under what situational and agent-level conditions is ToM engagement causally warranted in conflict? This paper presents a structural causal model formalized as a directed acyclic graph (DAG), treating ToM as a mechanism activated by situational and agent-level conditions rather than as an always-on capacity. The model specifies four exogenous variables capturing situational and agent-level conditions, five endogenous mediators, and a mechanistic ToM node producing engagement states through three distinct causal pathways: a tractability pathway, a reasoning-depth pathway, and an enabling-cause pathway.

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

PhysClaw-0: A Symbiotic Agentic System for Robot Autonomy via Language Corrections

arXiv:2607.14047v1 Announce Type: cross Abstract: Autonomous data collection governs the volume and quality of real-world trajectories for manipulation policy learning. Existing pipelines reduce human effort via self-resetting, VLM verification, or language-guided correction, yet episode-scoped fixes must be reissued whenever the same failure recurs, so oversight cost grows with session length rather than with the number of distinct problems. We present PhysClaw-0, a human-robot symbiotic agentic system in which corrections are retained and reused across rounds. The collection loop collects, verifies, and resets autonomously, pausing for a remote operator only when a phase exhausts an explicit retry budget. An LLM parser maps each natural-language utterance to a structured adjustment stored in Corrective Memory, so addressed failure modes typically need not be corrected again under the same conditions. On a real-robot desktop-clearing testbed, PhysClaw-0 matches teleoperation episode s

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects

arXiv:2607.13679v1 Announce Type: cross Abstract: AI agents are joining human teams, raising a basic question: when an automated agent becomes a regular participant, does group organization strengthen or weaken? We study this question in open-source software, where bots open pull requests, review code, and merge changes alongside people, leaving a public record of every interaction. Treating bots as participants rather than tools, we examine 2,991 GitHub projects for two years before and after each adopted its first bot. We measure three capabilities that institutional theory links to durable coordination - repeated engagement, social memory, and role differentiation - and two outcomes: conflict cascades and output distinctiveness. Bot adoption is followed by more repeated collaboration, greater recognition of specific bots in discussion, fewer conflict cascades, and more distinctive outputs. These changes cluster around adoption rather than accumulating gradually. Because we lack an u

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

arXiv:2607.13465v1 Announce Type: cross Abstract: LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. However, real-world user goals often span multiple devices: information may come from a phone, be processed on a desktop, and the result may need to appear on another device. Most existing benchmarks center on a single dominant execution environment, making it difficult to evaluate whether agents can acquire and integrate information across heterogeneous devices and complete end-to-end tasks with cross-device dependencies. We introduce DevicesWorld, a large-scale executable benchmark for cross-device collaborative operation. DevicesWorld contains 6,140 tasks and integrates three classes of device environments -- mobile, desktop, and IoT -- into a unified cross-device interaction and evaluation framework. Each task defines a natural-language user goal, participating devices and initial states,

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

Marker-free deformable registration and fusion for augmented reality-guided positive margin localization during tumor resection surgery

arXiv:2607.13343v1 Announce Type: cross Abstract: Positive margins in head and neck oncologic surgery require mapping specimen-side pathology findings to the patient resection bed. This is challenging because pathologists identify the positive margin on slices of the resected, deformed specimen, while surgeons must relocate the corresponding site on the resection bed using only verbal descriptions and no visual guidance. We present a marker-free augmented reality (AR) workflow for mapping a margin label from a three-dimensional specimen scan to the resection bed. The method combines contour-constrained deformation, residual alignment to a depth scan, surface-based fusion to a head-mounted display, and target projection onto the reconstructed bed. Bead-suture correspondences estimate specimen deformation, whereas patient-to-display fusion does not require external fiducial markers. Following formative experiments, five residents and surgeons performed cadaveric cheek and scalp re-resect

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science

arXiv:2607.13220v1 Announce Type: cross Abstract: Most AI-for-science systems focus on scaling a single reasoning process through better models, larger context windows, long-horizon agentic execution, or digital co-scientists working with one principal user. However, challenging scientific problems are rarely solved by one reasoner alone. They are solved by teams whose members bring different priors, experimental backgrounds, tacit knowledge, and domain-trained intuitions. The open problem is therefore not only how to scale models, but how to cultivate networked intelligence: scaling the connections between humans and AI systems so that a result or hypothesis produced in one context reaches another person, agent, instrument, or robot that can act on it. We introduce Mycelium, an active shared workspace that automatically connects researchers and AI agents as a multi-user co-scientist. As human users and agents work, the system captures important observations and hypotheses, tracks how

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

ExpressionCueLens: A Cross-Cultural Analysis of Human-AI Companion Conversations on Social Media

arXiv:2607.13924v1 Announce Type: new Abstract: LLM-based AI companion agents are increasingly being perceived not only as tools but also as social companions. On social media, people recount conversations where these agents comfort, negotiate and assert boundaries, reflecting a growing attribution of human-like qualities. To profile how agency is perceived in human-AI (HAI) interactions, we introduce the ExpressionCueLens framework, which organizes linguistic, cognitive, behavioral and perceptual cues into ten categories of anthropomorphism expressions. We apply this framework to $\sim$3500 Reddit and XiaoHongShu posts that discuss HAI companionship. Through iterative expert annotation and LLM-assisted labeling, our cross-platform analysis indicates patterns consistent with the hypothesis that XiaoHongShu users use significantly more expressions of vulnerability and emotions, and more non-perceptual cues. Reddit users employ more perceptual cues with temporality and embodiment express

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

ZipLine: Visual Analysis of Multivariate Graphs with Predicate Logic

arXiv:2607.13767v1 Announce Type: new Abstract: Multivariate graphs unite two distinct data perspectives: a topological structure defined by nodes and edges, and attribute data associated with each node. Analyzing such graphs therefore requires reasoning across two complementary spaces. However, existing systems typically emphasize the analysis of one space at a time, focusing either on topology or on attributes. As a result, exploration, analysis, and pattern discovery that depend on their interaction remain difficult. In this paper, we present ZipLine, a system designed to support integrative analysis of multivariate graphs by bridging both topology and attribute spaces. ZipLine introduces a predicate language that enables analysts to express patterns involving topology, node attributes, and neighborhood relations with a unified formalism. The system further provides a predicate-learning algorithm that maps analyst interactions across both topology (e.g., subgraph selection) and attr

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

Interaction Density as a Behavioural Signature of Exhibit Type: A Minimal-Log Study from a Two-Venue Science Experience Centre

arXiv:2607.13724v1 Announce Type: new Abstract: Understanding how visitors engage with interactive exhibits usually calls for either labour-intensive manual observation or invasive multimodal sensing -- eye-tracking, cameras, wearables -- that few science centres can deploy at scale. We ask how much can be learned instead from the handful of fields that most touch-enabled exhibits already log by default: a session's start time, end time, and press count. Analysing 2,816 visitor sessions across eight exhibits at two venues of a science experience centre in Bengaluru, India, we derive interaction density -- presses per second -- as a simple behavioural signature, and use it to distinguish fast-paced games from slower, deliberate quizzes. Density does so cleanly (Mann-Whitney r=0.556) and predicts exhibit type on its own with a cross-validated AUC=0.778. But the data complicates the obvious story: games are not just more intense, visitors also dwell on them longer (r=0.172), reversing the

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

VIP-MINGLE: A Corpus for Videoconference and In-Person Multimodal Interaction in Group Language Engagement

arXiv:2607.13614v1 Announce Type: new Abstract: Group conversations are a fundamental yet complex form of social interaction central to human cognition and telecommunication technology. While understanding and facilitating these interactions has been a long-standing goal, findings are often isolated within specific in-person or videoconferencing settings due to a scarcity of datasets that bridge the two. We introduce VIP-MINGLE, a multimodal dataset comprising 59 hours of recordings (32 groups, 105 participants), featuring paired within-subject sessions in both settings. The dataset includes raw audio/video, psychometric data, processed multimodal features (e.g., diarized speech, facial expressions, transcriptions), and time-resolved human annotations. Our analysis reveals significant behavioral distribution shifts across multiple modalities between settings, reinforcing the need for a cross-setting corpus. VIP-MINGLE serves as a critical resource for developing robust models of group

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

TANDE: Disentangling Verbal and Nonverbal Backchannels in Emotional AI-Avatar Conversations with Young Adults

arXiv:2607.13357v1 Announce Type: new Abstract: Embodied conversational agents (ECAs) need effective empathic grounding to foster social support and engagement. Expanding into emotional domains, ECAs now use Large Language Models (LLMs) and multimodal human-agent interactions to enhance their capabilities. Yet, understanding the impact of backchanneling modalities on young adults and their gender remains limited. We introduce TANDE, an LLM-powered ECA designed for emotional conversations with young adults, a population experiencing mental, personal, and social issues with limited tools to address them. In a within-subjects study with N=36 young adults, we explore nonverbal and combined verbal-and-nonverbal backchanneling modalities on rapport, empathy, and engagement and isolate for gender differences. Our research shows the importance of nuanced backchanneling cues with emotional ECAs with young adults, showing a preference for nonverbal cues. We derive design implications for more ef

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.HC

SoftBoard: A Multi-Agent Tool for the Creation and Evaluation of Low-Fidelity Prototypes

arXiv:2607.13179v1 Announce Type: new Abstract: User Experience (UX) is recognized as a critical factor for the success of digital products, particularly in software startups, environments marked by time constraints, limited resources, and low maturity in design practices. Building Minimum Viable Products (MVPs) through low-fidelity prototyping represents a well-established strategy for rapid validation cycles at reduced cost. A systematic literature mapping, however, revealed gaps in the ecosystem of available tools: a predominance of general-purpose solutions adapted for prototyping, the absence of integrated methodological guidance, and the incipient use of Artificial Intelligence in the design process. This paper presents SoftBoard, a web-based tool for the creation and evaluation of low-fidelity prototypes in the context of MVP development. The tool integrates a prototype editor, team-based project organization, and a multi-agent system based on large language models that supports

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation

arXiv:2606.02528v2 Announce Type: replace-cross Abstract: Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. We ask three questions: do LLMs systematically prefer certain financial instruments; can an internal representation with causal leverage over those preferences be identified; and does that representation affect downstream financial decisions? We develop a three-level audit protocol and apply it to Bitcoin. First, a behavioral audit of nine frontier LLMs shows that Bitcoin's ranking among money-like instruments is frame-dependent: models place it around rank 5 of 8 as "reliable money" but near the top under crisis and autonomous-agent frames, and an attribute-swap experiment shows that rankings track functional properties, not names. Second, we open a model's internals: a search across thousands of sparse-autoencoder features in Gemma 3 identifies a dominant Bitcoin-selective feature

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Temporal Shifts and Causal Interactions of Emotions in Social and Mass Media: A Case Study of the "Reiwa Rice Riot" in Japan

arXiv:2602.14091v2 Announce Type: replace-cross Abstract: In Japan, severe rice shortages in 2024 sparked widespread public controversy across both news media and social platforms, culminating in what has been termed the "Reiwa Rice Riot." This study proposes a framework to analyze the temporal dynamics and causal interactions of emotions expressed on X (formerly Twitter) and in news articles, using the "Reiwa Rice Riot" as a case study. While recent studies have shown that emotions mutually influence each other between social and mass media, the patterns and transmission pathways of such emotional shifts remain insufficiently understood. To address this gap, we applied a machine learning-based emotion classification grounded in Plutchik's eight basic emotions to analyze posts from X and domestic news articles. Our findings reveal that emotional shifts and information dissemination on X preceded those in news media. Furthermore, in both media platforms, the fear was initially the most

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Value Drifts: Tracing Value Alignment During LLM Post-Training

arXiv:2510.26707v2 Announce Type: replace-cross Abstract: As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems. Therefore, studying the alignment of LLMs with human values has become a crucial field of inquiry. Prior work, however, mostly focuses on evaluating the alignment of fully trained models, overlooking the training dynamics by which models learn to express human values. In this work, we investigate how and at which stage value alignment arises during the course of a model's post-training. Our analysis disentangles the effects of post-training algorithms and datasets, measuring both the magnitude and time of value drifts during training. Experimenting with Llama-3 and Qwen-3 models of different sizes and popular supervised fine-tuning (SFT) and preference optimization datasets and algorithms, we find that the SFT p

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Agentic Web Requires New Normative Infrastructure

arXiv:2606.10711v2 Announce Type: replace Abstract: The agentic web, in which users interact with the internet largely through agents acting on their behalf, is now technically feasible. However, many of the consumer and social benefits that could be realized by online AI agents acting scrupulously in their principals' interest are currently obstructed by outdated laws, terms of service, and other less formal practices which allow online platforms to block and degrade agent access, often in secret. Few distinctions are currently drawn between "malicious bots" and AI agents acting with the express delegated authority of a user. For the agentic web to realize its promise, it needs not only the technical infrastructure of protocols and interfaces, but the normative infrastructure of a broadly-accepted and socially-beneficial set of laws, norms and practices governing agentic access to online properties. Building that normative infrastructure requires a society-wide conversation. This pape

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Post-Deployment Accountability in AI Governance: A Cross-Regulatory Empirical Analysis of AI Incidents

arXiv:2605.16281v2 Announce Type: replace Abstract: Post-deployment accountability has become central to AI governance, yet little empirical evidence shows whether monitoring, incident reporting, and impact assessment obligations are visible when AI systems fail. This study analyzes real-world AI incidents from the AI Incident Database (2020--2026) and codes them against nine post-deployment provisions from the EU AI Act, the NIST AI Risk Management Framework, and the GDPR. The findings show substantial accountability gaps: 77.1\% of incidents lack evidence of EU AI Act post-market monitoring, and 99.6\% lack documented Data-Protection Impact Assessment evidence. Governance gaps are also systemic, with 9.8\% of incidents simultaneously non-compliant under two or more regimes. Incidents detected through internal monitoring show much higher compliance than externally detected incidents (87.5\% vs 5.3\% under the EU AI Act; 95.8\% vs 58.1\% under NIST), suggesting that monitoring capacity

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI Alignment Amplifies the Role of Race, Gender, and Disability in Hiring Decisions

arXiv:2605.13866v2 Announce Type: replace Abstract: Humans increasingly delegate consequential decisions to language models, yet whether these systems reproduce or reshape human patterns of discrimination remains unclear. Here, across 29 models and 177 occupations covering nearly half of U.S. employment, we show that language models incorporate demographics into hiring decisions, advantaging female and Black candidates while penalising disabled candidates, with effect sizes comparable to six months to one year of additional education. While pre-trained models show small demographic effects, post-training alignment, which adapts models to human norms and preferences, amplifies advantages for female and Black candidates by 396% and 413% and worsens the disability penalty by 152%. Compared with human employers in past correspondence experiments, language models reverse racial discrimination, substantially attenuate the disability penalty, and amplify the female advantage. Investigating th

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Efficiency Costs of Information Assurance in AI-Enabled Labor Markets: Evidence from LinkedIn's Policy Changes

arXiv:2511.01923v2 Announce Type: replace Abstract: Generative artificial intelligence (GenAI) systems rely heavily on user-generated data for training. As governments and platforms impose increasing restrictions on the use of personal data, an important question is whether limiting access to user data for AI training affects the performance of AI-enabled economic systems. We examine this question in the context of labor-market matching. Our setting exploits a unique sequence of LinkedIn policy changes: the quiet introduction of user data collection for AI training in August 2024, the restriction of Hong Kong user data from AI training in October 2024, and the subsequent restoration of data access in November 2025. Using employment and job-posting data from Revelio Labs and a Difference-in-Differences design comparing Hong Kong and Singapore, we find that the restriction significantly increased labor-market frictions: employee turnover increased and tenure declined, vacancies remained

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Early Adoption of Agentic Coding Tools by GitHub Projects

arXiv:2607.14037v1 Announce Type: cross Abstract: Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While prior studies have examined PR-level outcomes of agent-generated contributions, less is known about how agentic coding tools are adopted and managed at the project level. In this paper, we analyze 25,264 agentic PRs from 2,361 popular GitHub repositories to investigate (1) the adoption of agentic coding tools, (2) project-level agentic PR productivity, and (3) human-agent collaboration patterns. Our results show that the median repository generates only one to two agentic PRs during a three-month period, indicating that intensive adoption remains concentrated in a small subset of projects. At the same time, small projects (1-5 contributors) exhibit higher participation ratios and average levels of agentic PR activity than medium-sized and la

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Epidemic Informatics and Control: A Holistic Approach from System Informatics to Epidemic Response and Risk Management in Public Health

arXiv:2607.13914v1 Announce Type: cross Abstract: This paper presents a holistic systems informatics approach, i.e., Define, Measure, Analyze, Improve, and Control (DMAIC), for epidemic response and management through the intensive use of data, statistics and optimization. Despite the sustained successes of system informatics in a variety of established industries such as manufacturing, logistics, services and beyond, there is a dearth of concentrated review and application of the data-driven DMAIC approach in the context of epidemic outbreaks. First, we define specific challenges posed by epidemic outbreaks to populational health, health systems, as well as economic challenges to different industries such as retailing, education and manufacturing. Second, we present a review of medical testing and statistical sampling methods for data collection, as well as existing efforts in data management and data visualization. Third, we discuss the importance to realizing the full potential of d

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Design of policy digital twins incorporating multi-level agent based modelling

arXiv:2607.13766v1 Announce Type: cross Abstract: Digital twins are used across many industries to enable better decision making. However, while policy makers at all levels (including city, national and supranational scales) have expressed a desire to integrate digital twins into their workflows, this adoption has been slow to materialise. In this paper, we discuss the key issues associated with policy digital twins, and the ways in which they differ from, and are similar to, their counterparts in other areas. We describe how multi-level agent based modelling can be used within policy digital twins to include the effects of human behaviours on outcomes; an aspect that is often largely overlooked. We also describe how digital twins can be designed for policy use cases, and present as a case study the design of a policy digital twin incorporating multi-level agent based modelling to aid a UK city council (local authority) in delivering energy transition policy. After describing both the

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Extending Liquid Rank Toward Multi-Source Reputation Aggregation

arXiv:2607.13615v1 Announce Type: cross Abstract: In this paper, we present an extension of liquid rank reputation systems that enables the aggregation and blending of multiple heterogeneous reputation sources into a unified reputation score. The proposed framework supports the incorporation of external reputational signals alongside internally generated reputation, allowing influence to reflect participation and contribution across multiple contexts and subsystems. By introducing explicit weighting and blending mechanisms, the model provides fine-grained control over the relative impact of individual reputation sources, making it adaptable to diverse governance and coordination scenarios involving both human and machine agents. The resulting approach extends existing liquid rank systems and offers a flexible foundation for designing reputation-based governance mechanisms in complex socio-technical environments.

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized

arXiv:2607.13562v1 Announce Type: cross Abstract: Knowing when to say "I don't know" is fundamental to human judgment, yet AI assistants offer a fluent answer to almost any question. In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions and could always decline to respond. We engineered the questions so that AI advice was wrong, separating AI use from its accuracy. Merely having access to AI nearly eliminated participants' willingness to suspend judgment, and this held whether the advice was actively requested or simply displayed. Consequently, participants answered more questions but were correct about a third as often as when AI was unavailable-yet their confidence nearly doubled. Incentivizing accuracy and penalizing inaccuracy led participants to seek and follow AI advice less, answer more accurately, and suspend judgment more often, though still far less than when AI was unavailable. As AI suggestions grow ubiquitous

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

When Rubrics Change: Cross-Rubric Generalization for Critical Thinking Essay Scoring

arXiv:2607.13433v1 Announce Type: cross Abstract: Automated essay scoring (AES) research has largely focused on cross-prompt generalization, where essays from unseen prompts are scored while the scoring criteria are typically held constant. In practice, however, educators may revise or even introduce new rubrics in their scoring task, to evaluate different aspects of essays. We study cross-rubric generalization: training on essays labeled under one set of rubrics and evaluating on previously unseen rubrics, which target different aspects of the essay. We use a Large Language Model (LLM) fine-tuning framework with two components: rubric-agnostic intermediate representations, called traits, and target-essay supervision under seen rubrics during training. On an AES dataset augmented with multiple rubric-defined labels of student critical thinking skills, we find that traits improve macro F1 by 5.0% over a baseline without traits in the hardest setting, where both target rubrics and target

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

xChk: Bring Your Own Identity -- Heterogeneous Assurance with Verifier-Determined Sufficiency

arXiv:2607.13369v1 Announce Type: cross Abstract: We present xChk, a reference identity provider for Bring Your Own Identity (BYOI): users enroll via heterogeneous proofs (government KYC, corporate SSO, WebAuthn/FIDO2, professional networks, live verification, longitudinal activity, behavioral signals) and disclose them as portfolio claims in standard OAuth 2.0 / OpenID Connect (OIDC) tokens, while each relying party applies its own sufficiency policy - the IdP transports claims and may evaluate an RP-supplied evidence policy for consent, but does not adjudicate access. Enrollment depth varies by modality (some paths are user-initiated; org KYB and officer binding are operator-assisted). xChk also supports human-in-the-loop attestation for high-risk actions: humans can initiate attestations directly (browser UI / POST /api/attestations), and AI agents acting under those principals can trigger the same gateway via scope-gated authorize/attest - hash-chained human approvals on a shared v

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation

arXiv:2607.13230v1 Announce Type: cross Abstract: Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interact with third-party services. This paper develops an AI-native mathematical framework for underwriting, pricing, and contract design for agentic AI deployments. A deployment is represented by a risk state that captures autonomy level, operational authority, permission exposure, governance maturity, and dependency concentration. The framework maps the risk state to event probabilities, loss severities, governance costs, premiums, deductibles, coverage allocation, and policy covenants, and formulates an optimization problem for insurance contract design under participation, profitability, and incentive compatibility constraints. The paper establishes structural properties of insurability, including characterization of an insurability region, monotone deterioration of feasibility with increa

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Beyond AI-Generated Labels: Watermarking, Co-Creation, and Conflation of AI-Generation with Disinformation

arXiv:2607.13082v1 Announce Type: cross Abstract: Watermarking is often presented as a straightforward solution for distinguishing AI-generated from human-generated content, enabling platforms and regulators to trace synthetic content and detect AI-generated outputs at scale. This paper examines whether such mechanisms meaningfully address the epistemic and ethical challenges that arise in domains where the central concern is not the automation of content production, but the accuracy, intent, and deceptive potential of messages. We argue that extending watermark-based approaches to these settings is conceptually and practically misguided. Invisible watermarking encodes only model origin; when operationalized into visible AI-generated labels, it reduces complex creative processes to a misleading binary and provides no information about truthfulness. Such labels may stigmatize legitimate uses of generative tools while encouraging misplaced trust in unmarked content. Here we propose an al

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI-Augmented Human Resource Management? Insights from German companies

arXiv:2607.13839v1 Announce Type: new Abstract: This study examines the integration of AI into Human Resource Management in German companies. We ask if and how AI-based technologies are \enquote{augmenting} human resource management. Organisations employ generative AI or predictive analytics to transform traditional human resource functions, to streamline routine tasks and to reallocate resources toward strategic, people-centred activities. Our findings from interviews and group discussions and a survey (N=410) reveal that while AI tools enhance HR analytics capabilities, their adoption mainly serves efficiency and rationalising goals. The introduction of AI tools is shaped by organisational transformation factors such as digital infrastructure, co-determination frameworks, and ethical implications. The research highlights both the strategic potential for improved talent development and the challenges posed by data governance and algorithmic transparency. Overall, this work contributes

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation

arXiv:2607.13798v1 Announce Type: new Abstract: Generative AI tools are increasingly being piloted in public agencies, but limited evidence explains how employee acceptance changes after hands-on use. This study examines Microsoft 365 Copilot adoption during an eight-week pilot at a state Department of Transportation. A matched two-wave survey measured perceived usefulness, perceived ease of use, behavioral intention, and trust before and after participation. After matching and response-quality screening, the sample included 124 employees. Nonparametric tests assessed aggregate changes, k-means clustering identified baseline acceptance personas, and fixed-centroid assignment tracked migration. Open-ended responses were examined using keyword-based content mapping. Perceived usefulness declined significantly after use, suggesting recalibration of expectations, while perceived ease of use, behavioral intention, and trust showed only small, nonsignificant changes. Three baseline personas

Source ↗
technology Thu, 16 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Environmental Cost of Digital Sovereignty: Water, Energy, and Emissions Impacts of Sovereign AI Infrastructure in the Global South

arXiv:2607.13443v1 Announce Type: new Abstract: Sovereign AI has become a strategic priority across the Global South, with over \$200 billion in state-led commitments announced between 2024 and 2026. Yet the physical infrastructure that compute sovereignty demands, above all data centers, imposes water, energy, and carbon costs that fall hardest on countries least equipped to absorb them. This paper presents a comparative environmental stress analysis across four cases: the United Arab Emirates, Bangladesh, India, and Africa (with a focus on Kenya). Using publicly available water stress data, grid carbon intensity factors, and GPU power specifications, we model the water consumption, energy demand, and carbon emissions of hypothetical sovereign AI deployments under multiple cooling technology scenarios. We find that a 1,024-GPU cluster using evaporative cooling in the UAE would consume over 30 million liters of water annually in a country classified as ``extremely high'' water stress.

Source ↗
Showing 6551–6600 of 10879 signals
← Prev Page 132 of 218 Next →