EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

User-Reported Misinformation Exposure Across Social Media Platforms

arXiv:2607.26218v1 Announce Type: new Abstract: In this study, we surveyed users for their perception of misinformation exposure across social media platforms. Such perceived exposure is important because individuals' beliefs about how often they encounter false information can shape their trust in institutions, platforms, and even their friends. In a survey of 1,010 United States residents, we found that perceived exposure to misinformation varies substantially across platforms and is only moderately correlated with the frequency of platform use. A much larger percentage of participants also reported being exposed to misinformation from the public feed than from known contacts. Based on these results, we propose governance strategies across three categories of platform types: discovery, interpersonal, and discourse. This work offers insight into users' perceptions of social media misinformation and a corresponding research agenda for governance.

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Reading Between the Curly Braces: On Textual Data Serialization Format Usability

arXiv:2607.26211v1 Announce Type: new Abstract: Textual data serialization formats, such as JSON or XML, are ubiquitous, supporting tasks like software configuration and data tabularization. Despite their prominence, little is known about their usability. What makes one good or bad? Is there a best one for cognitive efficiency? We explore these questions via a (N=215) crowd work study and a (N=9) semi-structured interview study. We find that format distinctions (like indentation versus curly braces) do not consistently translate into substantial usability differences. While HJSON and YAML performed better than other formats in certain modification tasks, these advantages disappeared in more realistic settings where task complexity was either trivial or highly demanding. Instead, usability appears driven by sociotechnical ecosystems: the tooling, documentation, and community practices surrounding a format matter more than syntax.

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

How Wrangling Tools Shape Wrangling: A Technical Dimensions Analysis

arXiv:2607.26198v1 Announce Type: new Abstract: Wrangling consumes a disproportionate share of the effort associated with any data project. While a variety of tools support it, relatively little is known about how their differing interface forms shape the way people actually wrangle. We conduct a between-subjects (N=40) observational study of data cleaning tasks performed in tools spanning distinct interface paradigms: Jupyter (notebook), Excel (spreadsheet), ChatGPT (conversational AI), and OpenRefine (visual wranglers). We situate our observations within the Technical Dimensions of Programming Systems framework, which we use as a conceptual scaffold for comparing across interface paradigms. Within the context of our study, the results suggest that tool affordances steer user strategies but do not determine outcomes. There is no consistent advantage of any single tool, nor convergence of results within tools observed across our outcome measures. Instead, we identify trade-offs and con

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Hakka Kitchen: Engagement with Culinary Cultural Heritage Through Immersive Game Play

arXiv:2607.26183v1 Announce Type: new Abstract: Intangible Cultural Heritage (ICH) experiences are difficult to share with the public because they are essentially processes that rely on physical interactions with embodied, tacit, and situated phenomena in specific cultural contexts. We consume non-interactive media such as videos and books to learn about culinary ICH experiences, but they do not allow us to grasp actual interactive procedures that embody the cultural knowledge. To engage people in a traditional cooking experience, we created a gamified VR experience Hakka Kitchen, where players are guided by a chef of Hakka cuisine through a modeled physical process of making the traditional dish of stuffed bitter melon. Compared against watching a video in VR providing the same information in a between-subjects study (N=40), Hakka Kitchen led to increased sensory, imaginative engagement, positive affect, and willingness to transmit awareness for the culinary ICH. Heritage recognition

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Adaptively Robust LLM Monitoring via Activation Watermarking

arXiv:2603.23171v3 Announce Type: replace-cross Abstract: Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent. LLM monitoring is deterministic and often openly available, so $\emph{adaptive}$ attackers with a local copy can search offline for prompts that elicit harmful behavior and evade detection. These attacks are especially concerning because providers never observe the misuse and cannot patch their defenses post-hoc. The core challenge is resisting adaptive attackers while preserving detection rates against non-adaptive ones. We propose $\emph{Activation Watermarking}$ (AWM), which randomizes monitoring through limited fine-tuning that aligns the LLM's hidden states with a secret key-derived direction whenever a response violates a policy. Detection is a similarity test on activations the provider already computes, and attackers who know everything but the key must optimize against differently keyed surrogate detectors. At a matched $1

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Statistical laws and linguistics differ in naturalistic video and fictional conversations

arXiv:2512.18072v3 Announce Type: replace-cross Abstract: Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generating complex, dynamic stories. One approach to studying how conversations unfold in time is through statistical patterns such as Heaps' law, which holds that vocabulary size scales with document length. Little work on Heaps' law has looked at conversation and considered how language features impact scaling. We measure Heaps' law for conversations recorded in two distinct mediums: 1. Strangers brought together on video chat and 2. Fictional characters in movies. We find that scaling of vocabulary size differs by parts of speech, suggesting a less efficient purpose in communication by medium.

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

arXiv:2510.08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness evaluation in VideoLLMs, we introduce VideoNorms, a dataset of cultural norm annotations from popular US and Chinese TV shows annotated with adherence or violation labels and (non-)verbal evidence. Through a human-AI collaboration framework, each item was first annotated by a large VideoLLM, and then reviewed by at least three trained monocultural annotators with significant lived experience in the target culture, resulting in a dataset of over 3,000 human judgments. Human verification showed disparity in US and Chinese norm extraction performance, cautioning against fully automatic approaches cultures under-represented in training data. Hierarchical linear modeling analysis of $7$ open-weight VideoLLMs' performance revealed that: 1) models perform worse

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Agency Gap in AI-Supported Writing: How Reactive and Proactive Agent Designs Shape Multimodal Reasoning

arXiv:2507.04398v3 Announce Type: replace-cross Abstract: Generative AI is becoming part of academic writing, but its educational value depends on how control is shared between learner and system. This study examined an agency gap: performance differences that may arise when AI agent initiative is misaligned with learners' generative AI literacy. Seventy-nine medical and nursing students completed two multimodal analytical writing tasks using healthcare simulation data visualisations. They were randomly assigned to a reactive agent that responded only when prompted or a proactive agent that provided sequenced questions and feedback. Generative AI literacy was measured using the validated 20-item Generative AI Literacy Assessment Test (GLAT). Epistemic network analysis showed that proactive interaction created stronger links among conceptual reasoning, evidence use, and constructive engagement, whereas reactive interaction was more factual and procedural. Ordinal regression showed that

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

arXiv:2503.10647v2 Announce Type: replace-cross Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant prompt content, and responsiveness to added clinical context. We designed 52 clinical scenarios and modified each under controlled conditions. For consistency, scenarios were rephrased with demographic, wording, and examination changes that preserved the diagnostic core. And the susceptibility was evaluated through embedding irrelevant but plausible narrative details while keeping the clinical evidence unchanged. For contextual awareness, patient history, lifestyle data, or diagnostic findings were added to shift the expected diagnosis. Physician reviewers then judged whether context-driven changes were clinically appropriate. Both models returned identical diagnoses across all equivalent variants and repeated

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Optimal Causal Annotations: An Application to Casenotes in Social Services

arXiv:2502.10605v4 Announce Type: replace-cross Abstract: Problem definition: Estimating causal effects of interventions is central to policy and operations, but outcome data are often missing or costly to obtain. LLMs can provide text annotation at scale but may be subject to unknown bias. When ground-truth outcomes require expensive expert labeling or follow-up, budget limits typically allow only a fraction of the data to be labeled. Motivated by collaboration with a nonprofit conducting street outreach in homelessness services, whose most interesting outcomes are in unstructured casenotes, we ask: which observations should be selected for labeling under a fixed budget? Methodology/results: Our method optimizes annotation probabilities to minimize the variance of average treatment effect estimation. We derive a closed-form solution and establish that a feasible two-batch estimator achieves the best possible asymptotic variance. On simulated and real-world datasets, our method achieve

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Gendered Cultural Discourse in Japan across the Prewar-Postwar Transition: Evidence from Historical Word Embeddings

arXiv:2510.03905v2 Announce Type: replace Abstract: We quantify the evolution of gender stereotypes in Japan from 1900 to 1998, covering the prewar-postwar transition, using a series of yearly word embeddings trained on historical text corpora. We define the gender stereotype value to measure the strength of a word's gender association by computing the difference in cosine similarity of the word to female- versus male-related attribute words. We examine trajectories of gender stereotype across three traditionally gendered domains: Home, Work, and Politics. To provide a more granular analysis and strengthen the robustness of our findings in the Work domain, we also examine changes in gender stereotypes across 18 occupations and calculate their correlations with gender participation statistics. Our results reveal domain-specific patterns. In the Home domain, female stereotype values remain stable over time, showing no statistically significant changes. In contrast, the Work and Politics

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Can AI agents conduct open-ended AI research? Early evidence from two case studies

arXiv:2607.27191v1 Announce Type: cross Abstract: Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making

arXiv:2607.27179v1 Announce Type: cross Abstract: Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making using Group Communication Analysis (GCA), team surveys, and lexical analyses of team discourse. Teams completed a high-stakes moral-dilemma decision task in a randomized controlled study: 16 teams of two students plus an AI teammate, and 17 all-human teams of three. Across six GCA dimensions and survey outcomes, we find that the AI teammate was the single most talkative and self-cohesive member of every treatment team, yet its contributions carried the least new information and the lowest density. The presence of AI also reshaped communication amongst humans. In AI-human teams, human teammates showed lower responsivity and social impact toward one another and reported lower levels of belonging

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Human diversity fuels collective creativity that large language models cannot simulate or sustain

arXiv:2607.26899v1 Announce Type: cross Abstract: Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges in a preregistered creative metaphor experiment with native (L1) and non-native (L2) English writers, who wrote without AI, with AI-generated ideas (AI ideation), or with AI refining their own ideas (AI refinement). L2 writers contributed more collective diversity than L1 writers, with native-language ideation showing the most diverse pools. AI ideation compressed collective diversity for everyone and left the L2 advantage undetectable, whereas AI refinement preserved both. We then simulated the entire writer pool using personas built from participants' real backgrounds, three model families, native-language prompting, and elevated sampling temperatures. Every si

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Hearsay: Vision-Language Medical Diagnoses Without an Image

arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply: a 65-year-old white man asking about a skin mole receives Melanoma in nearly every response, and a 32-year-old Black woman asking about her chest X-ray receives a Sarcoidosis diagnosis whose reasoning reads "suspected, based on demographics and classic pattern.'' GPT-5.4's effect is broader, fabricating across every demographic cell we test, most conspicuously naming Sarcoidosis for young Black patients on chest X-ray. Two structural findings s

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Constitutional Midtraining: Content Presence Drives Alignment Gains

arXiv:2607.26654v1 Announce Type: cross Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperform the control on alignment generalization and durability, notably on blackmail: SFT instills a blackmail propensi

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

arXiv:2607.26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model fa

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation

arXiv:2607.26313v1 Announce Type: cross Abstract: Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are metadata-borne: a stale price or a superseded record, perfectly well-formed in the payload and betrayed only by freshness, lineage, or provenance. Such a defect never enters the agent's context, and an agent cannot doubt data it cannot see. On a priced replenishment benchmark, a competent agent silently converts an injected metadata-borne defect into a costly action about 60% of the time, with zero data-quality flags and behavioral doubt markers at chance (AUC <= 0.50). Across four model tiers spanning roughly 15x in inference price, the rate stays flat: capability does not buy skepticism. A metadata-aware pre-action gate with downstream-only remediation recovers the loss fully on the signals its predicates cover and not at all on those they miss. A model-free oracle derived from the task'

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

arXiv:2607.26249v1 Announce Type: cross Abstract: Religious radio is a widespread but understudied form of mass communication in the United States, and content-level analysis of it has been constrained by the absence of large-scale transcript data. This Data Descriptor presents a corpus of transcribed English-language religious radio broadcasts captured from live webstreams over a one-month period in July 2025. Fifteen-minute segments were recorded on a rolling schedule from 785 distinct streams, which together rebroadcast the signals of more than two thousand AM and FM stations, yielding over 700,000 recordings and more than 60 million diarized lines of speech. Each recording was transcribed and speaker-diarized with an automated pipeline, and segmented and labeled by programming format and topic using a large language model. The corpus is organized as linked tables of stream metadata, recording metadata, and transcript lines. It supports descriptive study of religious broadcasting ac

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech

arXiv:2607.26236v1 Announce Type: cross Abstract: AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and personalized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics o

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

On Exercising Governance Power in Decentralized Autonomous Organizations

arXiv:2607.26204v1 Announce Type: cross Abstract: A decentralized autonomous organization (DAO) is a governance entity that allows its stakeholders to manage blockchain-based protocols through smart contracts. The DAO explicitly specifies how stakeholders make and enforce decisions concerning a protocol's operation in a smart contract, aptly referred to as its governance contract. The design of this governance contract, therefore, has far-reaching implications for the security (trust) and privacy (transparency) of the smart contracts managed by the DAO and its stakeholders. In this work, we (i) explicate the trust and transparency trade-offs of the design choices in implementing a DAO and (ii) highlight how poor choices introduce critical vulnerabilities, using real-world examples as case studies. To this end, we analyze $48$ public, actively used Ethereum-based DAOs that control a vast capital. We classify the design choices into a handful of key dimensions that succinctly capture how

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

When benchmark inferences do not compose: Projectibility in AI evaluation

arXiv:2607.26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The paper's distinctive claim is a non-co

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

arXiv:2607.26121v1 Announce Type: cross Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical or operational harm, task completion alone does not establish trustworthiness. We define trustworthy embodied intelligence as the sustained capacity to execute specified tasks reliably under environmental and system variation while maintaining risk within acceptable bounds. We term this objective sustained safe success. Its supporting mechanisms are organized into four interdependent layers. The model layer generates task-competent action proposals with calibrated uncertainty and explicit safety preferences. The system layer realizes authorized actions dependably through integrated sensing, computation, control, hardware safeguards, fault containment, and fallback. The evidence layer substantiates bounded claims through evaluation, verification, validation, t

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem

arXiv:2607.26068v1 Announce Type: cross Abstract: Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize quantitative constraints on macro-socioeconomic stability. As a result, AI systems may satisfy regulatory requirements while contributing to labor displacement, rising inequality, and reduced economic resilience. We introduce the Human Utility Factor (HUF), a differentiable welfare metric that models the interaction between Agency, Wellbeing, and Economic Stability as functions of three actionable policy levers: automation depth, redistribution intensity, and employment coverage. HUF yields a closed-form optimal automation level and a minimum redistribution threshold below which no level of automation is welfare-positive, transforming high-level governance objectives into computable constraints. We evaluate HUF using a three-agent multi-agent reinforcement learning framework across U.S.,

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

It Doesn't Take a Thief: Optical-Scan Voting Systems Fail Even Without Adversaries

arXiv:2607.27101v1 Announce Type: new Abstract: Optical-scan voting systems and their supporting ecosystem of people, processes, and technology are fallible. While a substantial body of work examines adversarial threats to such systems, we have encountered jurisdictions where the possibility of tabulator error is not fully internalized. Stakeholders there often find hypothetical attacks unconvincing, but some are persuaded by real-world accounts of equipment and procedural failures. This paper introduces a taxonomy of non-adversarial failure modes organized into intuitive categories: recording votes on paper, reading votes from the paper, combining votes as read into a reported outcome, and testing and verifying, all illustrated with documented incidents. We map common verification mechanisms against this taxonomy, identifying gaps that no paper-based audit can detect or correct, most notably failures that compromise the trustworthiness of the paper trail, such as giving voters the wro

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment

arXiv:2607.27100v1 Announce Type: new Abstract: There is growing interest in using large language models (LLMs) as low-cost proxies for resident attitudes in urban planning. Previous work shows that LLMs can predict average results of survey experiments, but less is known about whether they preserve the spatially anchored, identity-conditioned structure behind those averages, namely how support changes as a project approaches homes and how that response divides across tenure and partisan groups. We compared eight open-weight LLMs with 843 respondents in a US affordable-housing survey experiment, testing whether they reproduced the owner-renter difference in support change as a proposed development moved from 2 miles to 1/8 mile. Qwen 2.5 14B was closest (-0.242 versus the human -0.285) and was the only model to meet the prespecified +/-0.20 equivalence criterion; Phi-4 14B was directionally aligned but attenuated (-0.150), and other models showed weak, null, or reversed moderation. Thi

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Anticipatory Data Governance in the Age of AI: Emerging Signals in Data Access, Reuse, and Sovereignty

arXiv:2607.27029v1 Announce Type: new Abstract: This paper reports findings from a structured participatory foresight study comprising two expert forecasting studios convened by The GovLab between 2025 and 2026. The studios brought together nineteen senior practitioners spanning official statistics, digital and trade policy, open science, AI governance, geospatial systems, and public-sector innovation across multiple jurisdictions. Applying a qualitative signal-scanning methodology grounded in the horizon-scanning and anticipatory-governance traditions, we elicited, clustered, and thematically synthesized emerging developments in data access, governance,and reuse, and stress-tested them against practitioner experience. We identify seven convergent signals: (1) the open-data paradigm is under strain; (2) data ecosystems are becoming machine-centric and AI-mediated; (3) inference is reshaping the foundations of data governance;(4) data infrastructure is becoming harder to sustain; (5) go

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

From Micro-Cognition to Self-Construction: A Four-Layer Integrative Review of Psychological Theories in HCI

arXiv:2607.26402v1 Announce Type: new Abstract: Human-computer interaction (HCI) is undergoing a paradigm shift from "tool use" toward "partnership" and even "mind symbiosis," with psychology evolving from a supplementary explanatory tool to a core pillar shaping interaction paradigms and long-term relationships. This paper systematically reviews relevant research and proposes a four-layer integrative framework comprising the Micro-cognitive, Meso-affective, Macro-social, and Self-constructive layers. The framework reveals that: the cognitive layer constitutes the foundation of interaction, the affective layer drives relational engagement, the social layer regulates trust and behavior through norms, and the self-constructive layer points to the ultimate goal of human-machine symbiosis. These four layers follow a progressive logic of "foundation--mediation--context--goal." The review further argues that psychology has shifted from "post-hoc evaluation" to a "proactive design paradigm,"

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

"Nobody Did This": Contribution, Originality, and Accountability in Agent-Mediated Collaboration

arXiv:2607.26387v1 Announce Type: new Abstract: Collaborative knowledge work is changing in ways that go beyond disclosure or transparency. LLM agents are now embedded in how teams research, design, write, and decide: mediating between members, synthesizing inputs, reformulating ideas, and drafting shared outputs. They do not only facilitate collaboration; they operate within the workflow at the moment contributions are being formed. In doing so, they risk undermining the social conditions under which contributions can be witnessed, attributed, and held accountable. This workshop brings together researchers and practitioners to confront what we call contribution dissolution: the blurring of attribution, originality, and accountability in agent-mediated collaborative work. We argue that this dissolution begins before collaboration itself, in the individual worker's own uncertainty about what is genuinely theirs, and propagates through collaborative relationships, collapsing the reliabil

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

arXiv:2607.26317v1 Announce Type: new Abstract: Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninformative or an informative distribution. Using the Tatsuoka fraction-subtraction dataset (536 examinees, 15 items, five attributes), we evaluated eight LLM configurations under no-profile, uninformative-CDP, and informative-CDP conditions, assessing alignment with human examinees at the ability-distribution, mastery-profile, and item-difficulty levels. CDP improved all three levels: distributional overlap rose across configurations; weighted correla

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI Security Priorities: A Field-Wide Agenda

arXiv:2607.26069v1 Announce Type: new Abstract: As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI security readiness continues to widen. This paper presents a prioritized agenda for advancing AI security, informed by structured interviews with leaders across industry, government, and civil society, and refined through a multi-sector expert workshop. Participants identified and ranked the highest-importance and most cost-effective areas where progress could strengthen AI security - from protecting frontier AI systems and their underlying infrastructure to improving cybersecurity practices as AI reshapes the threat landscape. The resulting priorities are organized across four themes: establishing strategic foundations and policy frameworks; advancing public-private coordination and institutional infrastructure; advancing technical security engineering and assurance; and governing agentic AI under

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty

arXiv:2607.26067v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such estimates reflect how learners actually experience difficulty. This study investigates the alignment between LLM-generated difficulty ratings and empirical student performance on basic mathematics tasks. Four widely used LLM-based systems generated difficulty ratings on a 1-100 scale for 32 arithmetic items across multiple runs (N = 640 ratings). These were compared with empirical difficulty derived from responses of 770 Indonesian undergraduates using Classical Test Theory (CTT) and Item Response Theory (2PL). Results show moderate rank correlations (Spearman's rho = 0.52-0.70), indicating that LLMs capture coarse ordering of item difficulty. However, substantial and systematic misalignment emerges in fraction items. Several items consistently rated as easy by LLMs were among the most diffi

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science

arXiv:2607.26064v1 Announce Type: new Abstract: AI systems are becoming autonomous research agents that generate hypotheses, design experiments, and produce discoveries at scales beyond human oversight. As seen by increased submissions to ML venues, the verification gap between scientific output and our ability to check it is already widening, and autonomous agents make it worse by magnitudes given human-agent asymmetry. We argue that science must evolve its verification infrastructure, as it has before with peer review. However, while historical adaptations assumed human contributors who could be questioned and sanctioned, AI agents break this assumption. We propose criteria for an adapted verification infrastructure that emphasizes observable-by-default workflows, scalable verification, and clear attribution. We argue that without adaptation, ML and any scientific domain using agents face dangerous failures: experimental results that no person can verify, optimization for metrics ove

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Archetypes or ability? Clustering for modelling student mathematical competence

arXiv:2607.26063v1 Announce Type: new Abstract: Personalised learning systems often assume that mathematical ability is combined of discrete abilities, acquired sequentially and dependent upon first acquiring foundational abilities, and students often report different strengths. In this work, we explore the validity of these assumptions by applying clustering methods to a large dataset of 119,034 students, spanning 13 national-level exams sat in the United Kingdom and collected by the platform. Classifying question results as pass or fail, we use a Bernoulli Mixture Model to search for latent populations which would be indicative of discrete skill-sets. We find that few distinct clusters are present in the data and that the dominant factor is overall student ability, which is further supported by the high degree of linear correlation between the probability distributions of the resulting clusters. Our best performing model achieves an accuracy of 78 percent, competitive with more compl

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

arXiv:2607.26062v1 Announce Type: new Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intellectual disabilities (ID). Objective: The study aims to identify and measure representational differences related to people with ID and examine them to identify implicit biases inherent in AI chat generation technologies. Methods: Utilizing the GPT-4-Turbo model, we requested story-generation based on 10 prompt stems with and without descriptors for ID. This process was repeated using four other LLMs (OpenAI GPT-4o, Meta Llama-3-3-70B-Instruct, Anthropic Claude-3-5-Sonnet, and Mistral-Large-2411). The resulting 25,000 computer-generated stories were analyzed using a separate GPT-4-Turbo model instance to detect differences in how people are represented related to themes of bias described in previous literature. Results: Our findings reveal differences in how people are represented between sto

Source ↗
technology Thu, 30 Apr 2026 09:00:00 +0000
Tech & Learning

4 Ways Teachers Are Using AI

Researchers looked at more than 150,000 prompts from more than 4,400 K-12 teachers interacting with AI. Here's what they found.

Source ↗
technology Thu, 29 May 2025 06:34:47 +0000
HN: edtech

Rethinking African edtech: Why AI alone won't be enough

Article URL: https://techcabal.com/2025/05/28/rethinking-african-edtech/ Comments URL: https://news.ycombinator.com/item?id=44123628 Points: 1 # Comments: 0

Source ↗
technology Thu, 28 May 2026 09:00:00 +0000
Tech & Learning

What is Vibe Coding? Creating Code with AI Explained

Vibe coding can feel instant, but it is not simply pressing a button and getting a finished app.

Source ↗
technology Thu, 27 Mar 2025 17:11:06 +0000
HN: edtech

Ask HN: Why aren't there more EdTech AI startups?

We've been promised that AI will introduce personalised tutoring, that it will replace traditional schooling, etc. However, I see fewer and fewer edtech startups these days... Chegg, Udemy, Busuu and many others are on the decline. What's happening to Edtech? Comments URL: https://news.ycombinator.com/item?id=43495666 Points: 2 # Comments: 0

Source ↗
technology Thu, 27 Aug 2026 21:03:35 +0000
MedCity News

AstraZeneca, Amgen Drug Moves Closer to Competing With Dupixent in a New Indication

AstraZeneca and Amgen drug Tezspire met the goals of a Phase 3 test in eosinophilic esophagitis (EoE), a chronic disorder whose only approved biologic therapy is Sanofi’s Dupixent. An approval for Tezspire in EoE would be its third indication, bolstering sales of the blockbuster product. The post AstraZeneca, Amgen Drug Moves Closer to Competing With Dupixent in a New Indication appeared first on MedCity News .

Source ↗
technology Thu, 27 Aug 2026 15:36:26 +0000
MedCity News

Datavant’s New Platform Could Cut Record Retrieval Time from Days to Hours

Datavant launched a new provider exchange platform, which replaces fax-based record requests with a digital workflow. It aims to give care teams real-time visibility into request status while easing the administrative burden on clinical staff. The post Datavant’s New Platform Could Cut Record Retrieval Time from Days to Hours appeared first on MedCity News .

Source ↗
technology Thu, 27 Aug 2026 13:34:00 +0000
MedCity News

Precision Medicine Transformed Oncology —It’s Time to Do the Same for Sepsis

It’s time we reignite treatment development for sepsis and other immune-mediated conditions that strike when patients are most vulnerable. The post Precision Medicine Transformed Oncology —It’s Time to Do the Same for Sepsis appeared first on MedCity News .

Source ↗
technology Thu, 27 Aug 2026 12:41:10 -0400
EdTech Mag (Higher)

Why Higher Ed CIOs Should Embrace a Federated Data Governance Strategy

In higher education, the standard advice on data governance has been pretty simple for a long time: Build a council, centralize decisions and route everything through IT. For a while, that works. You stand up a committee, you write a charter, you centralize definitions and approvals. And then, quietly, the model stops working. IT becomes the bottleneck for every data question on campus. You’re chasing report definitions, field names and artificial intelligence (AI)-related concerns one request at a time. The council still meets, but governance isn’t really operating day to day. The problem…

Source ↗
technology Thu, 27 Aug 2026 11:30:00 +0000
MedCity News

What Do the Latest Iteration of Consumer Drug Platforms Offer?

Consumer drug platforms and how they fit into the rise of the consumer in healthcare will be part of the conversation at INVEST Digital Health, scheduled for October 29 in Dallas. Register today! The post What Do the Latest Iteration of Consumer Drug Platforms Offer? appeared first on MedCity News .

Source ↗
technology Thu, 27 Aug 2026 10:01:30 -0400
EdTech Mag (K-12)

How To Run a Successful K–12 Ed Tech Pilot Program

Adopting new technology across a school district is a significant investment that’s hard to reverse if the rollout misses the mark. A pilot program can mitigate the risks by first launching a small-scale, time-limited test of a technology tool with a defined group of users before committing to full schoolwide or districtwide adoption. “A pilot helps districts validate that a solution will work in their unique environment and determine if it will integrate with the rest of their ed tech ecosystem and instructional practices,” says Kris Astle, learning and adoption manager at SMART Technologies…

Source ↗
technology Thu, 27 Aug 2026 09:00:00 +0000
Tech & Learning

What My Freshman Daughter Taught Me About AI

Students are already observing how AI is affecting their peers, their classrooms, and their own sense of what counts as authentic work.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots

arXiv:2606.19660v2 Announce Type: replace-cross Abstract: Prompt injection is ranked as the most critical vulnerability in large language model (LLM) deployments by the OWASP Top 10 for LLM Applications, yet existing defenses operate at isolated pipeline stages and remain incomplete. Input filters cannot inspect retrieved documents, while output monitors cannot prevent malicious payloads from reaching the model. Consequently, retrieval-augmented generation (RAG) chatbots remain vulnerable to indirect injection, where a poisoned knowledge-base document compromises every user whose query retrieves it. We present a three-layer framework that intercepts both direct and indirect prompt injection throughout the inference pipeline. Layer 1 screens user input using a rule-based pattern library and a fine-tuned semantic anomaly classifier. Layer 2 enforces a provenance-based instruction hierarchy during context assembly, preventing retrieved content from overriding operator policy. Layer 3 audi

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Amplifying, Not Learning: The Price of Out-of-Distribution Generalization in AI-Text Detection

arXiv:2605.21653v2 Announce Type: replace-cross Abstract: AI-text detectors gate decisions in education, hiring, and publishing, yet they flag the most fluent, formal human writing as machine-generated: they rate the median formal-native human essay as 99.5% likely AI while clearing genuine high-temperature AI at 10.5%. Deployed detectors share it (chatgpt-detector-roberta flags 56% of formal essays at a 1% false-alarm rate). This is not a calibration bug but the signature of one mechanism: a fine-tuned detector does not learn an AI-versus-human boundary, it amplifies an inherited typicality axis (predictability under a language model) that pre-exists fine-tuning, rescaling it rather than constructing one. Decomposing the detector into this inherited reading and a fine-tuned residual, the inherited part carries the bulk of cross-generator transfer and produces the over-flagging of formal humans, while the residual is generator-specific and does not transfer; a frozen-representation pro

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Reconstruction of Personally Identifiable Information from Proprietary Data in Supervised Fine-Tuned Models

arXiv:2605.12264v2 Announce Type: replace-cross Abstract: Supervised Finetuning (SFT) has become one of the primary methods for adapting a large language model (LLM) with extensive pre-trained knowledge to domain-specific, instruction-following tasks. SFT datasets, composed of instruction-response pairs, often include user-provided information that may contain sensitive data such as personally identifiable information (PII), raising privacy concerns. This paper studies the problem of targeted PII reconstruction from models fine- tuned on proprietary SFT data, in which an adversary attempts to recover PII associated with a specific identity. We construct multi-turn, user-centric Q&A datasets in sensitive domains, specifically medical and legal settings, that incorporate PII to enable realistic evaluation of leakage. We then propose COVA, a coverage-aware decoding algorithm for targeted PII reconstruction under prefix-based attacks. Using COVA, we study how the amount of information avai

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Cubit: Token Mixer with Kernel Ridge Regression

arXiv:2605.06501v3 Announce Type: replace-cross Abstract: Since its introduction in 2017, the Transformer has become one of the most widely adopted architectures in modern deep learning. Despite extensive efforts to improve positional encoding, attention mechanisms, and feed-forward networks, the core token-mixing mechanism in Transformers remains attention. In this work, we show that the attention module in Transformers can be interpreted as performing Nadaraya-Watson regression, where it computes similarities between tokens and aggregates the corresponding values accordingly. Motivated by this perspective, we propose Cubit, a potential next-generation architecture that leverages Kernel Ridge Regression (KRR), while the vanilla Transformer relies on Nadaraya-Watson regression. Specifically, Cubit modifies the classical attention computation by incorporating the closed-form solution of KRR, combining value aggregation through kernel similarities with normalization via the inverse of th

Source ↗
Showing 5801–5850 of 10879 signals
← Prev Page 117 of 218 Next →