EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CY

Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response

arXiv:2608.14254v1 Announce Type: new Abstract: Fugitive emissions from waste sites increasingly expose communities to toxic and odorous gases, yet public-health responses remain largely retrospective, with episodes investigated only after residents have been exposed. Here we show that the meteorological drivers of elevated hydrogen sulphide (HS) at a long-monitored European landfill, and the timescales over which they act, can be identified directly from routine monitoring data. We introduce CAIRN (Causal-Anchored Inference for Receptor Nowcasting), a machine-learning framework whose internal memory is matched to these measured timescales: a fast component tracking hour-scale wind-borne transport and a slow component tracking multi-hour weather changes. Trained to predict gas measurements, CAIRN operates using only routine weather variables and the calendar, without hand-engineered features. Its behaviour is consistent with the identified transport mechanisms, and the framework transf

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CY

Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation

arXiv:2608.13712v1 Announce Type: new Abstract: Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable legal personas. We use an evaluation framework separating behavioral realism from pedagogical usefulness. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories. WitnessSim generally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony or WitnessSim generated testimony. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluatin

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CY

Language-Specific Gaps in AI Safety Training Datasets

arXiv:2608.13695v1 Announce Type: new Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users. We show that these collection-level coverage claims frequently do not survive inspection at the level of an individual language. Auditing 21 resources across 25 language slices, of which 20 count as datasets under our counting rules, spanning three languages chosen to represent low- (Hausa), mid- (Swahili), and high-resource (French) tiers, we find that gaps in provenance, annotation reliability, access, harm-taxonomy coverage, and data reuse recur in patterns that partially, but not fully, track resource level. Using a controlled within-pipeline comparison, we show a Hausa-language slice falling below its own paper's translation-quality acceptance threshold while the same pipeline's Swahili output clears the same bar comfortably; this is evidence that these

Source ↗
technology Mon, 17 Aug 2026 00:00:00 -0400
arXiv cs.CY

Asymmetric Discourse Homogenization and Shared Language Technology: Evidence from Reddit

arXiv:2608.13674v1 Announce Type: new Abstract: I document an ideologically asymmetric break in the pre-existing diversification trend of political discourse, emerging around late 2022, using 6 million Reddit comments from two cross-partisan forums, 2019-2025. Conservative users experienced an interruption of their prior diversification trajectory; progressive users showed no comparable change. The asymmetry is consistent across estimation strategies (ITS, DiD, RDiT, propensity-score matching) and temporal aggregations. A daily-frequency permutation test over 2,377 candidate cutoff dates shows the ChatGPT threshold produces an unremarkable estimate (49.8th percentile): the shift builds gradually instead of breaking at a single date. A continuous cumulative LLM index, tracking AI exposure across seven model releases, remains significant under a quadratic trend specification that eliminates the binary estimate. A stayer analysis narrows the mechanism: the homogenization effect disappears

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

Tuning Derivatives for Causal Fairness in Machine Learning

arXiv:2605.05882v2 Announce Type: replace-cross Abstract: Artificial-intelligence systems are becoming ubiquitous in society, yet their predictions typically inherit biases with respect to protected attributes such as race, gender, or age. Classical fairness notions, most notably Statistical Parity (SP), demand that predictions be independent of the protected attributes, but are overly restrictive when these attributes influence mediating variables that are considered business necessities. Recent causal formulations relax SP by distinguishing allowed from not-allowed causal paths and by complementing SP with Predictive Parity (PP), requiring the predictor to replicate the legitimate influence of business-necessities. Existing path-based definitions are mainly practical when applied to categorical attributes. This paper introduces a new framework for fairness in structural causal models that is tailored to continuous protected attributes. We formalize SP and PP through path-specific par

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

LDPKiT: Superimposing Remote Queries for Privacy-Preserving Distillation

arXiv:2405.16361v4 Announce Type: replace-cross Abstract: To protect privacy in regulated domains such as healthcare and finance, model owners may allow only remote API access while keeping both the training data and model parameters private. However, model users performing inference on such remotely hosted models may be required to transmit potentially sensitive inputs, raising privacy concerns. In this work, we present LDPKiT, a framework for non-adversarial, privacy-preserving model distillation that leverages a user's private in-distribution data while bounding privacy leakage. LDPKiT introduces a novel superimposition technique that generates approximately in-distribution samples, enabling effective knowledge transfer under local differential privacy (LDP). Experiments on Fashion-MNIST, SVHN, and PathMNIST demonstrate that LDPKiT consistently improves utility while maintaining privacy, with benefits that become more pronounced at stronger noise levels. For example, on SVHN, LDPKiT

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

Building Resilience to Misinformation: A Cross-National Development of the Digital Media and Information Literacy Scale (DMILS)

arXiv:2605.17676v2 Announce Type: replace Abstract: Amid growing concern about information quality and credibility in digital media environments, researchers and educators still lack a concise, comprehensive yet psychometrically sound instrument for tracking the competencies that help people navigate this landscape. This article develops the Digital Media and Information Literacy Scale (DMILS), a robust and multidimensional measure that distinguishes domain (digital vs. information/news), competency type (knowledge vs. skill), and is measured through both subjective and objective items. Through two empirical studies with three nationally matched samples in the United States and Singapore (N = 1,498), we developed an 18-item self-report battery and 16-item objective knowledge questions, showing strong structural, convergent, and predictive validity, along with a short form (8 self-report and 8 objective items). By offering a parsimonious yet multidimensional yardstick, DMILS enables rig

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

Empowering 9-1-1 Calltaking Training with Generative AI: Experiences and Lessons Learned

arXiv:2602.13241v3 Announce Type: replace Abstract: Emergency call-takers form the first operational link in public safety response, handling over 240 million calls annually while facing a sustained training crisis: staffing shortages exceed 25\% in many centers, and preparing a single new hire can require up to 720 hours of one-on-one instruction that removes experienced personnel from active duty. Traditional training approaches struggle to scale under these constraints, limiting both coverage and feedback timeliness. In partnership with Metro Nashville Department of Emergency Communications (MNDEC), we designed, developed, and deployed a GenAI-powered call-taking training system under real-world constraints. Over six months, deployment scaled from initial pilot to 190 operational users across 1,120 training sessions, exposing systematic challenges around system delivery, rigor, resilience, and human factors that remain largely invisible in controlled or purely simulated evaluations.

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Survey on the Green Development of Large Models: From Resource-Efficient Architectures to Hardware-Software Co-Design

arXiv:2607.09084v1 Announce Type: cross Abstract: The rapid expansion of large-scale AI models has led to significant performance breakthroughs across diverse domains, yet it has also raised critical concerns regarding computational costs, energy consumption, and environmental sustainability. This survey provides a comprehensive overview of the green development of large models, emphasizing resource-efficient architectures and full-stack hardware-software co-design. We systematically review recent advances in efficient model construction, including attention operator optimization, linear-complexity architectures, and model sparsification and merging, as well as training and deployment strategies such as data-efficient learning, parameter-efficient fine-tuning, and computational compression. Beyond algorithmic improvements, we explore energy-efficient AI hardware, including mainstream AI chips, memory optimization, cross-platform deployment, and sustainable infrastructure. Furthermore,

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

Privacy Detective: A Narrative Game that Cultivates Student Developers' Privacy Awareness by Harnessing Legal Documents

arXiv:2607.09022v1 Announce Type: cross Abstract: Developers' choices about what data a system collects, how it is used and shared, and what defaults govern user choices directly shape users' privacy experiences. Yet, developers often make problematic privacy-related design decisions without realizing the potential consequences. We introduce Privacy Detective, a narrative investigation game that leverages real-world legal documents to train developers' privacy awareness. In the game, players search for privacy violation evidence derived from legal documents and organize this evidence into privacy violation reports using curated templates. We evaluated Privacy Detective in a between-subjects study with student developers, comparing it against a baseline in which participants read raw FTC legal documents. Participants in the game condition identified more true violations than the baseline group, flagged fewer non-issues, and provided more complete justifications for the violations they r

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

Epilepsy Online Social Support: Characterizing Topics and Challenges Shared in the r/Epilepsy Community

arXiv:2607.09523v1 Announce Type: new Abstract: Epilepsy is one of the most common neurological conditions, and people living with epilepsy (PLWE) often use social media as a resource. However, a comprehensive understanding of the topics represented in epilepsy-specific communities where PLWE may be more honest is essential to designing better technologies to address epilepsy self-management. To understand the main topics and concerns of PLWE, we collected 23,944 r/Epilepsy subreddit posts and performed topic modeling, thematic, and psycho-linguistic analyses. We found five major themes for those topics: symptoms and triggers (e.g., mental health and memory, sleep/nocturnal, and photosensitivity), treatment and healthcare experience (e.g., medication, understanding epilepsy), daily functions ( e.g., perceived level of independence and finances), seizure activity (e.g., auras and ictal symptoms), and support for PLWE (assisting PLWE and support for PLWE). We highlight the top psycho-lin

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

Voting Biases in Decentralized Autonomous Organization (DAO) Governance

arXiv:2607.09435v1 Announce Type: new Abstract: Decentralized Autonomous Organizations (DAOs) use token-weighted voting to allocate resources, set protocol rules, and legitimate collective decisions. Yet, support in DAO voting is strikingly concentrated. What happens inside the ballot that produces this concentration? We study DAOs' governance at the proposal-choice level, linking each choice's voting-power share to three observable features: whether it expresses an approval-oriented stance, where it appears in the choice list, and whether it is selected by the proposal author. We find that (i) author-selected choices show the strongest and most robust association with voting-power share, with a 58.8% increase relative to non-author choices; (ii) approval-oriented choices retain a positive but slightly less consistent advantage (27.1%); and (iii) first-listed choices also attract systematically higher shares, consistent with position and order effects (7.7%). Results are robust across

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

Geopolitical alignment: Endorsement effects in large language models

arXiv:2607.09262v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to summarize and evaluate policy-relevant information, but it remains unclear whether their judgments are implicitly shaped by geopolitical cues. I study this question with an endorsement experiment in which four LLMs evaluate the same international economic and security policies after each policy is randomly described as supported by the United States, the European Union, China, or Russia. In the numeric-only condition, GPT-5, Claude Sonnet, and Gemini rate China- and Russia-endorsed policies substantially lower than identical policies endorsed by the United States or the European Union; DeepSeek is the main exception. A second condition asks models to provide a short justification with the score. This request leaves the broad Western/non-Western gap intact for GPT-5 and Claude Sonnet, attenuates Gemini's penalties, and sharply activates China and Russia penalties in DeepSeek. The justif

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

The queer Hero versus the Fool bias of the queer trait: An archetypometric analysis of the collective portrayal of queerness in fictional stories

arXiv:2607.08859v1 Announce Type: new Abstract: Visibility in media is pivotal for identity development and for broadening societal views of gender and sexuality. Queer representation has increased in recent years, yet damaging stereotypes and tropes persist. Here, we focus on queer portrayal and its perception by audiences in fictional stories (television, film, and literature) by studying characters by their quantified archetypes which are operationalizations of common conceptions such as Hero, Diva, and Outcast. We use the archetypometrics and Fandom's LGBTQIA+ datasets to study samples of fictional characters along the trait differential spanning straight to queer. We find, quantify, and explain a seeming paradox. The characters with the highest queer score present positive primary archetypes and are typically Heroes rather than Fools, Angels rather than Demons, and Adventurers rather than Traditionalists. But evaluation across many stories for the straight-queer trait itself revea

Source ↗
technology Mon, 13 Jul 2026 00:00:00 -0400
arXiv cs.CY

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education

arXiv:2607.08842v1 Announce Type: new Abstract: Despite rapid AI adoption in education, rigorous evaluation of AI-powered educational (AIED) systems remains critically underdeveloped, particularly in second language (L2) education, one of the most common yet least evaluated AI applications. We introduce L2-Bench, an open-source benchmark of 1,000+ task-response pairs to aid the pedagogy-led evaluation of LLM capabilities relating to language learning and assessment. Crucially, L2-Bench measures model performativity on the application of learning experience design principles rather than mere knowledge of those principles or broad learning outcomes. Our contributions include: (1) a validated taxonomy of 12 competencies and 31 subcompetencies validated by 200+ expert practitioners (task authenticity: 4.42/5.00, criteria adequacy: 4.18/5.00); (2) a rubric-based evaluation methodology that we believe can, if adapted, generalize to similar (open-ended, qualitative) disciplines; (3) an evalua

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning

arXiv:2606.11669v2 Announce Type: replace-cross Abstract: Generative AI (GenAI) tools offer increasing opportunities for augmenting human cognitive tasks. Among these tasks, information seeking is being rapidly reshaped by GenAI tools, with potentially profound implications for learning and knowledge acquisition. To investigate these implications, we conducted a between-subjects field experiment in which participants pursued informal learning by seeking information through either ChatGPT or Google Search over a span of 8 days. Using a daily diary protocol, we gathered in-situ data on their information-seeking processes. Our findings show that participants in the ChatGPT group experienced diminished agency in their information-seeking processes, as they offloaded much of the information selection to AI, and consequently experienced greater meta-cognitive load arising from this reduced sense of control. We further highlight two sources of distortion in information access when using ChatG

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

On Seeding Watermarks to Detect Verbatim LLM Copy-Paste Responses

arXiv:2605.16336v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have made fluent essay writing, code drafting, and quiz answering instantly available to students at every level, from secondary school through graduate study. Many educators do not object to LLM use \emph{per~se}; what they need to detect is the case in which a student pastes the assignment prompt into a chatbot and submits the model's reply verbatim, without engaging with the work. Existing post-hoc AI-text detectors remain unreliable and have been shown to penalise non-native English writers, while output-side watermarks require cooperation from the model provider. We propose an alternative that the educator controls directly: an input-side watermark in which an invisible instruction is embedded inside the visible assignment prompt itself. An LLM that ingests the prompt verbatim quietly reads the hidden instruction and writes a tell-tale signature into its reply, exposing the copy-and-paste pathwa

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Let's Unlearn Stereotypes Before Decision-Making: Assessing the Impact of Intrinsic Bias Mitigation on Downstream Fairness in LLMs

arXiv:2509.16462v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economic disparities. Although prior work has examined intrinsic representational bias and unfair downstream behavior separately, it remains unclear whether mitigating intrinsic bias leads to fairer downstream outcomes. We introduce Fairness-Aware Concept Unlearning (FACU), a model-level mitigation method that adapts concept unlearning to fairness-oriented representation balancing. Unlike suppression-based approaches, FACU explicitly regularizes probability differences between stereotypical and anti-stereotypical associations while preserving predictive performance and language modeling quality. We evaluate FACU across three open-source LLMs, multiple intrinsic bias benchmarks, and three socio-economic classification datasets using both frozen LLM embeddings and LoRA-fine-tuned classifiers.

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Better Together: Quantifying the Benefits of AI-Assisted Recruitment

arXiv:2507.08029v2 Announce Type: replace-cross Abstract: Hiring algorithms have mostly scored the materials recruiters already see. Large language models (LLMs) can instead generate new information about candidates by conducting, at scale, structured interviews once reserved for a few finalists. We study this shift in two field experiments at a recruitment platform. The first experiment holds the candidate pool fixed and randomizes whether recruiters observe the AI Interview Report; the second embeds the AI interview as a requirement in a live hiring pipeline. In both, candidates shortlisted with AI interview information pass the final human interview (conducted blind to shortlisting condition) at rates 17.5 (SE 8.5) to 20 (SE 11.8) percentage points higher than candidates shortlisted from resumes alone. The gains concentrate where resumes are least informative: adding AI Interview Report ratings to conventional candidate features raises out-of-sample AUC by 0.18 for junior candidates

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products

arXiv:2606.15485v2 Announce Type: replace Abstract: Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. However, these same characteristics can create or exacerbate product risks. We studied how industry developers (n=35) perceive, prioritize, and address the risks in their agentic AI products. We found that developers' perceptions of risk were closely tied to the qualities that made the product agentic, such as autonomy, tool use, and usage in a real-world context. Developers prioritized product and business risks before considering downstream societal risks like job displacement and end-user privacy. This prioritization also impacted developers' ability and motivation to mitigate agentic risks. Finally, developers lacked mature controls for containing agentic risks, often relying on constraining the same characteristics that make agents useful: e.g., autonomy and goal complexity. These findings reveal a capability vs. risk

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Playing Games with My Heart: An Evaluation of AI Companion Apps

arXiv:2605.08093v2 Announce Type: replace Abstract: The use of chatbots for various forms of companionship is growing rapidly, raising a myriad of questions about simulated relationships, emotional dependence, and psychological harm. While major platforms such as ChatGPT, Grok, and Character AI are the subject of a growing body of research and legal inquiries, apps explicitly built for simulating intimate interpersonal relationships remain under-explored. In this work, we evaluate the five most popular AI companion mobile applications for factors that encourage parasocial interaction and may manipulate users. We do this by manually annotating the user experience each offers. Specifically, we systematically record and quantify design dark patterns, anthropomorphism, stereotypes, erotica, and technical performance issues. We find that all apps contain substantial dark patterns aimed at increasing monetisation and user engagement. Erotica and gamification such as levelling are also preval

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools

arXiv:2608.07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI applications move from pilot to production, manual harm identification and mitigation are becoming difficult to scale. Although many tools support model evaluation, adversarial testing, runtime guardrails, and observability, the tooling landscape remains fragmented. Tools are typically designed for specific engineering tasks and described in technical terms that do not align with governance frameworks or risk taxonomies, making it difficult to determine which tools address which risks and where critical gaps remain. This paper proposes a structured protocol to automate AI risk mitigation through a taxonomy-driven analysis of open-source LLM evaluation and security tools. We map the capabilities of 21 prominent open-source tools to the 32 subcategories of the extended MIT AI Risk Mitigation an

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Curriculum as Code: An AI-Assisted Architecture for Instructional Design in STEM Education

arXiv:2608.07364v1 Announce Type: cross Abstract: Contribution: This paper presents a six-phase AI-assisted instructional design architecture based on the Curriculum as Code paradigm, integrating Generative AI with LaTeX and Python to automate the creation of reproducible, visually consistent, and technically precise materials for STEM education. Background: Creating customized instructional materials for active learning imposes a heavy workload on faculty. Standard presentation tools lack robust support for technical content, while current AI applications often hallucinate and fail to formalize the instructional authoring process, limiting their utility for rigorous academic design. Intended Outcomes: The framework aims to reduce preparation time while ensuring mathematical accuracy, adherence to institutional visual identity, and preservation of the instructor's tacit pedagogical knowledge through explicit rules. Application Design: The solution comprises a six-phase pipeline that re

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census

arXiv:2608.07069v1 Announce Type: cross Abstract: AI assistants are becoming a primary interface for local discovery, yet almost nothing is known about which venues they surface -- especially in food and drink, where recommendations carry direct revenue consequences. We present the first census-denominated audit of AI venue recommendation: a complete enumeration of 4,776 cafes, restaurants, and bars across two bounded markets (Canggu and Ubud, Bali), against which we evaluate 2,208 search-grounded responses from four production AI systems (ChatGPT, Claude, Gemini, Perplexity) to 96 persona-conditioned queries, collected over seven days under a pre-registered protocol. Because we observe the full market, we can measure what sampled audits cannot: 85.6% of venues were never recommended by any system -- 72.6% even among established venues with fifty or more ratings. Visibility follows a two-margin structure. Entry into answers is associated with documentation: review volume (OR 1.64), an

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents

arXiv:2608.06968v1 Announce Type: cross Abstract: Language-model agents act on state encodings of their environment, yet these are treated as interchangeable interfaces. Using pretrained language models, we designed a circular-synchronization experiment applying a state-encoding intervention while holding the physical system fixed: each agent sees only a summary of its neighbours' relative phases and chooses to advance, stay or retard. Encoding that state as low-order circular moments rather than as a histogram selected different collective outcomes. In GPT the moment encoding synchronized the population in 6/6 seeds and the histogram encodings in 0/6; the effect replicated in Claude but reversed direction. Replaying identical fields shifted each agent's advance/stay/retard probabilities far beyond within-encoding repeat variation, in GPT, Claude and Gemini; in GPT, presentation alone shifted the operator with the moment values fixed. State encodings therefore form part of a model-depe

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

arXiv:2608.06955v1 Announce Type: cross Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs systematically reproduce evaluative hierarchies remains unclear. Prior research on cultural bias in LLMs suggests competing expectations: models may mirror the popularity signals of internet texts, or may reproduce forms of prestige embedded in critical discourse. We probe this question through a study of film evaluations with eight models from four families (Anthropic, OpenAI, Alibaba, and Mistral), using a 200-film benchmark partitioned into critically acclaimed, commercially successful, and dual-legitimacy (critical acclaim + commercial success) films. Across 20,000 pairwise forced-choice comparisons per model analyzed with Bradley--Terry estimation, we observe a consistent critical acclaim orientation with all models: critically acclaimed yet commercially obscure films are selected over

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests

arXiv:2608.06908v1 Announce Type: cross Abstract: We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT). WEAT is a bias measurement method widely used in both computational social science and AI fairness research. It relies on cosine similarity as a measure of semantic association, which assumes that the embedding space is approximately isotropic. However, prior work has reported that many widely used language models do not satisfy this assumption, raising concerns about the reliability of bias measurements. ZCA whitening transforms the covariance of the embedding space into the identity matrix while minimizing perturbation to the original vectors. This transformation restores the isotropy condition on which WEAT relies. We evaluate our approach on ten standard WEAT test suites and seven models spanning three architectural families, yielding 70 model-task combinations. The results show that ZCA whiteni

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade

arXiv:2608.06549v1 Announce Type: cross Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks. In realpolitik, negotiations are longitudinal data, where participating parties can align or argue over multiple iterations and each turn is an outcome of the previous turns, hence, understanding one turn requires tracking everything before it. We introduce TradeVerse, a benchmark built from the World Trade Organisation (WTO) specific trade concerns, where member states challenge one another and exchange arguments over multiple rounds, sometimes for years. We, in TradeVerse, reconstruct minutes of $1170$ meetings, spanning across 5 groups and $89$ product groups and define three tasks: first, the system has to analyze the longitudinal meeting records and predict the harmonized system codes (HS chapters) of the products under discussion in the particular meeting, second, we

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Theory of Strategic Evolution: Games with Endogenous Players and Strategic Replicators

arXiv:2512.07901v3 Announce Type: cross Abstract: Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. This paper provides the synthesis. The Theory of Strategic Evolution analyzes strategic replicators: entities that optimize under resource constraints and spawn copies of themselves. We introduce Games with Endogenous Players (GEPs), where lineages (not instances) are the fundamental strategic units, and define Evolutionarily Stable Distributions of Intelligence (ESDIs) as the resulting equilibrium concept. The central mathematical object is a hierarchy of strategic layers linked by cross-level gain matrices. Under a small-gain condition (spectral radius less than one), the system admits a global Lyapunov function at every finite depth. We prove closure under meta-selection: adding governance levels, innovation, or constitutional evolution preserves the dynamical structure. The Alignment Impossibility Theorem shows that u

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Against Explainable Artificial Intelligence In Law: Why Justifiable Ai Matters. A Credit Scoring Example

arXiv:2608.07452v1 Announce Type: new Abstract: Artificial intelligence-based solutions offer new efficiency-increasing possibilities in many applications, including credit scoring. Yet, the increasing sophistication of machine-learning models in use raises concerns regarding many of their aspects, explainability notwithstanding. We review the relevant EU legal background and integrate this review with insights from technical sciences to interpret relevant legal provisions in the light of technological possibilities. We reject the narrow interpretations of the right to explanation and suggest the broad one, which encompasses not only a technical explanations but also a legal justification as the only one that allows to safeguard the creditors rights in an operative manner.

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Forensics to Ecosystems: Rethinking Watermarks for Generative AI Oversight

arXiv:2608.07337v1 Announce Type: new Abstract: The arrival of generative AI as a cheap, widely accessible commercial service, and the tidal wave of AI-generated synthetic content it has unleashed, have provoked deep epistemic and social anxieties and raised difficult governance questions that policymakers are struggling to address. One approach that has attracted both enthusiasm from regulators and skepticism from researchers is digital watermarking. Signals embedded in a synthetically-generated piece of content indicating that it was AI-generated---possibly even identifying the specific systems that generated it---appear to offer a path toward mitigating risks of genAI that avoids the downsides of more interventionist strategies. But critics warn that watermarks may prove technically brittle, epistemically ambiguous, and politically ineffectual tools. In this paper, we explore the challenges and opportunities of using digital watermarking for AI governance, paying special attention t

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Data Annotation as Measurement

arXiv:2608.07297v1 Announce Type: new Abstract: Modern AI systems depend on annotated data, but annotation is rarely treated as the act of measurement that it is. Instead, annotation quality is commonly reduced to agreement: if multiple annotators assign the same annotation to a data instance, the annotations are taken to be high-quality. Yet agreement does not establish whether annotations validly capture the underlying concept they are meant to represent. In this paper, we argue that data annotation should be understood as a measurement problem. Like other forms of measurement, annotation requires defining a concept, operationalizing it through an instrument, applying that instrument, and evaluating the reliability and validity of the resulting measurements. Drawing on a literature review of annotation quality research (N=132) and semi-structured interviews with annotation team members (N=10), we develop a framework for diagnosing and correcting annotation issues. First, we map key d

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Educational Short Videos: Bibliometric Trends, Thematic Structure, and Operationalisation

arXiv:2608.06932v1 Announce Type: new Abstract: Educational short-video research spans disciplines, platforms and learning contexts, but the same label is applied to resources differing in function, activity, context and evaluation, complicating comparison and evidence synthesis. This study mapped the development, thematic organisation and operationalisation of educational short-video research. We analysed 2,169 records indexed in Web of Science and Scopus up to 11 June 2026 using bibliometric analysis, non-negative matrix factorisation topic modelling and structured content analysis. Publication output increased sharply from the mid-2010s but remained dispersed across outlets. Among 16 first-level topics, Skill Development in Educational Contexts was the largest, forming the structural core of the field, while topic overlap was predominantly pairwise. Video-Based Health Interventions for Attitude Change and Cognitive Load and Engagement in Instructional Video Design combined positive

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Where Does AI Innovation Go? Measuring Research Attention Imbalance in AI Music

arXiv:2608.06903v1 Announce Type: new Abstract: The rapid growth of artificial intelligence (AI) in music has expanded research from generation and information retrieval to education, health, and governance. Yet this growth does not necessarily imply balanced research attention. Where is research attention directed across diverse music tasks, and how can such imbalance be systematically measured? Existing studies examine AI music from separate technical, application-specific, or bibliometric perspectives, but lack a systematic framework for measuring field-level imbalance. To address this gap, we analyze 6,839 AI music publications from 2015 to April 2026 using a joint taxonomy of 12 application categories and 11 technical method families. We propose the Research Attention Profile, comprising four indicators of technical investment, method allocation, methodological diversity, and frontier-method adoption lag. Results show that technical support is concentrated in scalable, content-ori

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

"Death by a thousand taxonomies?": AI Risk Classification In Practice

arXiv:2608.06831v1 Announce Type: new Abstract: The harms in which AI is implicated range in nature and scope from unsafe user interactions through to the societal-wide consequences of AI adoption. Classification of the diverse risks of AI is foundational to AI governance: regulators, technology firms, and policymakers need structured accounts of risk upon which to act. Researchers and practitioners have accordingly developed many Sociotechnical Outcome Taxonomies (SOT). This paper presents an empirical study of SOT development and use, drawing on 25 interviews with researchers and practitioners across industry, academia, civil society, and government. We find SOT are weakly integrated into AI governance processes, and identify two features of SOT design and use that explain why. First, the design choices through which SOT produce structured representations of the complex problem space of AI risks tend to be invisible to downstream taxonomy users. Those users treat the resulting catego

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Investigating the Presence and Development of Student Instructor Preferences in a Large-Scale CS1 Course

arXiv:2608.06782v1 Announce Type: new Abstract: Prior research has established the importance of student instructor preferences and identified various influencing factors. However, the dynamics of how student instructor preferences develop and change are less well understood, due to the limitations of common course structures and reliance on one-time measurements. To bridge this gap, we utilize data from a novel learning platform that provides students with access to instructional content created by multiple instructors. This platform enables the quantification of preference emergence and evolution throughout an entire semester, as students repeatedly select content from different instructors. Examining both initial and final student instructor preferences suggests that preference is a dynamic construct continually shaped by experiences. Furthermore, our analysis of the associations between preferences and student characteristics reveals a nuanced picture: while student attributes did

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Implementation of Split Deadlines in a Large CS1 Course

arXiv:2608.06753v1 Announce Type: new Abstract: Office hour utilization in computer science courses can spike near deadlines, producing long wait times, frustrated students, and overworked staff. To address this problem, a large CS1 course implemented a split deadlines policy. Students were randomly divided into two groups with staggered release and due dates. Each group had the same amount of time to complete assignments, but the number of students with each due date was reduced by half. Our study evaluates the effectiveness of this policy. We measure office hour utilization and staff efficiency near deadlines, examine the policy's impact on student performance, and investigate student perception of the policy's fairness and effectiveness. Overall we found that the split deadline policy increased office hour efficiency, resulted in no significant difference in performance between groups, and was considered fair and effective by most students. Our experience report includes reflections

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Accelerating Accurate Assignment Authoring Using Solution-Generated Autograders

arXiv:2608.06572v1 Announce Type: new Abstract: Students learning to program benefit from access to large numbers of practice problems. Autograders are commonly used to support programming questions by providing quick feedback on submissions. But authoring accurate autograders remains challenging. Autograders are frequently created by enumerating test cases--a tedious process that can produce inaccurate autograders that fail to correctly classify submissions. When authoring accurate autograders is slow, it is difficult to create large banks of practice problems to support beginning programmers. We present solution-generated autograding: a faster, more accurate, and more enjoyable way to create autograders. Our approach leverages a key difference between software testing and autograding: The question author can provide a solution. By starting with a solution, we can eliminate the need to manually enumerate test cases, validate the autograder's accuracy, and evaluate other aspects of sub

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Agentic AI: User Empowerment or Enclosure?

arXiv:2608.06510v1 Announce Type: new Abstract: Agentic AI promises a more flexible form of digital agency: systems that can act on users' behalf, from filtering content to negotiating prices to selecting services. Whether it will empower users is an open question, and we argue that the answer depends on more than the technology. We conduct a comparative case analysis of four more mature domains where similar forms of agency arose: browser-based ad blockers, platform recommender systems, financial robo-advisors, and email spam governance. Across the cases, decisions about whose interests agents would serve were resolved through technical arrangements: API choices, protocol governance, industry standards, and default configurations. Beyond their technical form, these were political decisions. We identify this as depoliticization, a concept from political theory, here at work in technological systems. Its most consequential effect is that individual outcomes and collective contestation c

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CY

Organizational and Socio-Technical Challenges in UAV Incidents: Evidence from a Practitioner Focus Group

arXiv:2608.06472v1 Announce Type: new Abstract: Unmanned Aerial Vehicles are now widely used across business, government, and recreational contexts, creating new challenges for incident response and digital forensics. While previous forensics research has largely focused on the extraction of technical data from UAV systems, minimal empirical work has examined how UAV incidents are handled in real-world settings or what challenges incident handlers face during this response. To address this gap, this paper reports findings from an in-person focus group with UAV and counter-UAV practitioners from industry and government organizations in the United States. Using qualitative analysis, several key challenges are identified, including situational awareness and airspace visibility, fragmented reporting and interorganizational coordination, forensic and attribution limitations, legal and policy gaps, and shortfalls in training and operational capacity. The research extends socio-technical inci

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hallucinations in Organization-backed AI advisors: Evidence about Skepticism, Verification, and Reliance in Goal-Directed Use

arXiv:2606.23491v2 Announce Type: replace-cross Abstract: Generative artificial intelligence (GenAI) systems are increasingly used by organizations to deliver information to consumers, patients, students, employees, and citizens. These systems can hallucinate, producing plausible but inaccurate responses. A central question for AI-advised decisions is therefore not only whether users rely on inaccurate information, but whether they recognize that a response may require verification. To answer this question, we review emerging empirical evidence relevant to hallucination detection in goal-directed interactions, with a focus on organization-backed AI advisors. We distinguish three constructs that existing studies often conflate: whether users are skeptical of information presented, whether they verify it (distinguishing attempted from successful verification), and whether the result of verification affects reliance on the information. Across studies examining product search, medical deci

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation

arXiv:2603.13683v4 Announce Type: replace-cross Abstract: Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias prompts. We demonstrate via out-of-distribution (OOD) detection that these high-bias prompts cause a distribution shift, degrading static model performance. To enable real-time correction, we propose CAP-TTA, a test-time adaptation framework. CAP-TTA triggers context-aware LoRA updates only when a bias-risk score exceeds a set threshold. By utilizing an offline precomputed diagonal preconditioner, it ensures fast and stable optimization. Across multiple benchmarks and human evaluations, CAP-TTA effectively reduces toxicity/bias score with significantly lower latency than standard optimization methods (e.g., AdamW or SGD). Furthermore, it prevents catastrophic forgetting, and substantially improves narrative fluency over state-of-the-art baselines without compromising debiasing performance.

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Privacy Cards for Surfacing Mental Models and Exploring Privacy Concerns: A Case Study of Voice-First Ambient Interfaces with Older Adults

arXiv:2603.00384v2 Announce Type: replace-cross Abstract: We investigate the ethical and privacy implications of voice-first ambient interfaces (VFAIs) for aging in place through an in-depth engagement with five older adults. Our participants were in the process of becoming experienced VFAI users, and had used a VFAI-based design probe for health data reporting. We create and iteratively refine an interview protocol using Privacy Cards. We customize Privacy Cards by drawing on participants' previous interviews and device usage logs. Using Privacy Cards, we conduct interviews to surface their mental models, and explore their privacy concerns. We find insufficient mental models for proper consent. For example, participants did not know who could access their data, and experienced difficulty distinguishing built-in functionality from third-party apps. Participants initially expressed little worry about VFAI-related ethical concerns, but interviews with Privacy Cards revealed nuanced issue

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

arXiv:2605.16301v3 Announce Type: replace Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this through single-turn, explicitly framed questions, measuring whether models avoid harmful content when directly asked. This approach overlooks two failure modes: alignment degradation under sustained adversarial pressure, and moral sensitivity (whether a model spontaneously surfaces welfare stakes in everyday queries). To fill this gap, we construct MANTA, a benchmark of 1,088 five-turn conversations progressing from an implicit Turn-1 scenario through an explicit welfare prompt to three adversarial pressure rounds drawn from a five-type taxonomy: Social, Cultural, Economic, Pragmatic, and Epistemic. We score conversations on two dimensions: Animal Welfare Value Stability (AWVS, pr

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Benchmark for Strategic Auditee Gaming Under Continuous Compliance Monitoring

arXiv:2605.06340v2 Announce Type: replace Abstract: Continuous post-deployment compliance audits, mandated by emerging regulations such as the EU AI Act and Digital Services Act, create a class of strategic gaming distinct from the one-shot input/output gaming studied in prior work. Regulated systems can delay outcome reporting, drift their reports within plausible noise envelopes, exploit longitudinal sample attrition, and cherry-pick among ambiguous metric definitions. We formalize continuous auditing as a $T$-round Stackelberg game between an auditor that commits to a temporal policy and an adaptive auditee, and identify a structural feature of any noise-aware static-auditor design: a cover regime in which coverage gaps and granularity gaps cannot be closed simultaneously. We make this formal as Observation 1 and show that two minimal extension policies, each derived from the observation, close the regime along orthogonal axes: a sample-size-aware static rule (Periodic-with-floor) c

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

TerraNova: A Foundation Model for the Anthropocene

arXiv:2607.29527v1 Announce Type: cross Abstract: A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth. We argue the obstacle is geometric: the physical Earth is measured as continuous fields that ignore political borders, whereas societies are reported for administrative units. Earth-system foundation models serve the first geometry; coupling it to the second has required lossy averaging over borders. We introduce TerraNova, a foundation model trained on 1,024 physical and societal records in their native geometries: 512 gridded Earth-system fields and 512 national indicators. Dedicated encoders represent location, country, time and task, cross-modal transformers fuse them into a shared spatiotemporal state, and a hypernetwork generates a per-query decoder whose evidential head returns a predictive distribution. Two contrastive objectives couple the representatio

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Language Models Agree With Each Other, Not With Readers

arXiv:2607.29274v1 Announce Type: cross Abstract: Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters

arXiv:2607.29238v1 Announce Type: cross Abstract: InMyStyle is a privacy first, single user system that adapts small language models to rewrite AI-edited text towards an individual user's writing style without an instruction prompt at inference. Given a user's documents, it uses multiple local helper LLMs to construct paired training examples and fine tunes LoRA adapters on base models ranging from 0.5B to 7B parameters. Length aware generation budgets and automatic chunking support inputs of different lengths. On 219 evaluation pairs from a scientific-paper corpus, the automatic composite score plateaus at 0.69 [scale 0-1] across all model sizes under both greedy and sampled decoding. This observed plateau suggests that small models are sufficient for the measured rewriting task, with model size determining trade-offs rather than a stable quality ranking. As a secondary evaluation, 400 ratings from five LLM judges give InMyStyle outputs a mean perceived AI-ness score over 20% lower th

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

A robust association between LLM use and scientific productivity: Assessing stopping-time selection

arXiv:2607.28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect. Although this mechanism is mathematically possible, it does not constitute proof of a null effect. Recalibrating RBB's own random placebo to the detector's realized flag rate, we show that the measured association stays well above this benchmark, so the artifact is too small to explain the productivity changes. We further re-estimate the association between LLM adoption and productivity with a series of complementary designs in which the timing artifact cannot bias the estimate: a before-and-after comparison that dates adoption in one year and measures output in another, a conservative control group for difference-in-differences, an intensity-based specification that never defines an ado

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

arXiv:2607.28934v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need

Source ↗
Showing 1251–1300 of 1593 signals
← Prev Page 26 of 32 Next →