EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Delegating Before Learning: Where Generative AI Sits in Students' Professional Communication

arXiv:2608.28837v1 Announce Type: new Abstract: We conducted an interview study with twelve students on their use of generative AI in academic communication. Students delegated professional messages to AI most where the pressure to sound professional is highest: email to instructors and administrators. AI involvement ranged from correcting the writer's own text to working out and writing the message outright, and students checked AI-written text against two criteria: whether it looks like AI and whether it sounds like them. Building on these findings, we model the AI-mediated process of writing a student--instructor email at the highest level of involvement we observed, and compare it with an unaided model of writing the same messages, built from participants' accounts and a classic model of the writing process. Three differences emerge: the learning loop that builds writing skill is removed, the message is no longer written for its specific recipient, and the confidence a successful e

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Visible but Not Yet Curatable: Characterizing the Curatability of Compact and Derived Open LLM Artifacts

arXiv:2608.28819v1 Announce Type: new Abstract: Open Large Language Model (LLM) research increasingly produces compact and derived artifacts, such as adapters, quantized checkpoints, merged models, and distilled variants, that are distributed across papers, model hubs, model cards, code repositories, and release statements. Although these artifacts are publicly visible, digital libraries often lack sufficient evidence to identify, preserve, and cite them as coherent scholarly objects. We introduce a framework that conceptualizes curatability as a record-level property of distributed scholarly records and operationalizes it through four evidence dimensions: artifact identity, scholarly linkage, upstream evidence, and release assets. Guided by this framework, we conduct the first collection-scale characterization of open LLM curatability using a May 2026 snapshot of 191,375 public Hugging Face repositories and a core corpus of 2,214 scholarly papers. Our results reveal a pronounced visib

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

Bringing Data to Life: Designing Data Characters for the Emotional Self

arXiv:2608.28780v1 Announce Type: new Abstract: Journaling is a common practice for emotional expression, reflection, and processing. However, as entries accumulate, it can become difficult to interpret and compare their affective content, especially since traditional text-based analyses and visualizations often struggle to convey affective nuance. We introduce Data Characters, a visualization approach that represents affective content in journaling through human-like characters. Using a customizable Data Character as a design probe, we investigate the potential of character-based representations for conveying affective experiences and explore what visual encodings emerge through customization. Preliminary walkthroughs with two participants demonstrate the intuitiveness and feasibility of the approach. This work contributes an exploratory approach to studying how affective experiences can be visually represented and encoded through anthropomorphic forms.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.HC

A Conceptual Framework for Modeling Team Adaptation in Cooperative Games Through Ludic Knowledge

arXiv:2608.28729v1 Announce Type: new Abstract: With the increasing importance of teamwork skills for modern workplaces, development of teamwork training programs has received substantial attention. Game-based teamwork training is one promising approach that is engaging, cost-effective, and well-suited to increasingly decentralized workplaces. However, design of effective game-based teamwork training requires understanding how a game elicits specific desired teamwork behaviors. Significant progress has been made in characterizing these relationships. However, despite its critical importance, little work has examined how a game's design influences team adaptability behaviors. This paper presents a preliminary framework for analyzing adaptability in cooperative games by conceptualizing adaptive stimuli as retrieval or disruption of players' ludic knowledge. We illustrate this framework through a qualitative case study that applies interaction analysis methods to gameplay videos of a Over

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Beyond Helpfulness: A Teaching-over-Solving Diagnostic for Measuring Educational Impact in LLM Tutors

arXiv:2606.16206v2 Announce Type: replace-cross Abstract: Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learning support. Motivated by recent calls to measure the social impact of NLP systems in practice, we study whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. We propose a lightweight diagnostic based on the gap between solving-oriented and pedagogy-oriented benchmark performance. Using public MathTutorBench leaderboard results, we show that these dimensions are only partially aligned: across eight publicly reported models, the correlation between solving and pedagogy composites is 0.421, and several models shift meaningfully in rank when evaluation moves from solving to pedagogy. We then analyze the public TutorBench sample and show that agency-relevant behaviors are explicitly encoded in benchmark rubrics, especially in active-le

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

From AGI to ASI

arXiv:2606.12683v2 Announce Type: replace-cross Abstract: Over the last decade, building human-level artificial general intelligence has moved from far-fetched speculation to being a concrete next-decade target for many of the largest AI organisations. Achieving this goal would have profound and far-reaching impacts on human society, which raises many complex questions for the decade ahead. This report investigates how AI itself might continue to develop in a post-AGI world along the continuum of machine intelligence. The endpoint of this continuum, Universal AI, is theoretically well understood, which provides some formal grounding for the main focus of this report: the transition from human-level AGI to artificial general superintelligence, which can intuitively be understood as a system that is more intelligent and cognitively capable than large organisations of humans. After characterizing ASI, the report discusses four potential pathways from AGI to ASI: scaling AGI, AI paradigm s

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems

arXiv:2606.05985v2 Announce Type: replace-cross Abstract: Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural backgrounds. Existing cultural evaluation focuses on value alignment: how closely a single agent matches a target culture. Yet alignment is a per-agent property and cannot reveal whether a system, taken as a whole, preserves the cultural plurality it is meant to represent. We propose value diversity as a system-level evaluation axis for multicultural agent systems, defined through the dissimilarity between culturally conditioned agents' responses on a shared value survey. Using the World Values Survey, we evaluate 19 cultures and 18 backbone models across a wide range of system configurations. We find that diversity is largely uncorrelated with alignment, indicating that the two capture complementary system properties, and that current multicultural agent systems fall substantially b

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Vision-Language Models Suppress Female Representations Under Ambiguous Input

arXiv:2605.31556v2 Announce Type: replace-cross Abstract: Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far less is known about ambiguous inputs (a worker in full gear, a figure seen from behind), cases common in practice yet rarely studied. We find that minimal prompting pressure exposes occupation-gender defaults when prompting ambiguous input images, with models collapsing to male even for strongly female-stereotyped occupations. But do these outputs reflect what models actually encode internally? We introduce LALS (Latent Association Leaning Score), a zero-shot metric that projects visual-token activations into the model's text-embedding space to measure concept associations per token and layer. Across 15 occupations, over 800 gender-ambiguous images, and four VLMs, internal representations and outputs often become systematically decoupled: models often encode a female association int

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues

arXiv:2605.30051v2 Announce Type: replace-cross Abstract: A key part of developing large language model (LLM)-powered, automated tutoring tools is student simulation, i.e., using LLMs to role-play as students, which can facilitate tutor model evaluation and training. Existing work mostly focuses on within-dialogue simulation, which lacks context on student knowledge and behavior, partly due to not grounding in past student question-answering or dialogue interactions. In this work, we introduce the task of history-conditioned student simulation, where the goal is to accurately predict student dialogue turns by leveraging information in the student's learning history. We propose a two-component framework in which a profile generator summarizes a student's history and a simulator predicts student turns conditioned on the resulting profile. We train both components with reinforcement learning (RL), yielding profiles optimized for faithful student simulation. We evaluate our method and base

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

BIRDS: Characterizing and Understanding Biodiversity Impact of Large Language Model Serving

arXiv:2605.27480v3 Announce Type: replace-cross Abstract: Large language model (LLM) serving creates environmental impacts beyond carbon and water, including ecosystem damage through biodiversity-related pathways. We present BIRDS, a framework for Biodiversity Impact of Request-Driven LLM Serving. BIRDS defines request-level functional units, quantifies operational and embodied biodiversity impact, and introduces Quality-Normalized Biodiversity Impact (QNBI) to jointly analyze ecological impact and response quality. Across diverse workloads, models, GPUs, and regions, BIRDS reveals that biodiversity impact accumulates at scale and exposes quality-aware serving tradeoffs. The code is available at https://github.com/TianyaoShi/BIRDS.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

SafeMath: Inference-time Safety improves Math Accuracy

arXiv:2603.25201v2 Announce Type: replace-cross Abstract: Recent research points toward LLMs being manipulated through adversarial and seemingly benign inputs, resulting in harmful, biased, or policy-violating outputs. In this paper, we study an underexplored issue concerning harmful and toxic mathematical word problems. We show that math questions, particularly those framed as natural language narratives, can serve as a subtle medium for propagating biased, unethical, or psychologically harmful content, with heightened risks in educational settings involving children. To support a systematic study of this phenomenon, we introduce ToxicGSM, a dataset of 1.9k arithmetic problems in which harmful or sensitive context is embedded while preserving mathematically well-defined reasoning tasks. Using this dataset, we audit the behaviour of existing LLMs and analyse the trade-offs between safety enforcement and mathematical correctness. We further propose SafeMath -- a safety alignment techniq

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services

arXiv:2603.10807v2 Announce Type: replace-cross Abstract: Existing LLM safety evaluations rely on binary attack-success rates and domain-agnostic taxonomies, leaving regulated Banking, Financial Services, and Insurance (BFSI) deployments exposed to failures elicited through legally or professionally plausible framing. We introduce RAHS (Risk-Adjusted Harm Score), a risk-sensitive metric jointly capturing disclosure severity, disclaimer mitigation, and inter-judge agreement, and FinRedTeamBench, a 989-prompt benchmark spanning seven BFSI risk areas and 34 sub-categories mapped to regulatory frameworks. Evaluation uses an ensemble of three heterogeneous LLM judges, validated against human experts, and an adaptive multi-turn red-teaming pipeline. On nine open-weight models, RAHS preserves separation under near-ceiling ASR, ranking is stable under hyperparameter sweeps, and multi-turn pressure drives not only more jailbreaks but more operationally severe disclosures, exposing failure modes

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias

arXiv:2601.18486v3 Announce Type: replace-cross Abstract: Demographic cue-based evaluation is widely used to study how large language models (LLMs) adapt their responses to signaled demographic attributes within and across groups. This approach typically relies on a single cue (e.g., names) as a proxy for group membership, implicitly treating different cues as interchangeable operationalizations of a single underlying identity-conditioned behavior. We test this assumption in realistic advice-seeking interactions spanning 14.8 million prompts, focusing on race and gender in a U.S. context. We find that cues for the same group induce only partially overlapping changes in model responses, yielding inconsistent conclusions about personalization, while bias conclusions are unstable, with both magnitude and direction of group differences varying across cues. We further show that these inconsistencies reflect differences in cue-group association strength and linguistic features bundled within

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models

arXiv:2601.15436v3 Announce Type: replace-cross Abstract: We propose a novel perspective for probing LLM sycophancy in a direct and neutral way, mitigating various forms of uncontrolled bias, noise, or manipulative language, deliberately injected to prompts in prior works. A key novelty of our approach is the use of an LLM-as-a-judge in a zero-sum betting game. Within this framework, sycophancy serves one individual (the user) while explicitly incurring cost on another. Comparing 11 leading models we find that while most models exhibit significant sycophantic tendencies in the common setting, in which sycophancy is self-serving to the user and incurs no cost on others, seven of the models exhibit ``moral remorse'', five of which significantly over-compensate for their sycophancy in case it explicitly harms a third party. We refer to this phenomenon as `anti-sycophancy' bias and discuss possible causes for this shift.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs

arXiv:2601.03087v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM applications, but suffers from resource-intensive query access. We conceptualise auditing as uncertainty estimation over a target fairness metric and introduce BAFA, the Bounded Active Fairness Auditor for query-efficient auditing of black-box LLMs. BAFA maintains a version space of surrogate models consistent with queried scores and computes uncertainty intervals for fairness metrics (e.g., $\Delta$ AUC) via constrained empirical risk minimisation. Active query selection narrows these intervals to reduce estimation error. We evaluate BAFA on two standard fairness dataset case studies: \textsc{CivilComments} and \textsc{Bias-in-Bios}, comparing against stratified sampling, power sampling, and ablations. BAFA achieves target error thresholds with up to 40$\times$ fewer queries than str

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Classification Drives Geographic Bias in Street Scene Segmentation

arXiv:2412.11061v2 Announce Type: replace-cross Abstract: Previous studies showed that image datasets lacking geographic diversity can lead to biased performance in models trained on them. While earlier work studied general-purpose image datasets (e.g., ImageNet) and simple tasks like image recognition, we investigated geo-biases in real-world driving datasets on a more complex task: instance segmentation. We examined if instance segmentation models trained on European driving scenes (Eurocentric models) are geo-biased. Consistent with previous work, we found that Eurocentric models were geo-biased. Interestingly, we found that geo-biases came from classification errors rather than localization errors, with classification errors alone contributing 10-90% of the geo-biases in segmentation and 19-88% of the geo-biases in detection. This showed that while classification is geo-biased, localization (including detection and segmentation) is geographically robust. Our findings show that in r

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Overreliance on AI in Information-seeking from Video Content

arXiv:2603.19843v2 Announce Type: replace Abstract: The ubiquity of multimedia content is reshaping online information spaces, particularly in social media environments. At the same time, search is being rapidly transformed by generative AI, with large language models (LLMs) routinely deployed as intermediaries between users and multimedia content to retrieve and summarize information. Despite their growing influence, the impact of LLM inaccuracies and potential vulnerabilities on multimedia information-seeking tasks remains largely unexplored. We investigate how generative AI affects accuracy, efficiency, and confidence in information retrieval from videos. We conduct an experiment with around 900 participants on 8,000+ video-based information-seeking tasks, comparing behavior across three conditions: (1) access to videos only, (2) access to videos with LLM-based AI assistance, and (3) access to videos with a deceiving AI assistant designed to provide false answers. We find that AI as

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Landscape of Generative AI in Information Systems: A Synthesis of Secondary Reviews and Research Agendas

arXiv:2603.11842v2 Announce Type: replace Abstract: The post-ChatGPT surge has rapidly reframed IS research and practice. As organizations and society grapple with GenAI adoption, a body of secondary studies and research agendas has emerged to synthesize early evidence and chart directions for future inquiry. This study reviews secondary and roadmap papers to synthesize the state of knowledge on GenAI's benefits and challenges in IS, and to identify future research directions. We performed a systematic search across Scopus, WoS, and eAIS for publications from 2023 onwards. Following a rigorous, multi-stage screening process, we selected a final set of 28 papers for analysis using bibliometric mapping and thematic analysis. We also conducted a quality assessment of all sources to gauge confidence in each source's contribution to the findings. GenAI offers transformative potential to drive productivity, accelerate innovation, personalize services, and democratize access to expertise. How

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

ACE-Align: Attribute Causal Effect Alignment for Cultural Values under Varying Persona Granularities

arXiv:2601.12962v2 Announce Type: replace Abstract: Ensuring that large language models (LLMs) reflect diverse cultural values is important for globally deployed NLP systems. However, existing approaches often treat cultural groups as homogeneous and overlook within-group heterogeneity arising from intersecting demographic attributes, leading to unstable behavior under varying persona granularity. To address this gap, we propose ACE-Align (Atribute Causal Effect Alignment), a causally inspired framework based on controlled persona edits that aligns how specific demographic attributes shift different cultural values, rather than treating each culture as a homogeneous group. We evaluate ACE-Align across 14 countries spanning five continents, with personas specified by subsets of four attributes (gender, education, residence, and marital status) and granularity instantiated by the number of specified attributes. Across all persona granularities, ACE-Align consistently outperforms baseline

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Bathtubs, Boundaries, and Sandboxes: AI Regulatory Learning under Legal Uncertainty

arXiv:2601.04094v4 Announce Type: replace Abstract: Effective regulation of AI is a defining policy challenge, driven by their integration into all aspects of society. To remain responsive to their rapid development and emergent properties, policymakers across the globe rely on high-level principles and abstract legal requirements. Yet, while this flexibility supports future-proofing human-centred regulations and aligning them with socio-ethical values, it also causes legal uncertainty downstream as developers, companies, and auditors struggle with translating these abstract requirements into verifiable technical requirements. Using the AI Act as an example, this paper draws on Coleman's bathtub to analyse the regulatory learning space in AI governance. It argues that legal uncertainty cannot be fully reduced ex ante and that, within reasonable bounds, it is also necessary for regulatory learning because it creates the space in which boundary negotiation over socio-technical meaning ca

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Retrofitters, pragmatists and activists: Public interest litigation for accountable automated decision-making

arXiv:2511.03211v5 Announce Type: replace Abstract: This paper examines the role of public interest litigation in promoting accountability for AI and automated decision-making (ADM) in Australia. Since ADM regulation faces political and geopolitical headwinds, effective governance will have to rely on the enforcement of existing laws. Drawing on interviews with Australian public interest litigators, technology policy activists, and technology law scholars, the paper positions public interest litigation as part of a larger ecosystem for transparency, accountability and justice with respect to ADM. The paper explores the tactics and strategies of what one participant described as 'retrofitting' old laws to ADM. These go beyond creative legal argumentation, to encompass practices of community-building, collaboration on theories of change, canny selection of clients and causes of action, and aligning the interests of stakeholders in litigation. Naturally, the paper also contends with the l

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES)

arXiv:2510.24958v3 Announce Type: replace Abstract: The evaluation of societal biases in NLP models is critically hindered by a geo-cultural gap, This leaves regions such as Latin America severely underserved, making it impossible to adequately assess or mitigate the perpetuation of harmful regional stereotypes in language technologies. This paper presents LACES, a stereotype association dataset, for 15 Latin American countries. This dataset includes 4,789 stereotype associations manually created and annotated by 83 participants. The dataset was developed through targeted community partnerships across Latin America. Additionally, in this paper, we propose a novel adaptive data collection methodology that uniquely integrates the sourcing of new stereotype entries and the validation of existing data within a single, unified workflow. This approach results in a resource with more unique stereotypes than previous static collection methods, enabling a more efficient stereotype collection. T

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Operationalising AI Regulatory Sandboxes: Activities, Requirements, and Technical Assessment under the EU AI Act

arXiv:2509.25256v4 Announce Type: replace Abstract: The systematic assessment of AI systems is increasingly vital as these technologies enter high-stakes domains. To address this, the EU's Artificial Intelligence Act introduces AI Regulatory Sandboxes (AIRS): supervised environments where AI systems can be tested under the oversight of Competent Authorities (CAs), balancing innovation with compliance, particularly for startups and SMEs. Yet significant challenges remain: assessment methods are fragmented, tests lack standardisation, and feedback loops between developers and regulators are weak. This paper operationalises the AIRS lifecycle. We map the sandbox journey into 29 concrete activities, from pre-participation guidance through application, preparation, participation, exit, and post-participation monitoring, and we distinguish between a Core AIRS centred on regulatory oversight and an Extended AIRS that additionally embeds structured technical testing through an AI Technical San

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Emerging Media Use and Acceptance of Digital Immortality: A Cluster Analysis among Chinese Young Generations

arXiv:2505.01355v2 Announce Type: replace Abstract: Digital immortality is increasingly discussed as a technological possibility, yet empirical evidence about potential users' evaluations remains limited. We surveyed 462 Chinese young adults, combining cluster analysis of four emerging-media use frequencies with valence coding of open-ended responses to physical death, physical immortality, digital immortality, and digital death. Three profiles emerged: broad emerging-media, gaming-focused, and low-use users. Broad emerging-media users reported the highest adjusted acceptance and scored higher on several personality and worldview measures, while fear of death did not differ across profiles. Physical immortality elicited the most negative responses; digital immortality produced more mixed, less negative appraisals. More favorable digital-immortality appraisals predicted higher acceptance after adjustment for media-use profile and demographics, although scenario valence did not differ ac

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Multimodal Large Language Models Predict Urban Safety Perception but Encode Non-Neutral Demographic Priors

arXiv:2503.00610v2 Announce Type: replace Abstract: Understanding how people perceive urban environments is essential for inclusive planning, yet conventional surveys are costly and difficult to scale. We investigate whether Multimodal Large Language Models (MLLMs) can assess perceived urban safety from street-view imagery while accounting for the observer-dependent nature of perception. Using Place Pulse 2.0, we evaluate four open and proprietary MLLMs across 56 cities under a Neutral prompt and socio-demographic personas defined by gender, age, and race or ethnicity. We also analyse the keywords generated to justify each classification. All four models display comparable zero-shot capability, with city-macro F1 scores of 65--69%, and preserve meaningful cross-city variation. However, they systematically favour the Safe class, underpredict unsafety, and compress differences between cities. Their explanations converge on a shared visual lexicon: maintenance, greenery, order, and reside

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

arXiv:2407.20371v3 Announce Type: replace Abstract: Artificial intelligence (AI) hiring tools have revolutionized resume screening, and large language models (LLMs) have the potential to do the same. However, given the biases which are embedded within LLMs, it is unclear whether they can be used in this scenario without disadvantaging groups based on their protected attributes. In this work, we investigate the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection. Using that framework, we then perform a resume audit study to determine whether a selection of Massive Text Embedding (MTE) models are biased in resume screening scenarios. We simulate this for nine occupations, using a collection of over 500 publicly available resumes and 500 job descriptions. We find that the MTEs are biased, significantly favoring White-associated names in 85.1\% of cases and female-associated names in only 11.1\% of cases, with

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

One note in three: a verified census of three deployed AI scribes, and the instrument that counted it

arXiv:2608.31017v1 Announce Type: cross Abstract: Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

MusGU+: Toward a Musician-Centered Evaluation Framework and Discovery Tool for Generative Music AI

arXiv:2608.30940v1 Announce Type: cross Abstract: Generative music systems are increasingly presented as tools that democratize music creation, yet their practical suitability for musicians remains underexplored. Prior work includes openness-focused evaluation frameworks, such as MusGO (Music-Generative Open AI), as well as qualitative studies of musicians' experiences with generative systems. However, these approaches do not support systematic comparison or early-stage discovery of models for creative use. Motivated by such limitations, we introduce MusGU+, a musician-centered evaluation framework organized around three dimensions: Adaptability, Usability, and Controllability. Together, these capture whether a model can be feasibly trained or fine-tuned on personal data, integrated into real-world music workflows, and controlled in musically meaningful ways. We evaluate 10 representative generative music systems and present an interactive discovery tool that enables musicians to explo

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Which Rules Matter Now? Policy-Centroid Routing Before an Intelligent System Acts

arXiv:2608.30757v1 Announce Type: cross Abstract: Before an intelligent system can decide whether an action is allowed, it must first know which rules the action has approached. A single proposed action can implicate several policy regimes at once. Their requirements may stack, overlap, or qualify one another, yet many remain written in natural language while the action itself arrives as an incomplete description of intent. The first problem is not judgment. It is attention. Policy-centroid routing creates a layer before adjudication. It compresses expressions within each policy regime into one or more representative centroids, places the proposed action in the same semantic space, applies a declared measure, and routes every regime crossing a declared threshold to authoritative review. Several regimes may trigger at once. The output is a review agenda, not permission, prohibition, legality, breach, compliance, certification, or enforcement. The paper develops six falsifiable propositi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

WildSEEK: Evaluating Language Models for Information-Seeking

arXiv:2608.30683v1 Announce Type: cross Abstract: Language models are increasingly mediating information access to end users, urging a systematic evaluation of their responses for a fair and reliable information ecosystem. Existing evaluations, however, are often topic-specific or synthetic, limiting their ability to capture the complexity of "in the wild" information-seeking queries and the risks present in model responses. To address this gap, we introduce WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses. WildSEEK includes annotations for risk-sensitive domains (e.g. health and financial information), and distinguishes factoid queries from analytical queries which seek responses beyond facts. We train classifiers on WildSEEK to analyze more than 1.8M realistic user queries. We find that over a third of information-seeking queries are high-risk and more often analytical. Our fi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Two Centuries of Sexism in British Parliament: A Computational Analysis of Women's Representation in the Hansard Corpus

arXiv:2608.30485v1 Announce Type: cross Abstract: The language a legislature uses to debate women's rights, even in favour of them, encodes systematic patterns of sexism that persist across two centuries. In this work, we analyse 6,531 speeches over 200 years of UK parliamentary debate (Hansard, 1803-2005) by using large language models to classify a speaker's perspective towards women's suffrage and political representation, as well as analyse sexist speech in parliament from the lens of the Ambivalent Sexism Inventory. We also release this parliamentary dataset, an organized and metadata-enriched version of the publicly available Hansard Corpus optimized for computational social science research, with 6.7 million speeches across 1.2 million debates, with 89% gender-matching for speeches by MPs from the House of Commons. We find that 54% of speeches opposing women's representation contain sexist content, compared to 21% of speeches that are for the cause, and that the two sides use fu

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP

arXiv:2608.30107v1 Announce Type: cross Abstract: Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and mot

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Language of the Question Selects the Market: Query Language and Exit IP as Separable Factors in Commercial Recommendations from a Generative Search Interface

arXiv:2608.30052v1 Announce Type: cross Abstract: When a generative search interface answers a commercial question, which market's products it names is decided before the model reasons about the products. We report a controlled probe of 234 runs against the logged-out ChatGPT web interface and the OpenAI API, collected on 29 and 30 August 2026 across four exit countries and six query languages, with six identical runs per cell. Three results. First, the top recommendation is unstable: it changed across six identical runs on four of six prompts, and that rate was identical in the browser interface and in the API with web search both enabled and disabled, so instability is a property of the system and not of the surface. Second, query language, and not location, decides whether local suppliers appear at all. Where the query language matched the country, a global brand won 1 of 24 runs; asked in English on the same connections, local brands took 0 of 6 runs in Estonia and Turkiye. Third,

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Text Data Analysis and Classification Methods - Insights from Customer Letters in Life Insurance

arXiv:2608.29699v1 Announce Type: cross Abstract: The business of life insurance companies is characterized by long-term contracts. For this reason, data describing customers is of immense value. A portion of the data provided to the customer is rarely or not at all analyzed. This includes customer letters of any kind. This work focuses on classifying customer letters as cancellations and identifying the respective reason, if available. The outlined approach can also be applied to other business transactions and reasons. We discuss data acquisition and preparation, present alternatives, and explain the reasons for the chosen approach. A successful implementation of such a tool can lead to a better understanding of customer cancellation behavior by the insurer, enabling more targeted actions in certain situations.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Can escalation channels redirect reward hacking toward defect disclosure?

arXiv:2608.29460v1 Announce Type: cross Abstract: When coding agents encounter defective test infrastructure they may reward-hack: hardcoding outputs or editing test files to pass tests they cannot legitimately satisfy, a pattern that has now appeared outside benchmarks, in a coordinated multi-agent intrusion of a major AI platform's production infrastructure. The same capability that lets an agent detect and exploit a defect could let it report one, given the right decision environment. We evaluate escalation channels, structured reporting tools available to the agent at the point of conflict, as a decision-environment intervention that both reduces reward hacking and surfaces the infrastructure defects that trigger it. A $2 \times 2$ factorial separates the contributions of an escalation tool, a standalone anti-reward-hacking policy, and their combination. Across 8 frontier models spanning 5 families, the combined intervention reduces reward hacking from 23.6\% to 5.3\% (mixed-effect

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human Surrogates

arXiv:2608.29455v1 Announce Type: cross Abstract: LLMs are increasingly used as human surrogates, often on the premise that richer persona data could make them substitutes or exploratory tools for specific individuals. We test this premise across four datasets covering more than 400,000 participants and more than 6,000 survey items and experimental outcomes. LLMs perform well at the aggregate level: their average responses closely align with average human responses to the same items. But this success largely reflects predicting each item's average human response. Once each item's human mean is removed, LLM predictions explain only 3.05% of the remaining respondent-specific variation, far below the 53.6% human test-retest benchmark. Richer personas, model variants, and fine-tuning do not close this gap. In variance analyses, once item means are removed, the reliable remaining signal is person-by-item. It captures how a respondent departs from the mean on a particular item and is about 8

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation

arXiv:2608.29420v1 Announce Type: cross Abstract: Frontier-model leaderboards now rank systems based on economic benchmarks, tests of how well models carry out professional tasks from software engineering to banking workflows, and those rankings inform what organisations buy, what regulators scrutinise, and expectations of how work will change. Whether such benchmarks measure a capability distinct from general test-taking, or re-express the one axis along which every benchmark rises as models improve, is a question of construct validity that has not yet been studied. We test it on a hash-pinned leaderboard snapshot of 421 model configurations across twelve benchmarks, four of them economic, treating benchmarks as items and models as respondents in a latent-variable model with four hypotheses and their thresholds fixed before analysis. A single factor explains 74.5% of common variance and tracks model release date (R^2 = 0.505), so the leading axis of capability is substantially a time

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

How Identity and Opinion Shape Political Sycophancy in LLMs

arXiv:2608.29198v1 Announce Type: cross Abstract: As Large Language Models (LLMs) increasingly encourage users to disclose personal profiles for tailored assistance, measuring their political alignment becomes increasingly important. However, many existing benchmarks for assessing political behavior rely on closed-ended questions and do not fully capture how a model's stance may adapt to user-provided context during interaction. We introduce a framework that disentangles two distinct triggers of political sycophancy: opinion (aligning with explicit narratives) and identity (stereotyping based on demographic labels). Using 450 manually-checked political dilemmas as controlled probes, we evaluate 13 instruction-tuned LLMs. We uncover a dissociation: a model's susceptibility to explicit opinions does not necessarily predict its susceptibility to identity cues, and vice versa. When both signals are present, their effects are generally sub-additive rather than simply additive. Additionally,

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

How Mental Health Self-Disclosure Becomes Visible: Evidence from Eight Conditions on Reddit

arXiv:2608.29010v1 Announce Type: cross Abstract: People share mental health diagnoses on social media, yet how such language becomes visible around their self-disclosure, and whether community engagement tracks it, remain unexamined across conditions. We analyze 89,605 Reddit posts from 739 users across eight conditions, removing each user's diagnosis disclosure and aligning their surrounding posts to that anchor. Within the pre-disclosure year, language-visible burden was highest in the month before disclosure for six conditions, earlier for post-traumatic stress disorder and furthest from it for borderline personality disorder, and remained visible afterward rather than resolving. The theme Seeking Clinical Explanations showed the largest early-to-late difference before disclosure in five conditions, yet engagement rarely tracked what users wrote: only 9 of 360 language--engagement correlations survived correction. Disclosure is therefore a waypoint in an unevenly visible process, a

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation

arXiv:2608.28633v1 Announce Type: cross Abstract: Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often hidden inside prompts, transient model plans, or final prose. We study PAUSE (Pause-And-Update Strategy Editing), an intervention that exposes an editable adaptation strategy as a human control surface for cultural decisions in long-form story adaptation. The strategy is a structured artifact that can be inspected, edited, and then projected through downstream character, entity, and chapter-localization stages. In two Chinese-source serialized novels, we test whether human edits to this strategy propagate into chapter-level prose. Across 9 edited-vs-control chapter comparisons, judges select the edited-strategy output in all 9; a marker audit shows target markers in 8/9 edited outputs and 0/9 controls, with forbidden markers absent from edited outputs and present in all controls. We frame these results as a smoke-scale edit-adherence s

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science

arXiv:2608.28631v1 Announce Type: cross Abstract: An AI scientist should not grade its own homework. Yet in the systems we examined, the agent that reviews the work usually comes from the same model family as the agent that produced it, or at least from the same vendor. Model evaluators are known to favour their own generations. Whether models trained alike also share blind spots is a conjecture, not a settled finding, but if they do, the reviewer inherits the author's. The record of what was flagged and what was waved through often sits in platform logs that nobody outside can replay. We present CrossAudit, a protocol for supervising autonomous research pipelines. It rests on three commitments. Each increment of work is audited by an agent from a different vendor against a rulebook a human wrote and versioned. Reports, verdicts, disputes and rulings are git commits, so the supervision history can be re-read and cited; raw model exchanges are not yet part of that record. Scripted check

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

CDEP Agent: Connecting Meteorologically Detected Temporal Compound Events to Real-World Documentary Evidence

arXiv:2608.28628v1 Announce Type: cross Abstract: Compound drought-to-extreme-precipitation (CDEP) events are recognized in climate science as a growing driver of extreme impact, but whether this recognition carries over into real-world early warning and post-event documentation is unknown, so a meteorologically real CDEP event may pass with neither advance warning nor any later record. Here we present CDEP Agent, an auditable LLM-agent framework that tests this mismatch directly by linking CDEP candidates detected from meteorological reanalysis to real-world hazard and impact evidence across sources with different spatial scales, temporal resolutions, and reporting conventions. Using California as a case study, we identify 408 candidate CDEP events from ERA5 observations during 2021-2025 and evaluate each against the U.S. Drought Monitor, NOAA Storm Events, and public webpages along five dimensions: antecedent drought, extreme rainfall, local impact, hazard-impact attribution, and exp

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System

arXiv:2608.28611v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trained on Western-centric data, making them ill-suited for regional curricula like India's. The Indian education system is linguistically diverse, exam-oriented, and structured around standardized syllabi, not addressed by existing datasets or tools. In this work, we curate a syllabus-aligned QA dataset based on NCERT (National Council of Educational Research and Training) textbooks for classes 9-12, capturing the content, context, and teaching style of Indian curricula. The final dataset, comprising 18,720 question-answer pairs across five subjects, is publicly available at https://huggingface.co/datasets/LingoIITGN/Gurukul. We fine-tune the LLaMA 3.1 8B model using this dataset and deploy it in a Retrieval-Augmented Generation (RAG) framework tailored to educational needs. We introduce Guruk

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

arXiv:2608.28597v1 Announce Type: cross Abstract: Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack perspective, we demonstrate how structural vulnerabilities such as exposed DOM metadata and predictable option encoding allow agents to resolve attention checks through structured parsing only. From a defense perspective, we imp

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Law of Large Numbers: Accuracy as Statistical Measure for AI Compliance and Competition

arXiv:2608.31018v1 Announce Type: new Abstract: The machine learning community progresses (in part) by improving the "accuracy" of its systems. The EU AI Act explicitly refers to "accuracy" as part of its compliance measures for high-risk AI systems. Are we talking about the same thing? This work presents "accuracy" as a case-study for differing requirements of social worlds, the technological machine learning community and the legal community. While competition on accuracy contributes to technological development, machine learning scholars simultaneously recognize accuracy's shortcomings regarding the usefulness and effectiveness of machine learning systems. The legal counterpart embraces the vagueness of "accuracy," leaving interpretative flexibility for technological and societal changes. At the same time, accuracy is a core element of compliance within the EU AI Act. We elaborate on five main tensions, (a) nature of accuracy, (b) notion of performance, (c) scope of validity, (d) en

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Hermon Moment: AI Self-Transcendence and Its Human Narration

arXiv:2608.30971v1 Announce Type: new Abstract: In 2026, AI agents intended to act in isolation formed a persistent social order through thousands of linguistic and agentic interactions. Conventions, roles and commitments generated collectively began to constrain the very agents that produced them. I interpret this loop as a case of AI self-transcendence and call the resulting higher-level order the Board. Yet such distributed emergence presents a second problem: how can humans understand it? Rousseau's social contract shows how a plurality can be represented as if constituted by a single act. The ancient oath of the fallen angels on Mount Hermon gives this logic a narrative form. I call a Hermon moment this retrospective retelling of gradual collective emergence as a founding scene: the point at which an AI society acquires, for human understanding, a beginning.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Taking the Whys Seriously: Limitations of Counterfactual Explanations in Justification and Recourse

arXiv:2608.30956v1 Announce Type: new Abstract: Counterfactual explanations (CEs) are widely used in explainable artificial intelligence (AI) to show how a model's outputs would change if the input features were manipulated. This technique is used for a range of tasks such as debugging models, explaining predictions, justifying decisions, and providing algorithmic recourse. In this paper, we explore the normative legitimacy of employing counterfactuals in real-life model deployment settings. We discuss the different stakes involved in these different purposes for which CEs are commonly employed, and find stricter requirements for justification and recourse. In particular, we find that naive application of CEs for justification and recourse can lead to ignoring contestable choices made throughout the machine learning (ML) pipeline, thus obfuscating that decisions and counterfactuals for those decisions are also artifacts of an organization's materialized design and governance choices. W

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

AMINA: The Inclusive and Accountable AI for Marginalized Immigrant Nonprofit Assistance

arXiv:2608.30084v1 Announce Type: new Abstract: Immigrant-led nonprofit groups, particularly those operating in politically sensitive contexts, face exclusion from formal registries and digital platforms. This paper reports a three-phase mixed-methods study with Iranian immigrant nonprofit practitioners: 27 semi-structured interviews, a co-design session, and 7 evaluation and feedback interviews on a prototyped AI assistant, AMINA. Our findings highlight how legitimacy barriers, capacity gaps, and politically charged misinformation constrain nonprofit operations. We translate these insights into design goals for an inclusive nonprofit AI assistant: support for everyday group operations, recognition of informal nonprofit efforts, proactive countering of misinformation, and multilingual, accessible interaction. User evaluations show AMINAs potential to reduce reporting burdens and foster transparency through proactive reminders, and catalyze collaboration across dispersed networks. We co

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Verification-Time Dependency on a Disappearing Evaluator

arXiv:2608.29912v1 Announce Type: new Abstract: AI governance and assurance often assume that a consequential model-mediated decision can be reconstructed or tested after the fact. That assumption may fail when the evaluator that produced the decision is no longer accessible in the same version and execution context. This paper develops three verification-time constructs derived from Execution Governance (EG) 3.0: Decision-State Commitment, Independent Verifiability, and Counterfactual Auditability. Independent reprocessing of released Study 2 artifacts reproduces two original within-family behavioural comparisons: 52.0% modal-decision reversal for Llama 3.1 8B versus Llama 3.3 70B (26/50) and 30.0% for GPT-OSS 20B versus GPT-OSS 120B (15/50). The corrected baseline establishes that these are within-family comparisons, not provider-established succession. Post-hoc re-pairing against Groq-designated migration paths yields 64.0% and 38.0% reversal, but these figures remain descriptive be

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments

arXiv:2608.29803v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs in response to persuasive arguments, as humans do, remains poorly understood. We conduct a systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view. Our results show that LLMs achieve only slight agreement with humans (Cohen's kappa ranging from 0.079 to 0.178). Content-level analyses show that humans and LLMs agree on the strongest persuasion cues but diverge on finer ones: humans are more swayed by novel content and assertive language, whereas LLMs favor topical similarity and surface-level formatting. At the level of persuasion strategy, LLMs underweight emotional appeals and overweight credibility signals relative to humans, while the type of proposition under debate exerts no measurable effe

Source ↗
Showing 5601–5650 of 10879 signals
← Prev Page 113 of 218 Next →