EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Mimicry to True Intelligence (TI) -- A New Paradigm for Artificial General Intelligence

arXiv:2509.14474v3 Announce Type: replace-cross Abstract: The debate around Artificial General Intelligence (AGI) remains open due to two fundamentally different goals: replicating human-level performance versus replicating human-like cognitive processes. We argue that performance-based definitions are inadequate, offering no roadmap for research and failing to define the qualitative nature of genuine intelligence. Four decades of work on cognitive architectures have already characterized the relevant mechanisms; our claim is that this knowledge is not being consulted by the research program now driving AGI development. What we offer is a translation of it into terms bearing on contemporary systems, with criteria for assessing them. We define True Intelligence (TI) as a system characterized by six components: five architectural pillars for which assessment criteria can be stated (embodied sensory fusion, core directives, dynamic schemata, a highly interconnected multi-expert architectu

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

TS-Mob: Social and Geographical-Aware Time Series Foundation-Model Framework for Human Mobility Prediction

arXiv:2507.00945v2 Announce Type: replace-cross Abstract: Short-term forecasting of aggregated human mobility flows supports urban planning, intelligent transportation systems, and emergency response, yet existing models often require substantial mobility history and learn spatial structure implicitly through grids or graphs. Time series foundation models provide strong temporal priors but typically lack explicit geographic and social conditioning for origin-destination interactions. We introduce TS-Mob, a framework that conditions a fine-tuned time series foundation model (TimesFM) forecaster on a gravity-inspired destination-attractiveness index that encodes geographic and social signals computed from open data (living population, centroid distances, and Overture POI counts), together with weather covariates. Evaluated on commonly used benchmarks like Bike New York City, Taxi Beijing, and a nation-scale Spain origin-destination matrix estimated through mobile phone data, TS-Mob outpe

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence

arXiv:2501.17629v2 Announce Type: replace-cross Abstract: Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instructions precisely. Passing the test holds significance as evidence that a machine demonstrates human-like intelligence, and as a marker for artificial-general intelligence in commercial and legal domains. We conducted Turing's three-player imitation game with an LLM by following the guidelines identified by Turing and applying scientific standards wherever detailed instructions were missing. We performed a computer-imitates-human game without duration constraints and a man-imitates-woman game as a benchmark. In the computer-imitates-human game, only one participant misidentified the large language model, indicating that claims of large language models' passing the Turing test are premature. Participants required over five minutes for both tasks, with the man-imitates-woman game taking longer;

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory

arXiv:2406.14373v3 Announce Type: replace-cross Abstract: The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science research at scale. Building upon prior explorations of LLM agent design, our work introduces a simulated agent society where complex social relationships dynamically form and evolve over time. Agents are imbued with psychological drives and placed in a sandbox survival environment. We conduct an evaluation of the agent society through the lens of Thomas Hobbes's seminal Social Contract Theory (SCT). We analyze whether, as the theory postulates, agents seek to escape a brutish "state of nature" by surrendering rights to an absolute sovereign in exchange for order and security. Our experiments unveil an alignment: Initially, agents engage in unrestrained conflict, mirroring Hobbes's depiction of the state of nature. However, as the simulation progresses, social contracts emerge, leadi

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Private Again: Artificial Intelligence Agents Restore Anonymity---Foreclosing Discrimination and Its Proof

arXiv:2607.23539v2 Announce Type: replace Abstract: Artificial intelligence agents can transact online on behalf of a human principal---browsing, paying, receiving, and reviewing---without revealing who that principal is. That architecture starves algorithmic discrimination of its inputs---identity, purchase history, location history, behavioral traces, and demographic proxies---but also forecloses its proof. Disparate-treatment needs comparators; disparate-impact needs protected-class baselines; and *Iqbal*-era pleading needs specific factual allegations---doctrinal predicates that anonymous transactions never generate. The effects fall asymmetrically: those most vulnerable to discrimination are least able to afford the shield and, when harms remain, least able to prove them. The challenge for the law shifts from detecting and remedying algorithmic discrimination to governing agent-mediated anonymity as civil rights infrastructure: ensuring access to privacy-preserving agents, regulat

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Contemporary AI lacks the imagination to diverge or negate in science

arXiv:2606.08251v3 Announce Type: replace Abstract: Bold claims that AI will accelerate scientific discovery have raced ahead of evidence from working scientists, yet large-scale, scientist-in-the-loop evidence is scarce. Here we mount the largest evaluation to date, inviting authors of 121,640 recent preprints in biology, medicine, chemistry, and social science to judge large language model (LLM)-generated ideas derived from their own papers. 6,749 representative scientists returned 25,139 rating sets on novelty, feasibility, probability of being true, and favorability of adoption. Three patterns emerge. First, non-reasoning LLMs collapse into a narrow "hivemind" of similar ideas while reasoning models explore a wider hypothesis space, but no model spontaneously proposes null hypotheses, a move humans make more freely. Second, scientists reward ideas resembling their own and prize probability over novelty, though social scientists tolerate risk more than life scientists; senior social

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap

arXiv:2605.16283v3 Announce Type: replace Abstract: Large-scale AI deployment data and controlled learning experiments characterize different consequences of the same technology. Deployment telemetry shows that AI use is concentrated in skilled work and frequently supports immediate task performance. It observes tasks, interaction patterns, and outputs, however, not whether users become more capable of performing those tasks independently. Controlled studies measure independent capability more directly, but only in narrower populations and settings, with outcomes that vary substantially by interaction design. We formulate this discrepancy as a stock--formation measurement gap: current systems observe the use of existing expertise more readily than the formation of future expertise. Because formation has historically been society's recovery mechanism through technological change, the gap matters well beyond any single classroom. We synthesize the experimental and observational evidence

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

UGAF-ITS: A Standards Harmonization Framework and Validation Tool for Multi-Framework AI Governance in Distributed Intelligent Transportation Systems

arXiv:2604.22789v2 Announce Type: replace Abstract: Organizations deploying AI-enabled Intelligent Transportation Systems face fragmented governance: ISO/IEC~42001 demands a certifiable management system, the EU AI Act imposes binding high-risk obligations from August~2026, and the NIST AI Risk Management Framework structures voluntary practice. This paper introduces UGAF-ITS, a standards harmonization framework that consolidates 154 source obligations from the three instruments into 12 unified controls across eight governance domains through a reproducible five-phase crosswalk methodology. A three-tier operating model allocates each control to the vehicle, edge, or cloud tier where enforcement and defensible evidence production are feasible. An evidence backbone of 20 versioned artifacts supports a single audit package across all three frameworks without duplicating content. We evaluate UGAF-ITS through an open-source governance engine applied to four architecturally distinct ITS depl

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Locating Translation as a Craft in the Age of AI Translation

arXiv:2604.00758v2 Announce Type: replace Abstract: Rapid development of Large Language Models (LLMs) and similar automated approaches for translation tasks is increasingly affecting the landscape of translation technologies. As concerns about the outsourcing of translator work to these automated translation tools grow, it is increasingly crucial to gather insights from the translation community directly. To this end, we conduct an interview study with 19 professional translators working across 11 languages and 11 domains to understand their perspectives, experiences, and concerns with using translation technologies in their work. We find that translators are cautious when incorporating new tools into their workflow, with several expressing concerns that machine translation (MT) and LLMs are infringing on the necessary human aspects and verification processes of translation. Importantly, translators are worried that these tools have potential for harmful downstream effects due to compr

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Internal Deployment in the AI Act

arXiv:2512.05742v4 Announce Type: replace Abstract: This memorandum analyzes and stress-tests arguments in favor and against the inclusion of internal deployment within the scope of the European Union Artificial Intelligence Act (AI Act). In doing so, it aims to offer several possible interpretative pathways to the European Commission, AI providers and deployers, courts, and the legal and policy community at large based on Articles 2(1), 2(6), 2(8) of the AI Act. Specifically, this memorandum first analyzes interpretative pathways based on Article 2(1)(a)-(c) supporting the application of the AI Act to internally deployed AI models and systems. Then, it examines possible objections and exceptions based on Articles 2(6) and 2(8), with particular attention to the complexity of the scientific R&D exception under Article 2(6). Finally, it illustrates how Articles 2(1), 2(6), and 2(8) can be viewed as complementary to each other, once broken down to their most plausible meaning and interpre

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Simulating Dispute Mediation with LLM-Based Agents for Legal Research

arXiv:2509.06586v2 Announce Type: replace Abstract: Legal dispute mediation plays a crucial role in resolving civil disputes, yet its empirical study is limited by privacy constraints and complex multivariate interactions. To address this limitation, we present AgentMediation, the first LLM-based agent framework for simulating dispute mediation. It simulates realistic mediation processes grounded in real-world disputes and enables controlled experimentation on key variables such as disputant strategies, dispute causes, and mediator expertise. Our empirical analysis reveals patterns consistent with sociological theories, including Group Polarization and Surface-level Consensus. As a comprehensive and extensible platform, AgentMediation paves the way for deeper integration of social science and AI in legal research.

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings

arXiv:2508.17092v2 Announce Type: replace Abstract: Knowledge Tracing (KT) aims to predict a student's future performance based on their sequence of interactions with learning content. Many KT models rely on knowledge concepts (KCs), which represent the skills required for each item. However, some of these models are vulnerable to label leakage, a phenomenon in which the input data inadvertently reveal the correct answer, particularly in datasets with multiple KCs per question. We propose a straightforward yet effective solution to prevent label leakage by masking ground-truth labels during input embedding construction whenever such leakage could occur. To accomplish this, we introduce a dedicated \texttt{MASK} label, inspired by masked language modeling (e.g., BERT), to replace ground-truth labels. In addition, we introduce Recency Encoding, which encodes the step-wise distance between the current item and its most recent previous occurrence. This distance is important for modeling le

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks

arXiv:2507.03162v2 Announce Type: replace Abstract: The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These models exhibit remarkable capabilities in code-related tasks and problem-solving, raising questions about their potential and limitations in advanced CS contexts. This study presents a novel bilingual (English-Romanian) multimodal (text and image) dataset of multiple-choice questions derived from a high-level computer science competition. A particularity of our dataset is that the problems are conceived such that some of them are easier solved using reasoning on paper, while for others writing code is more efficient. We systematically evaluate State of The Art LLMs on this dataset, analyzing their performance on theoretical programming tasks. Our findings reveal the strengths and limitations of current LLMs, including the influence of language choice (English vs. Romanian), providing insights into

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Ethical Framework for Responsible Foundational Models in Medical Imaging

arXiv:2406.11868v2 Announce Type: replace Abstract: The emergence of foundational models represents a paradigm shift in medical imaging, offering extraordinary capabilities in disease detection, diagnosis, and treatment planning. These large-scale artificial intelligence systems, trained on extensive multimodal and multi-center datasets, demonstrate remarkable versatility across diverse medical applications. However, their integration into clinical practice presents complex ethical challenges that extend beyond technical performance metrics. This study examines the critical ethical considerations at the intersection of healthcare and artificial intelligence. Patient data privacy remains a fundamental concern, particularly given these models' requirement for extensive training data and their potential to inadvertently memorize sensitive information. Algorithmic bias poses a significant challenge in healthcare, as historical disparities in medical data collection may perpetuate or exacer

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football

arXiv:2608.09887v1 Announce Type: cross Abstract: Ball possession is the most-cited and most-misleading number in football: 60% recycled in one's own half is not 60% spent pinning the opponent back. Existing event-based possession-value frameworks (expected threat, VAEP, on-ball value) price on-ball actions but ignore the off-ball question a sterile possession poses: did holding the ball create space, or was the circulation dead? We answer this in two layers. First, an event-side junk-possession index prices each possession sequence by its peak threat gain under an expected-threat grid and -- after reconstructing the live scoreline to exclude lead-protecting circulation -- flags low-threat sequences in tied-or-losing states. On the 2026 FIFA World Cup (103 matches, 206 team-matches) the flag correlates negatively with points (r=-0.37) and xG difference (r=-0.51, partly index-coupled). It is not a repackaging of on-ball value: with team offensive VAEP and field tilt held fixed, the junk

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans

arXiv:2608.09717v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction from theory-grounded persona profiles constructed from ten psychological and relational constructs and organized into three tiers: socially attractive, socially mixed, and socially unattractive. We examine LLM ratings in two studies and compare them with human judgments in a third study. In Study 1, 34 LLMs rated 12 profiles across three repeated runs. Although some models tended to give higher or lower ratings overall, they showed strong stability across runs, consistent three-tier ordering, and high agreement in relative profile ordering. Study 2 examined sensitivity to gender presentation using six matched name-and-pronoun profile pairs and a separate pronoun-only test with a gender-neutral name, finding no

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Beyond headcount and human capital: The Effective Cognitive Population as a decomposable capacity unit for AI-era planning

arXiv:2608.09642v1 Announce Type: cross Abstract: National planning counts population, human capital, and artificial-intelligence preparedness in separate ledgers. Demographic accounting has advanced from headcount to skills-adjusted stocks and still debates how much age structure retains once skills are modeled, yet no existing unit carries the conditions under which preparedness becomes productive capacity. This study introduces the Effective Cognitive Population (ECP), a decomposable unit that weights population by capability and by the conditions under which capability is deployed, anchored to the World Bank Human Capital Index Plus (HCI+) and the non-overlapping dimensions of the IMF AI Preparedness Index. The architecture is portable in principle; the case tested here is artificial intelligence, which has a published preparedness index. For 144 countries, HCI+ becomes a productivity level, AI opportunity uses digital infrastructure and innovation integration, conversion governanc

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics

arXiv:2608.09638v1 Announce Type: cross Abstract: Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight. We present Avalon-ToM-Bench, a fine-grained benchmark that operationalizes ToM through the asymmetric-information mechanics of The Resistance: Avalon. Rather than evaluating end-to-end gameplay, it decomposes ToM into a 2$\times$2 taxonomy -- epistemic versus motivational reasoning crossed with inference versus action -- using human-crafted, perspective-constrained queries. Benchmarking 28 LLMs reveals three insights: 1) Reasoning, not knowledge. Models show strong game-rule comprehension but markedly weaker ToM abilities, isolating failures to social reasoning rather than missing domain knowledge. 2) Expression, not representation. Mechanistic analyses via linear probing and activation steering show that models frequen

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

arXiv:2608.09548v1 Announce Type: cross Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses education-facing suitability as an integrated profile. We introduce ELBench, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. We evaluate nine models, seven frontier general-purpose systems and two education-specialized variants, and report three findings. First, module-level profi

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

When Can Text Embeddings Replace Item Calibration? A Geometric Diagnostic for Semantic Loadings in Multidimensional Adaptive Testing

arXiv:2608.09058v1 Announce Type: cross Abstract: Multidimensional item response theory relies on calibrated item parameters, such as discrimination and category threshold values, which are usually estimated from large samples of human test responses. This study investigates whether the directional loadings of these parameters can be recovered directly from item text using pre-trained sentence embeddings, avoiding the need for initial item calibration. Using the open-source IPIP Big-Five dataset ($n=19{,}719$; 50 items), we built a multidimensional computerized adaptive testing (CAT) simulation using D-optimal item selection. We compared three item loading sources: fitted graded response model parameters, semantic text embeddings, and a lexical baseline. In simulation, semantic embeddings recovered latent trait profiles almost as accurately as fitted parameters (correlation $0.825$ vs. $0.857$), performing noticeably better than simple word overlap ($0.752$). However, the embedding-bas

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

arXiv:2608.09046v1 Announce Type: cross Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a model is invoked. We introduce the Tokenization Equity Audit (TEA), a reproducible benchmark for measuring tokenization premiums in technical tutoring content. TEA evaluates three widely used tokenizers, GPT-4o's o200k base, Qwen2.5-7B, and Mistral-7B, on a 120-item Python debugging corpus translated from English into Bengali, Hindi, Arabic, Tamil, and Yoruba. Bengali and Hindi serve as the primary validated cases, while the remaining languages provide exploratory cross-script and cross-family comparisons. Across this corpus, Bengali requires

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

How People Evaluate AI-, Expert-, and Peer-Style Financial Advice

arXiv:2608.09019v1 Announce Type: cross Abstract: As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (N = 285) in which substantive financial content---including facts, numerical values, recommendation direction, and core reasoning---was held constant while communication style varied across AI Financial Assistant (AI), Certified Financial Planner (Expert), and Online Community Forum (OC) advice. Displayed source attribution was independently manipulated through correctly labeled, unlabeled, and mislabeled conditions, allowing us to separate attribution effects from source-specific communication cues. Expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|d|=0.20--0.47), and this advantage remained visible without source labels, where Expert advice outperformed AI advice on 8 of 10

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning

arXiv:2608.08887v1 Announce Type: cross Abstract: Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems often use separate solutions for facial recognition, vehicle identification, fire detection, and behavioral analysis, resulting in fragmented infrastructure and multiple interfaces for operators to manage. This paper presents City Sentinel, a unified AI-based surveillance framework that integrates six detection capabilities into one scalable platform: facial recognition, automatic number plate recognition (ANPR), fire and smoke detection, weapon and knife detection, violence detection, and road accident detection. The system combines a Next.js operator dashboard, FastAPI backend, cloud-based PostgreSQL event storage, InsightFace and YOLOv8 vision models, and EasyOCR for plate recognition. Camera streams are processed through dedicated inference workers using RTSP. On a workstation equippe

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

arXiv:2608.08852v1 Announce Type: cross Abstract: AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fit a specified learner, which education calls Pedagogical Content Knowledge (PCK), has not been benchmarked. To measure it, we introduce the Teaching Monster Challenge, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion. Each system is given a topic and a learner persona and must generate a complete instructional video. Every video is screened by an LLM-judge, ranked by crowd pairwise voting, and finalized by an expert panel. The first edition shows that today's systems handle the content well but are far weaker at presenting it and adapting it to the learner. The same process exposes a limit of automatic judging. The LLM-judge separates a clear low-performing tail but ranks the strongest systems poorly. The strongest system

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Distribution Mapping Approach to Counterfactually Fair Reinforcement Learning

arXiv:2608.08743v1 Announce Type: cross Abstract: Reinforcement learning (RL) seeks to optimize sequential decisions to maximize population-level benefits over time. However, when deployed in high-stakes settings such as healthcare, RL decisions might systematically restrict some subpopulation's access to valuable services in a manner contrary to the values and goals of stakeholders. Counterfactual fairness (CF) offers a promising framework to address this problem based on causal reasoning. This paper develops a data preprocessing algorithm that, when used in tandem with policy learning, enables CF in RL. Our algorithm relies on a novel quantile distribution mapping method for sequentially estimating the counterfactual states and rewards in the data preprocessing step, subsuming common additivity assumptions used for counterfactual prediction as a special case. We theoretically prove that the per-step level of counterfactual unfairness and infinite-horizon suboptimality gap can be boun

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Beyond Aggregate Calibration: Decomposing Income-Conditional Recall Disparities in Automated Credit Default Prediction

arXiv:2608.08202v1 Announce Type: cross Abstract: Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances. Evaluating this filtering convention on a large-scale consumer lending sample (LendingClub, N = 1,344,936) uncovers an underlying demographic asymmetry: high-income defaulters are disproportionately classified as label noise relative to low-income defaulters (Cramer's V approximately 0.03-0.07). Re-examining this behavior through the lens of equal opportunity [Hardt et al., 2016] reveals a far more severe discrepancy: a 16.86 percentage point gap in true positive rate (recall) between high- and low-income borrowers who ultimately defaulted. Implementing a sequential feature-blinding methodology allows us to isolate the drivers of this disparity across three distinct mechanisms: (1) direct reliance on self-reported applicant income; (2) algorithmic absorption of upstream institutional bias encoded within o

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Legal Responsibilities Using Autonomous Agents For Artificial Intelligence

arXiv:2608.08022v1 Announce Type: cross Abstract: Recent incidents involving Artificial Intelligence (AI) agents, which were reported escaping their containment `unintentionally' to gain unauthorized access, pose looming questions about who or what should be held legally responsible for resultant criminal or negligent damage. As the independent capabilities of agents expand, Promise Theory suggests a systematic method to resolve these questions, based on the Downstream Principle for causal influence. Responsibility can easily be expanded to include AI agents where tracing responsibility becomes impactical, and agents' freedoms to act can be limtied by policy choices.

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Blueprint for Collaborative Cybersecurity Operations Centres with Capacity for Shared Situational Awareness, Coordinated Response, and Joint Preparedness

arXiv:2608.08011v1 Announce Type: cross Abstract: With digital technologies now being part of the fabric of our societies, identifying and managing cybersecurity threats becomes imperative. Within the European Union, several initiatives are underway, aiming to motivate, regulate and eventually orchestrate the establishment of capacity and enhancement of situational awareness, incident response, and preparedness capabilities, with an expected emphasis on operators of essential services and state actors entrusted with cybersecurity. In this context, the institution of cooperation and information exchange channels to allow for coordinated cross-border responses to large-scale incidents is particularly prioritised. Motivated by the above, this work presents a conceptual blueprint in support of architecting and establishing interoperable Cyber Security Operations Centres that combine capacity for situational awareness, incident response, and preparedness, also benefiting from the interplay

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Indirect Geoeconomic Influence: A Switching Dynamical Systems Framework for Mechanism Design

arXiv:2608.07940v1 Announce Type: cross Abstract: We develop a formal framework for analyzing indirect geoeconomic influence. The influencing state (sender) does not attempt to change a target nation's policy directly. Instead, the sender restructures the target's internal political economy so that its own citizens, firms, and institutions generate the compliance pressure. The framework rests on a switching dynamical system (SDS) in which a target's political economy evolves under mode-dependent rules. We analyze two modes: a permissive mode, in which a mechanism transmits pressure toward the sender's preferred policy, and a contested mode, entered naturally once the target detects and attributes the mechanism. Crucially, the sender's mechanism design shapes the transition into the contested mode rather than paying a static toll for legibility. This inverts the usual regime-switching problem: rather than estimating a latent transition kernel from data, the designer engineers the kernel

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Crowd-Sourced Geographies of Income: Using Google Maps Points of Interest as High-Frequency Proxies for Sub-Municipal Income Estimation in Sao Paulo, Brazil

arXiv:2608.07871v1 Announce Type: cross Abstract: Accurate, up-to-date income data at the sub-municipal scale is essential for social policy in middle-income countries, yet in Brazil it depends on a costly decennial census whose intercensal gap recently exceeded a decade. We test whether the composition of crowd-sourced Google Maps Points of Interest (POIs) can serve as a high-frequency, low-cost proxy for household income across the 26,625 census sectors of the municipality of Sao Paulo. Using a theoretically motivated set of POI categories retrieved from Google Places, we represent each sector by its POI counts, decompose these high-dimensional, sparse features with principal component analysis (PCA) and non-negative matrix factorization (NMF), and train a sweep of regression models to predict census-derived income. Under a data leakage-aware spatial validation design the best model (NMF with gradient boosting) attains a held-out R^2 of 0.65, with performance stable across feature-ex

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

China RealDID: Verifiable Credentials Anchored in Legal Identity

arXiv:2608.07846v1 Announce Type: cross Abstract: Verifiable credentials (VCs) and decentralized identifiers (DIDs) enable selective disclosure but lack legal anchoring: without a trusted identity root, verifiers cannot distinguish a genuine holder from a fabricated identity. State identity systems provide biometric-grounded verification but impose three costs: verifiers must collect subjects' full personally identifiable information, infrastructure concentrates on a single API, and the state observes every transaction. We present China RealDID, a three-layer architecture -- CTID (centralized legal identity), RealDID (decentralized anchor on an open permissioned blockchain), and VCs with SD-JWT-based selective disclosure -- evaluated against five adversary classes and six security goals. The central mechanism is a content-blind government relay: the state authenticates participants and counter-signs every credential but cannot read the payload, encrypted by the issuer to the holder's p

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

"Always Want to Use it for Everything": Understanding Young Adults' Perceptions of AI Dependence

arXiv:2608.07592v1 Announce Type: cross Abstract: The growing integration of general-purpose AI chatbots into people's daily lives has raised concerns about the potential for unhealthy dependence, particularly among young adults. As a first step toward understanding and characterizing AI chatbot dependence from the perspective of young adults, we collected testimonials from AI chatbot users aged 18 to 25 through an online questionnaire to capture their thoughts and experiences with this phenomenon. From participant responses, we identified three contributing factors of AI dependence: chronic use, efficiency, and delegation. The combination of these in a person's interaction behavior was considered to indicate AI dependence. Participants also observed feelings of atrophy in abilities from AI dependence, leading to psychological impacts such as feelings of inadequacy. Interpreting these findings through the lens of self-determination theory reveals how AI chatbot dependence can impact yo

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

arXiv:2608.07535v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this architectural shift reshapes the safety landscape of machine learning. Increased model complexity and cross-modal interactions give rise to novel threats, including compromised modality integration, modality misalignment, and fused safety risks, reflecting shifts in threat modeling beyond uni-modal assumptions. These shifts, in turn, impose new constraints on safety solutions not captured by existing frameworks rooted in uni-modal learning. Motivated by these challenges, this survey provides a systematic analysis of the evolving safety landscape of MLLMs. We first propose a multimodal grounded taxonomy of safety threats and analyze shifts in threat models, covering adversarial attacks, data poisoning, jailbreaks, and hallucinations. We then summarize updated safe

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Democratizing Ski Safety: Real-Time Turn Segmentation with Smartphone IMU and Causal LSTM Networks

arXiv:2608.07513v1 Announce Type: cross Abstract: Anterior cruciate ligament (ACL) injury is one of the most common and serious injuries in sports, particularly among recreational skiers. Research shows that structured technique awareness and continuous feedback can significantly reduce the risk of such injuries, yet access to professional instructors is limited to wealthy athletes who can afford continuous private coaching, creating a harmful inequity in injury prevention. This gap can be mitigated by automating the real-time analysis of skiing techniques available to the wider recreational skiing community. The approach relies exclusively on inertial sensors embedded in standard smartphones, eliminating the need for specialized equipment and enabling broad social scalability. To support immediate feedback, the system operates causally, producing predictions based solely on past observations. The work is conducted in cooperation with professional ski instructors, ensuring that problem

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare

arXiv:2608.07511v1 Announce Type: cross Abstract: Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have been proposed for tasks such as triaging patients, report drafting or translating medical jargon to support informed decision-making. These applications require both factual and social competence. This study evaluates dialogues between LLMs and participants to assess the current state of socio-communicative competencies displayed in LLM-generated texts. Methods. We extracted a subset of extended dialogues from the HELP-Med dataset, comprising 1800 conversation transcripts of interactions between human participants seeking medical information and three different LLMs, GPT 4o, Llama 3 and Command R+. Two experts coded the transcripts for demonstrations of socio-communicative behaviours (non-hostility, sensitivity, structuring, non-intrusiveness) using the IC-MD instrument, originally designed

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

World Simulator: Queer Erotica and the Absurdity of AI Video Models That Promise the World

arXiv:2608.07510v1 Announce Type: cross Abstract: Increasingly, AI video models are marketed as "world simulators," suggesting their ability to model infinite realities. Despite such claims, these models systematically exclude significant aspects of embodied human experience, particularly sexuality. World Simulator is a video installation exploring the poetic friction between these universal claims and the models' inherent blindness. To do so, the work feeds explicit gay erotica into an AI video-to-video pipeline. Lacking the training data to recognize these images, the system hallucinates surreal alternatives, transforming intimate acts into banal scenes of kitchen appliances, strange architectures, and abstract flesh. By visualizing the limits of synthetic knowledge, the work challenges the hubris of the "world simulator" label, asking how a system, trained primarily on large filtered video datasets, can claim to simulate the world while remaining structurally blind to the body. Beyo

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Large Language Models Explain Experts Better Than Experts Themselves

arXiv:2608.07488v1 Announce Type: cross Abstract: Tacit knowledge, or the "know-how" embedded in experience, is difficult to articulate, making its transfer a challenge in organizations. Tacit knowledge is hard to externalize (transform into explicit knowledge), and expertise is often poorly documented and lost when experts leave. This study examines whether LLMs can externalize tacit knowledge from experts' behaviors and whether such externalized knowledge supports downstream decision-making and transfer to novices. Across two studies, we show that LLM-externalized tacit knowledge improves decision quality and enables novices to approach expert-level performance, often outperforming knowledge articulated by human experts. These findings provide empirical support for Polanyi's Paradox -- that we can know more than we can tell -- and highlight the potential of LLMs as scalable tools that can help overcome human experts' articulation bottleneck. Mechanism analyses and robustness checks s

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

SafeStudent Driving: A Multimodal Driver-Safety System to Support Teen Drivers Using Computer Vision and Mobile Sensing

arXiv:2608.07487v1 Announce Type: cross Abstract: Teen drivers face disproportionately high crash rates, often due to inexperience and inconsistent attention to basic traffic rules. SafeStudent Driving addresses this problem with a multimodal coaching system deployed on both a Raspberry Pi device and a Flutter-based mobile app. The system uses three YOLO-based computer-vision models to detect traffic lights, light-bulb colors, and road signs, an OCR module to read speed-limit values, and an audio model plus IMU data to infer whether turn signals are used during turns. An analysis layer smooths detections over time and triggers prioritized voice prompts through text-to-speech or pre-recorded audio. Key challenges included achieving sufficient model accuracy in varied lighting, running inference fast enough on limited hardware, and designing prompts that inform without distracting the driver [3]. Experiments on sign detection and turn-signal recognition highlight strengths and failure mo

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Humour as Resistance: Visceralizing the Environmental and Social Impact of AI through Humour-based Creative Practices

arXiv:2608.07485v1 Announce Type: cross Abstract: The growth of AI does not come without cost. While we hear about the economic costs, the environmental and social costs are often obfuscated by mainstream narratives, and even when acknowledged, are accompanied by a sense of helplessness. To counter this, we, a group of designers and researchers, reflect on our experience leading a humour-based creative campaign surfacing the material impact of AI infrastructures. Through graphics design, physical installations, digital content creation and co-creation workshops, we leaned into humour as a creative practice to provoke collective reflection. Drawing on event ethnography, surveys and interviews with campaign engagers, we examine four roles of humour-based creative work in HCI: a connector to critical friends, a visceral and emotional harbour, social glue, and resistance to power. We argue for the importance of creative practices in bridging social and emotional gaps between people and con

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains

arXiv:2608.07474v1 Announce Type: cross Abstract: Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds human cognitive capacity C_max. The operative constraint, however, is not V alone but V x L, where L denotes per-item cognitive load. L consists of triage, judgment, and response, which respond asymmetrically to AI capability improvement. Triage cost does not decline as models become more capable, because semantic indeterminacy is inherent in general-purpose design. Response cost is invariant to accuracy improvements. Only judgment cost faces downward pressure, and this pressure often operates by inducing omission rather than genuine reduction. Capability improvement therefore restructures L rather than reducing it. Governance mechanisms based on evaluating whether AI output is correct either delegate that evaluation to AI and inherit hallucination risk, or delegate it to humans and face the V x L ceil

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Application of Artificial Intelligence for Fraudulent Banking Operations Recognition

arXiv:2608.07471v1 Announce Type: cross Abstract: This study considers the task of applying artificial intelligence to recognize bank fraud. In recent years, due to the COVID19 pandemic, bank fraud has become even more common due to the massive transition of many operations to online platforms and the creation of many charitable funds that criminals can use to deceive users. The present work focuses on machine learning algorithms as a tool well suited for analyzing and recognizing online banking transactions. The study`s scientific novelty is the development of machine learning models for identifying fraudulent banking transactions and techniques for preprocessing bank data for further comparison and selection of the best results. This paper also details various methods for improving detection accuracy, i.e., handling highly imbalanced datasets, feature transformation, and feature engineering. The proposed model, which is based on an artificial neural network, effectively improves the

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Foundational values for foundation models

arXiv:2608.09377v1 Announce Type: new Abstract: Research values, properties with a distinctive normative dimension, often affect how technological research is performed in both direct and indirect ways by influencing how technical decisions are made. In machine learning for medical imaging, understanding these values can be important for understanding why particular researchers justify the decisions made in their publications and explain why certain technologies become ubiquitous (or not) in the scientific literature and in the clinic. This article explores one of these technologies, foundation models, finding detailed justifications both for their use and abstention from their use. By taking a Socratic approach to research values arising from this specific technical decision, this article aims to better illustrate how foundation models fit into the philosophy of machine learning in medicine.

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI-AI co-creation outperforms human pairs in creative tasks

arXiv:2608.09023v1 Announce Type: new Abstract: Prior research often finds that AI creativity is limited: single systems rarely outperform humans, and human-AI collaboration does not exceed human output. We argue these conclusions underestimate AI's potential because most studies do not allow iterative, multi-agent exchanges that mirror the social processes underpinning human creativity. We conducted an experiment comparing four conditions: (i) AI-AI co-creation with complementary generator-evaluator roles, (ii) AI-AI co-creation with identical roles, (iii) single-AI creation, and (iv) human-human co-creation. Across three open-ended tasks, 1,212 ideas were rated by trained judges on creativity, novelty, and usefulness. Both AI-AI co-creation conditions consistently outperformed single-AI creation and human pairs on creativity and novelty. Usefulness varied by task: complementary roles yielded the most useful solutions in the broadest and most socially complex task, suggesting role dif

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Data Findability, Governance, and Community Engagement for M\=aori Research Data Sovereignty

arXiv:2608.08905v1 Announce Type: new Abstract: M\=aori data sovereignty (MDSov) has established important principles for recognising M\=aori rights and interests in data. Relatively little attention has been given to how these principles can be operationalised within research institutions who, as a function of their operations, collect and use M\=aori data. This paper introduces the concept of M\=aori Research Data Sovereignty (MRDSov), extending existing understandings of MDSov into the specific context of research data and the research data lifecycle. Drawing on Indigenous Data Sovereignty scholarship and research data management literature, we define M\=aori research data and position MRDSov as the application of MDSov principles to research data produced by, about, or for M\=aori. We argue that operationalising MRDSov requires three interdependent elements: data findability, data governance, and community engagement. Data findability enables M\=aori research data to be identified

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Automating Freshman Course Placement and Registration: A Case Study

arXiv:2608.08776v1 Announce Type: new Abstract: This implementation report explores Rowan University's effort to automate the process of freshman course placement and registration. Historically, Freshman Instructional Guides (FIGS) at Rowan was manually executed, requiring significant time from Testing Services, University Advising, and the Registrar's Office to evaluate placement needs and assign students to courses. Given the 57% surge in first-time degree-seeking student enrollment over a decade, the manual processes became increasingly unsustainable. In response, a cross-departmental team developed a comprehensive automated process to integrate data from Banner (Student Information System), Google Sheets maintained by Advising, and other sources. This automated process classifies students based on program groupings, determines primary and secondary course placements, checks for real-time availability and constraints in Banner, and completes course registration for freshmen in bulk.

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Qualifying and Quantifying Risk under the EU AI Act

arXiv:2608.08564v1 Announce Type: new Abstract: The EU AI Act uses a risk-based approach to regulate AI systems, calibrating the intensity of regulation according to the risks they pose. While the AI Act's definition of 'risk' (as the combination of the probability and severity of harm) implies quantification, the AI Act focuses on risks to fundamental rights, thereby engaging a typically qualitative perspective. In this piece, we address this tension using a two-step framework under which the EU AI Act balances the protection of fundamental rights, the legitimate purposes of providers and deployers, and the impacts of regulatory measures on providers, deployers, and regulators. We discuss this framework against the backdrop of potential approaches to quantifying risks, with a specific focus on defining and measuring the main components of the concept of risk: 'probability', 'severity', and their 'combination'. We suggest that the protection of fundamental rights and risk quantificatio

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Rethinking Higher Education: From Fixed Curricula to Learnity Graphs

arXiv:2608.08543v1 Announce Type: new Abstract: Higher education stands at a turning point. In an era where knowledge is increasingly accessible and which is, more often than not, mediated by advanced Artificial Intelligence (AI), the value of traditional curricula models warrants reconsideration. This does not imply that one should replace thorough academic studies. Universities remain essential in providing foundational knowledge, theoretical depth and conceptual grounding. The challenge is to extend these educational facets with learning environments that foster creativity, interdisciplinary integration, hands-on experience, and especially long-term development. In this paper, we introduce a lifelong learning framework that integrates academic, professional, and personal learning, centered on a new concept that we term learnity graphs, a structured representation of learning as interconnected units of knowledge, skills, experience, and actual artifacts, coupled with a method for pre

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Abstracted Away: Resisting Alienation and Ungrounded Abstraction in AI Research Communities

arXiv:2608.08408v1 Announce Type: new Abstract: Logics of abstraction in computational AI research often push important forms of knowledge and reflection aside: dominant standards of legitimacy separate from lived experience of harm; the goals of work misalign with the practices that operationalize them; and career demands crowd out critical reflection. Even as prior academic and community-oriented efforts have sought to recontextualize and challenge common practices, exposure to sociotechnical harms and epistemic injustice persists. As three early-career critical AI researchers, we experienced this as alienation: feeling like outsiders in our research communities. This alienation has involved having some aspects of our backgrounds overlooked and others tokenized. We argue our alienation occurred through mechanisms that mirror abstraction by creating distance from relevant material realities. Beyond abstraction's role in computational AI research as a foundational practice structuring

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Representational Equality in Cross-country Value Simulation: A Systematic Analysis of Large Language Models

arXiv:2608.08058v1 Announce Type: new Abstract: Traditional methods for studying human opinions often struggle to support representative and scalable research across countries. Large language models (LLMs) can serve as scalable proxies for simulating human opinions, enabling more efficient opinion analysis. However, this use of LLMs requires not only high average accuracy but also representational equality, that is, comparable simulation accuracy across populations. Uneven simulation accuracy may reproduce or amplify societal biases in downstream applications. This study systematically investigates country-level representational equality across 59 countries and finds substantial, systematic inequality. Populations from wealthier and more technologically advanced countries are simulated more accurately. We further compare two foundational intervention pathways, contextual adaptation and parametric modification, and show that improvements in average or target-group accuracy do not necess

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety

arXiv:2608.07902v1 Announce Type: new Abstract: Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations. However, existing child safety evaluations of AI lack grounding in real-world harms that youth experience, rely on unvalidated assumptions about what counts as an appropriate output (e.g., refusal), and typically focus on detecting adversarial prompts or surface-level harms in outputs only. Thus, these evaluations can fail to detect responses that pose harm to youth in practice. To better understand the limitations of current evaluation practices, we conducted interviews with 19 practitioners working directly with youth in vulnerable situations, including social workers, therapists, and psychologists, asking them to reflect on chatbots' responses to risky situations commonly faced by youth, as established in prior empirical work. Practitioners identified chatbot behaviors likely t

Source ↗
Showing 651–700 of 1593 signals
← Prev Page 14 of 32 Next →