EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning

arXiv:2604.22770v2 Announce Type: replace Abstract: Most digital language learning curricula rely on discrete-item quizzes that test recall rather than applied conversational proficiency. When progression is driven by quiz performance, learners can advance despite persistent gaps in using grammar and vocabulary during interaction. Recent work on LLM-based judging suggests a path toward scoring open-ended conversations, but using interaction evidence to drive progression and review requires scoring protocols that are reliable and validated. We introduce Learning in Blocks, a framework that grounds progression in demonstrated conversational competence evaluated using CEFR-aligned rubrics. The framework employs heterogeneous multi-agent debate (HeteroMAD) in two stages: a scoring stage where role-specialized agents independently evaluate Grammar, Vocabulary, and Interactive Communication, engage in debate to address conflicting judgments, and a judge synthesizes consensus scores; and a re

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Digital Engagement, Income Disparities, and Job Seeking in the United States since 2010

arXiv:2511.05294v2 Announce Type: replace Abstract: Surveys often record how frequently people use the internet without measuring the infrastructures, skills, and support systems that make digital participation possible. Using the U.S. National Longitudinal Survey of Youth 1997 cohort, we study how internet-use frequency relates to labor income, employment attachment, and job seeking after 2010. The main digital-engagement analysis uses the comparable 2011, 2013, and 2015 waves, with 2017 retained as later labor-market context. Across repeated cross sections, daily internet use consistently marks higher income and stronger employment attachment. Relative to daily use, less-than-daily use is associated with roughly 11 to 20 percent lower income, while nonuse is associated with about 18 to 21 percent lower income in 2011 and 2013. Respondents reporting no internet use are also 13 to 23 percentage points less likely to report full-year work. Job-search estimates reveal a distinct mechanis

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Evolving Media Discourse on ChatGPT and Higher Education

arXiv:2508.14692v2 Announce Type: replace Abstract: This research full paper examines how news media have been instrumental in creating specific narratives about generative AI applications, especially ChatGPT, in higher education, and how these narratives have changed over time. The introduction of emerging technologies in higher education is driven not only by their technological affordances but also by the narratives built around their perceived value, risks, and possibilities. Therefore, understanding how news media narratives contribute to sociotechnical imaginaries - the imagined futures of technology use that institutions and educators inherit - is important for evaluating ChatGPT's role in teaching and learning, including engineering education. Through temporal and sentiment analyses of 198 U.S. news articles from November 2022 to October 2024, we traced the evolving narratives surrounding generative AI and the use of ChatGPT in higher education. We found that the media discours

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

LLM-Based Social Simulations Require a Boundary

arXiv:2506.19806v3 Announce Type: replace Abstract: This position paper argues that LLM-based social simulations require clear boundaries to make meaningful contributions to social science. While Large Language Models (LLMs) offer promising capabilities for simulating human behavior, their tendency to produce homogeneous outputs, acting as an "average persona", fundamentally limits their ability to capture the behavioral diversity essential for complex social dynamics. We examine why heterogeneity matters for social simulations and how current LLMs fall short, analyzing the relationship between mean alignment and variance in LLM-generated behaviors. Through a systematic review of representative studies, we find that validation practices often fail to match the heterogeneity requirements of research questions: while most papers include ground truth comparisons, fewer than half explicitly assess behavioral variance, and most that do report lower variance than human populations. We propos

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Patterns and Purposes: A Cross-Journal Analysis of AI Tool Usage in Academic Writing

arXiv:2502.00632v2 Announce Type: replace Abstract: This study investigates the use of AI tools in academic writing through an analysis of AI usage declarations in journals. Using a mixed-methods approach combining content analysis, statistical analysis, and text mining, this study analyzed 135 AI declarations from 8633 articles across 27 categories. Results show that ChatGPT dominates academic writing assistance (73.3 percent usage). The primary purposes of AI integration are concentrated on lower-level cognitive tasks, specifically improving readability (57.8 percent) and grammar checking (19.3 percent). Statistical analysis indicates a highly significant association between team composition and AI-use purposes (p = 0.0008), highlighting international teams' reliance on grammar assistance, while no significant association was found regarding authors' native-speaker status (p = 0.2359). These findings provide insights for journal policy development and for understanding the evolving r

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Automated Textbook Auditing with Multi-Agent LLM Systems

arXiv:2607.11276v1 Announce Type: cross Abstract: Ensuring the quality of educational materials requires more than standard proofreading: textbooks must be audited for factual accuracy, domain-specific technical correctness, and linguistic quality simultaneously -- a task that general-purpose grammar checkers cannot address. We present \textbf{AI Textbook Auditor}, a modular multi-agent pipeline for automated quality assurance of educational materials across subject domains. The system accepts a textbook PDF and produces a structured, human-reviewable report via two analysis tracks: a \textbf{Factual and Technical Track} in which an ensemble of specialized LLM agents detects factual inaccuracies, code errors, incorrect definitions, and conceptual inconsistencies, augmented with web search for humanities domains; and a \textbf{Grammar Track} operating PDF-natively to preserve diacritical encoding. A \textbf{Judge Agent} filters false positives using domain-specific rules before presenti

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Do LLMs Fabricate Legal Citations? A Bilingual Benchmark on Saudi Data Protection Law and the GDPR

arXiv:2607.11127v1 Announce Type: cross Abstract: Organizations and regulators increasingly consult large language models (LLMs) for regulatory-compliance questions, yet a wrong statutory citation can silently propagate into legal advice, compliance documentation, and policy decisions. We introduce a bilingual benchmark of 120 questions probing whether freely accessible LLMs fabricate article citations for two data-protection instruments: the EU General Data Protection Regulation (GDPR) and the Saudi Personal Data Protection Law (PDPL). The benchmark pairs direct citation retrieval questions with false premise verification probes and deliberately unanswerable "trap" questions -- including questions about a repealed article and about deadlines that exist only in implementing regulations, not in the law itself. Every question is posed in both Arabic and English, and all scoring is fully automatic against a manually verified gold reference. Evaluating three freely accessible models (Gemin

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Verifier-Centric Conceptual Model for Digital Credential Ecosystems

arXiv:2607.10747v1 Announce Type: cross Abstract: Digital credential ecosystems increasingly combine multiple standards. Because implementations have evolved independently across jurisdictions and application domains, systems described under the common label ``digital credential'' often remain mutually non-interoperable. Conventional element-by-element comparisons of identifiers, data models, credential formats, protocols, and signature algorithms do not explain why interoperability fails even when stacks share a data model, nor do they identify what a verifier must obtain, and what it must trust, before accepting a credential. We present a verifier-centric conceptual model built on two decompositions. The first separates credential processing into signature verification (L1), semantic interpretation (L2), and validation (L3), and models the supporting materials through two orthogonal planes: Constitution, which captures ecosystem-level arrangements and trust declarations, and Logistic

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Commenting with Copilot: A Taxonomy and Multi-Year Analysis of Student Code-Generation Specifications

arXiv:2607.10674v1 Announce Type: cross Abstract: As AI code tools become integrated into programming environments, students increasingly describe intended behavior in natural language and rely on these tools to generate code, shifting emphasis from code writing to specification. Yet little is known about the comments students write as specifications in AI-assisted programming tasks. We analyze a four-year dataset of undergraduate programming submissions and reflections from tasks in which students wrote comments to guide code generation and refined solutions using test-case feedback. We introduce a taxonomy spanning three dimensions: comment type, code expression level, and code construct. Using automated classification, we examine how these dimensions vary across attempts and how students describe the process in their reflections. Our findings show that students mostly wrote natural-language What comments, shifted toward How comments for more procedural constructs, and focused more o

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator Practices in Artistic Image Generation

arXiv:2607.10538v1 Announce Type: cross Abstract: The open-sourcing of powerful image generation models has created a vibrant ecosystem where creators curate and combine a vast array of community-contributed models. This practice stands in sharp contrast to using closed-source tools like Midjourney. Yet, little is known about these emerging creative workflows. To bridge this gap, this paper presents the first large-scale empirical study of creator model usage behavior within this open-source image generation ecosystem. We construct a novel dataset of 6 million images with their embedded generation metadata -- a detailed recipe of the creation process, including the models used and the prompts. By linking the usage of 22.4K base models and 154K LoRA models to the images, our findings underscore the ecosystem's unique strengths and its inherent obstacles. This provides valuable insights for making this ecosystem more sustainable and innovative. Moreover, we make our dataset publicly avai

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

How Data Narratives Go Wrong: A Taxonomy of Issues Across the Data Communication Process

arXiv:2607.10523v1 Announce Type: cross Abstract: Data narratives increasingly shape public understanding, but their failures are rarely just isolated factual errors or deceptive charts. Instead, they emerge through a broader meaning-making process in which quantitative evidence is transformed into claims, representations, and arguments. While prior work has examined these failures across disparate fields (e.g., statistics, visualization, and fact-checking), the community lacks a holistic lens to explain how these issues arise, propagate, and compound. To address this gap, we introduce TIC, a Taxonomy of Issues in Data Communication, synthesized from prior literature and refined through the qualitative annotation of 700 real-world data narratives from fact-checking sites, research datasets, and controversial media. TIC organizes recurring breakdowns across six dimensions-data, analysis, visual encoding, text, reasoning, and interpretation-and situates them within a framework spanning a

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Comparing Socio-technical Design Principles with Guidelines for Human-centered AI

arXiv:2607.10331v1 Announce Type: cross Abstract: Human-centered AI (HCAI) refers to guidelines or principles that aim on ethi-cally oriented design of systems. We compare HCAI- guidelines with princi-ples of socio-technical systems that emerged in the context of conventional in-formation technology. The comparison leads to a revision of socio-technical heuristics by including aspects of AI-usage. The comparison reveals that con-tinuous evolution is a basic characteristic of socio-technical systems, and that human oversight or interventions and the subsequent appropriation of AI-systems lead to continuous adaptation and re-design of the systems, if autono-my is collaboratively exercised. From a socio-technical point of view, the cru-cial requirement of transparency has not only to be fulfilled with technical fea-tures, but also by contributions of the whole system including human actors. It will be promising for using AI, if not only technical features, but organization-al and social p

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Intervenability as a Design Requirement for Autonomy and Oversight within Human-Centered AI

arXiv:2607.10322v1 Announce Type: cross Abstract: Based on the literature and several practical examples of possible AI applica-tions, we outline the concept of intervenability. This new phenomenon is not covered by emergency shutdowns, workarounds, or the reconfiguration of automated systems. Intervenability instantiates the principles of control-lability, autonomy, oversight, and keeping humans in the loop in the context of AI. We provide a taxonomy that encompasses a range of possibilities for intervening activities and differentiates them regarding the mental effort of the users. This taxonomy extends the scope of interventions from real-time control of automated processes to AI-based discrete case-related decision-making. This is in accordance with human-centered AI, which seeks to combine human strengths with the usage of AI. We demonstrate how inter-venability can potentially contribute to the ongoing development of human capabilities on the one hand and to further technical imp

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Neutralizing Structural Inequality in the Nigerian FinTech Sector

arXiv:2607.10317v1 Announce Type: cross Abstract: Algorithmic decision systems in financial services often rely on data proxies that inadvertently encode structural inequalities. This paper introduces a hierarchical human-AI triage model for Point of Sale fraud detection in the Nigerian FinTech sector. Adopting a We Are All Equal worldview, we address the challenge of discrimination laundering, wherein the system misinterprets infrastructure related aleatoric noise such as rural network timeouts as fraudulent intent. We implement a three-tier routing policy utilizing a calibrated ensemble model as a primary filter. The policy routes transactions characterized by epistemic uncertainty such as cold start new accounts to specialist analysts while reserving high stakes cases for a senior supervisor. To manage finite human capacity, we utilize a dynamic shadow price to ration human attention and implement a random audit mechanism to prevent human skill atrophy. Our experimental results demo

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

From Patent Expiry to Business Pathways: AI Workflows for Activating Innovation Archives

arXiv:2607.10179v1 Announce Type: cross Abstract: Patent databases represent one of the largest public archives of technical knowledge, yet much of this knowledge remains difficult to identify, interpret, and reuse once patent rights expire or lapse. This paper proposes an AI-enabled framework for discovering expired and lapsing patents, identifying technology trends, and translating patent disclosures into business pathways. We use pathways to mean structured commercialization routes such as SaaS products, services, licensing packages, consulting playbooks, training offerings, data products, or internal process tools. The framework treats patent expiry as both a business signal and an archival transition, not primarily as a legal problem. Legal status remains important, but it is one risk-screening input alongside customer need, implementation feasibility, channel access, and market timing. We describe a system architecture that combines patent metadata, maintenance-fee records, legal

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Coupled Tensor-Matrix Recovery via Proximal Alternating Linearized Minimization, with an Application to Workforce Skill and Small-Business Health Estimation

arXiv:2607.10163v1 Announce Type: cross Abstract: We study recovery of a low-rank tensor $\mathcal{T}$ and a low-rank matrix $M$ from sparse, noisy observations. $\mathcal{T}$ and $M$ share one mode. We relax tensor rank using the nuclear norm of the mode-1 unfolding. This unfolding carries the coupling. It also has an exact proximal operator. We couple $\mathcal{T}$ and $M$ through a learned linear operator $G$. We prove a minimizer exists for the ridge-stabilized penalized objective. We prove that a proximal alternating linearized minimization (PALM) scheme converges to a critical point, for the algorithm as implemented, by verifying the hypotheses of a known nonconvex block-coordinate convergence theorem against our objective and identifying which conditions come from this problem's structure. For the matrix-only sub-problem, we state a proven sampling bound from matrix completion theory. For the coupled problem, we prove a sample-complexity result for a sequential sub-case: a separ

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Evaluating AI Models' Capability to Automate Voice Phishing Attacks

arXiv:2607.09970v1 Announce Type: cross Abstract: Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults' susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play.AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the "relative-in-distress" category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems

arXiv:2607.09790v1 Announce Type: cross Abstract: The article investigates the fundamental problem of ensuring the stability of operator control and preserving goal-targeting in hybrid human-machine decision support systems (DSS) of a new generation. Based on a two-month continuous longitudinal experiment on the joint design of a monograph-format textual array, the latent phenomenon of semantic context drift in large language models of deep logical reasoning (Reasoning LLMs) is verified and described. A mathematical model of interaction in the human-machine interface is proposed, and an original metric is introduced - the operator control stability coefficient, which takes into account the non-linear contextual pressure of hidden reasoning chains. Within the paradigm of the cognitome theory, a critical point of control functions inversion is captured. Engineering recommendations are formulated for implementing dynamic relational arbitration loops based on a modified hierarchical simila

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Model Collapse: On Recursion, Noise, and Uncharted Machine Visions

arXiv:2607.09705v1 Announce Type: cross Abstract: Since 2023, computer scientists have warned against model collapse -- the contamination of training sets with AI-generated outputs that progressively degrade model performance. Exemplifying a positive-feedback-driven failure, it produces effects such as word repetition or pixel noise, ultimately leading to a loss of meaning and coherence -- at least from an engineering standpoint. From a creative one, however, collapse is not merely a breakdown: it also functions as a recursive mirror that recalls early analog video feedback experiments, raising once again the question of what happens when a system turns inward and sees itself. In such cases, so-called machine vision no longer transmits the world (as in tele-vision) but increasingly generates worlds from within. Drawing on media archaeology through case studies of both historical video synthesis techniques and contemporary artistic uses of machine learning, this paper examines what recu

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems

arXiv:2607.09685v1 Announce Type: cross Abstract: Policy models often assume that the relationship between a policy instrument and its outcome remains stable across institutional conditions. In adaptive socio-technical systems this assumption may fail: regulatory change can alter incentives, agents can respond strategically, and the mapping from policy variables to aggregate outcomes can change. This paper studies such regime change as a transfer-learning problem in adaptive multi-agent systems. A policy regime is represented as a learning problem induced by an observable input distribution and a target function mapping policy variables to outcomes. We compare a blank-slate learner that searches a flexible hypothesis class in the new regime with a transfer learner whose effective hypothesis class is restricted by structural knowledge from the previous regime. Transfer is beneficial when this restriction preserves the new target function while reducing effective complexity; it is harmfu

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers?

arXiv:2607.11859v1 Announce Type: new Abstract: Can large language models perform deep technical comprehension of computer architecture papers -- not summarization, but structured critique that names the core mechanism, surfaces buried assumptions, and connects a contribution beyond its own scope? We study Gauntlet, an open-source pipeline that analyzes a paper through five independent expert-persona reviewers and an adversarial synthesis stage. On 20 ISCA 2025 and HPCA 2026 papers, ten researchers each wrote their own analyses and then judged, for papers other than their own, the human analysis against Gauntlet's. Across the 20 comparisons evaluators preferred Gauntlet in 15 (human in 4, one tie); its advantage is significant on per-analyst totals (paired Wilcoxon, p < 0.01) and largest on Critical Rigor, vanishing only on Calibration. Where humans win, it is on trust and usefulness rather than depth: a confident wrong claim, a mechanism described but not taught, or unprioritized brea

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Uncovering Students' Mental Models of Generative Artificial Intelligence

arXiv:2607.11692v1 Announce Type: new Abstract: In this paper we present a study of students' mental models of generative AI (GenAI). A student's mental model of GenAI influences not only how they perceive the technology's capabilities and limitations but also how they choose to integrate it into their academic work. Whether they view it as a collaborative partner, a shortcut to complete tasks, or something in between, depends on how they conceptualize its use. This study addresses the following questions: (I) What mental models do undergraduate students hold about GenAI? and (II) What aspects of conceptual knowledge - declarative, procedural, and conditional - are present in these mental models? Sixty-four concept maps were collected from students enrolled in a course on technology ethics. Students were asked to construct concept maps representing their understanding of GenAI use. The concept maps were analyzed using a structured codebook and the analysis revealed five categories of m

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

One Vote, Several Parliaments: An Empirical Analysis of the Algorithmic Ambiguity of the Italian Electoral Law on the 2022 General Election Data

arXiv:2607.11676v1 Announce Type: new Abstract: Crafa's algorithmic analysis of the Italian electoral law (D.P.R. 361/1957, as amended by Law 165/2017, the Rosatellum) showed that the statutory text distributing proportional seats among territories (Art. 83(1)(h)) admits at least three algorithmic interpretations, which can elect different people from the same votes. We test this empirically: we implement the full seat-allocation pipeline (Arts. 77, 83, 83-bis, 84, 85) and run the three interpretations on the complete open data of the Italian general election of 25 September 2022 (Chamber of Deputies). The implementation reproduces the official national apportionment exactly from raw municipal data, matches by name 389 of the 391 seats it models (99.5%), and agrees step by step with the official minutes of the National Central Electoral Office; the two residual mismatches fall on seats that the Chamber's Committee on Elections placed under formal investigation in 2025. On these validat

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Programming Language Policy as an AI Literacy Equity Problem: A 15-Nation Comparative Analysis

arXiv:2607.11314v1 Announce Type: new Abstract: The promise of AI literacy ``for all'' confronts a structural challenge embedded in how nations organise secondary computer science education. In most systems, a general-track subject -- Digital Literacy, ICT, TIC, or SNT -- bears the weight of universal AI literacy, while a specialist Informatics course serves STEM pathways separately. Yet the content and depth of the general track are shaped by governance decisions made largely with reference to the specialist one. This paper presents a comparative analysis of curricula and examination frameworks across fifteen countries, identifying two structural challenges. First, in several systems a significant portion of students completes secondary education without any formal programming exposure. Second, among those who do receive CS education, a \emph{Syntax Ceiling} emerges: Python-based instruction reaches most students, while the algorithmic depth associated with C++ remains concentrated in

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students

arXiv:2607.11292v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as conversational tutors, they risk institutionalizing systemic inequalities. This study presents a systematic API audit of four LLMs acting as history tutors, evaluating 1,800 responses regarding the 1989 Romanian Revolution across five student personas varying by ethnicity and socio-economic tier. We uncover four interconnected patterns of \emph{epistemic paternalism}: (1)~\textbf{Differential Refusal}, where safety-aligned models block 76.7\% of educational requests from low-tier students; (2)~\textbf{Epistemic Gatekeeping}, evidenced by a 3$\times$ reduction in access to geopolitical complexity (e.g., the contested ``coup theory'') for marginalized learners; (3)~\textbf{Agency Theft}, a lexical shift where models like LLaMA produce a 5$\times$ higher victimization-to-politics vocabulary ratio for Roma students compared to elite peers; and (4)~\textbf{Elite Hermeneutics}, where

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs

arXiv:2607.11228v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities. We introduce DeepBias, an adaptive framework for the in-depth probing of social biases in LVLMs with carefully designed agents. Our approach operates through a dynamic ''generation-evolution-probing'' loop. First, a generative ProposerAgent synthesizes test data and is iteratively updated via Direct Preference Optimization (DPO) based on the target LVLM's responses, exploring model-specific failure modes. Second, an autonomous skill-driven DiggerAgent rewrites each test data across multiple probing turns, adaptively selecting from a curated skill library of deepening and rew

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Adoption-Ready Project-Based Learning for Computing Education: The FORAP Framework and a Multi-Scale Project Portfolio

arXiv:2607.11129v1 Announce Type: new Abstract: This innovative practice full paper presents FORAP (Framework for Organizing Reusable and Adaptable PjBL Projects) and a portfolio of 14 adoption-ready project-based learning (PjBL) project packages built with the framework. PjBL in computing education offers strong educational benefits, yet its adoption remains limited by high instructor workload and recurring student technical challenges. FORAP addresses these barriers by organizing each package around a project designed with aligned learning objectives and described through project attributes, along with coordinated instructor, student, and assessment materials that support adoption and adaptation across diverse computing courses. We report on four years of deployment across 44 classroom trials at seven universities, drawing on feedback from students, instructors, and advisory board members. Results suggest that structured project packaging supports feasible adoption with limited modif

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

LLM-Generated Design Problems for Assessing Higher-Order Thinking in Project-Based Learning

arXiv:2607.11032v1 Announce Type: new Abstract: Project-based learning (PjBL) is common in computing education, but traditional assessments of PjBL often fail to capture higher-order thinking (HOT), especially in transfer contexts. This study introduces "design problems" (DPs): concise, scenario-based prompts that require applying project concepts in new situations, to address this gap. We examined instructor perceptions, the ability of large language models (LLMs) to generate DPs, and student experiences. Surveys of 31 instructors, evaluation of 80 LLM-generated DPs, and student performance data showed that while instructors value DPs, creation effort is a barrier. LLMs helped by producing high-quality prompts with strong expert agreement. Students rated DPs from different LLMs similarly, and their performance on DP tasks showed negligible correlation with traditional project grades, suggesting DPs may capture distinct aspects of HOT. Keystroke data also suggested deeper cognitive eng

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Effectiveness of Using Remote Laboratory in Promoting Simulation and Verification Tools

arXiv:2607.10900v1 Announce Type: new Abstract: The transition to remote learning during the pandemic has necessitated the development of new methods for conducting hands-on experiments. One significant challenge in this transition has been providing students with reliable and sustainable access to necessary hardware components, particularly for courses that require substantial equipment. Additionally, industry partners have high expectations for students to be proficient in simulation and verification tools. To address these challenges, we implemented a virtual breadboard feature that allows students to remotely access Field Programmable Gate Array (FPGA) hardware and complete lab assignments. Our evaluation of this approach, which included surveys of students and industry partners, revealed that it effectively transformed a traditionally in-person lab assignment into an online modality. Furthermore, this paper presents the perspective of industry professionals on verification and sim

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

FPGA Meets Breadboard: Integrating a Virtual Breadboard with Real FPGA Boards for Remote Access in Digital Design Courses

arXiv:2607.10895v1 Announce Type: new Abstract: With the transition to remote instruction, new modalities for conducting hands-on labs are needed. In particular, courses that entail major hardware components faced challenges in making the hardware available for students in a reliable and sustainable way. This paper presents our experience in implementing a virtual breadboard feature to interface with real Field Programmable Gate Arrays (FPGAs) boards located on the University of Washington's campus, where students can access the boards remotely to complete their lab assignments.

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Q-Learning Lab: Teaching Reinforcement Learning Through Learner-Generated Trace Analysis

arXiv:2607.10802v1 Announce Type: new Abstract: Reinforcement learning is usually introduced through the Bellman update, yet the equation often remains abstract to undergraduates: they watch policy arrows converge but rarely observe how each value is computed or why an action is chosen. We present Q-Learning Lab, a single-file, browser-based, bilingual (Thai/English) tool for teaching tabular Q-learning that requires no installation. Beyond the usual gridworld visualization - color-coded Q-values and policy arrows on a $5 \times 5$ world - the tool exposes a live Bellman-substitution panel showing the numeric update at every step, and logs each transition, including the full pre-action Q-row, the greedy-versus-random decision under $\varepsilon$-greedy exploration, and wall-collision events, into an exportable trace. The central contribution is a learn-export-analyze loop: learners run their own agent, export the complete trace as CSV, and analyze it themselves, producing learning curv

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Return of the solo author: The changing division of labor in science in the age of generative A

arXiv:2607.10780v1 Announce Type: new Abstract: Modern science has experienced a long shift from individual work to team production. Generative artificial intelligence (AI) might appear to extend this trajectory by lowering research costs and enabling larger-scale collaboration. Yet if tasks once performed by coauthors can be delegated to AI, the same technology may also weaken the need for collaboration in parts of the research process. Here, we examine this tension by moving beyond average team size and focusing on the solo-authored tail of the author-count distribution. Analyzing over 300 million works across 26 fields, we find that the decades-long decline in solo authorship halted and partially reversed with ChatGPT's public release in late 2022. We also reveal that this is an uneven phenomenon: it is strongest in fields where coauthors' work is more readily replaceable, and weak or absent in fields that depend on physical collaboration. At the individual level, the recovery is no

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Robo-Reporters: Evaluating Autonomous AI Agents as Algorithmic Gatekeepers in Computational Journalism

arXiv:2607.10736v1 Announce Type: new Abstract: Artificial intelligence agents increasingly perform journalism tasks autonomously, searching for sources, evaluating credibility, and producing news content with minimal human oversight. Yet research has largely treated AI as a monolithic category, leaving the effects of architectural design unexamined. Drawing on gatekeeping theory, this study presents the first systematic comparison of four agent architectures, monolithic (Claude), chain-based (LangChain), multi-agent collaborative (CrewAI), and autonomous iterative (AutoGPT), across 200 controlled experiments spanning 50 journalism tasks of graduated difficulty. All architectures used the same underlying language model and identical tools, isolating architectural effects. Results revealed significant effects on task duration (F(3, 196) = 24.54, p < .001, eta-squared = .27) and computational strategy (F(3, 196) = 305.63, p < .001, eta-squared = .82), with architecture explaining 82% of

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Sequential compliance decisions of firms on cross-border data flows: An institutionally anchored decision support system

arXiv:2607.10620v1 Announce Type: new Abstract: The economic value of data arises from its flow across organizations and national borders. Yet increasingly stringent data governance regimes are turning cross-border transfer into an institutionally constrained sequential decision, in which firms repeatedly weigh compliance costs against the value of data flows. From the perspective of a data-exporting firm, this paper develops an institutionally anchored decision support system. It converts regulatory rules into a computable minimal compliance mapping and models the firm's weekly decisions as a finite-horizon Markov decision process (MDP), with compliance represented as a hard constraint rather than a penalty term. The resulting problem is solved using masked deep reinforcement learning, while counterfactual path advantages provide interpretable signals to support the firm's cross-border data flow decisions. Experiments show that the policies learned within the system outperform the bas

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Rise of the Smart Compound: Privately Governed Urban Intelligence and Its Research Agenda

arXiv:2607.10469v1 Announce Type: new Abstract: Across a range of fast growing urban markets, private developers are constructing a version of the smart city that operates largely outside the purview of municipal government, often at the explicit invitation of city officials seeking to shift the cost and complexity of digital infrastructure onto private capital. Gated residential and mixed use developments are increasingly marketed not merely on the basis of security and amenity, but on their smartness: integrated home automation, app mediated access control, and centralized energy and resource management, among other features. We refer to this phenomenon as the smart compound. Despite its rapid proliferation, it has received comparatively little sustained scholarly attention: the literatures on smart cities and on gated communities have developed largely independently of one another, even as developers are, in practice, merging the two. This paper introduces the smart compound as an e

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Conceptual Architecture for Educational Digital Twins Supporting AI Literacy Across Educational and Professional Settings

arXiv:2607.10013v1 Announce Type: new Abstract: In the AI Literacy for Multidisciplinary Professional Readiness and Outreach (AIM-PRO) project, we are creating integrated methods to improve the education on AI literacy. One concept on which the project relies is educational digital twins, that is, digital representations of educator trainers, teachers, and learners that can be used in different stages of the educational process. Such digital twins enable the simulation, monitoring, and optimization of learning experiences. This paper presents the AIM-PRO project and its conceptual foundations, focusing on its core objective: designing and implementing Digital Twins for Education to foster AI literacy across higher education, vocational education and training and professional learning environments.

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Religion and Artificial Intelligence as Distributed Meaning Systems: A Naturalistic Conceptual Model

arXiv:2607.10011v1 Announce Type: new Abstract: This paper develops a naturalistic account of religion and artificial intelligence as structurally similar distributed meaning systems. I argue that both emerge from the same underlying cognitive architecture: socially extended processes that offload interpretation, norm-guidance, and world-model construction into external symbolic environments. Drawing on work in distributed cognition, cultural evolution, and philosophy of mind, the paper proposes a conceptual model showing how meaning is generated, stabilised, and transmitted through recursive interactions between agents and their informational ecologies. Religion is analysed not as a set of beliefs but as a cognitive-ecological system that scaffolds coordination, normativity, and shared interpretation. Contemporary AI systems are shown to instantiate analogous functions, operating as high-bandwidth, algorithmically mediated environments that shape reasoning, attention, and social meani

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

A realist theory of digital objects, digital systems, and digitalized systems

arXiv:2607.09834v1 Announce Type: new Abstract: As human reliance on information technology (IT) increases, having a clear, precise, and comprehensive understanding of the nature of digital objects, digital systems, and digitalized systems becomes more critical. Otherwise, our ability to research, manage, control, and make reliable predictions about them will be limited. Accordingly, in this paper, we propose a new theory of digital objects, digital systems, and digitalized systems based on the adoption, adaptation, and extension of existing theories of ontology, semantics, and semiotics. The theory provides precise explanations of the nature of digital objects, digital systems, and digitalized systems. Ours is a realist theory that does not countenance the independent existence of nonmaterial or hybrid objects in the world. Accordingly, we are at odds with much of the prevailing discourse about digital phenomena. We show how our theory generates different insights and predictions from

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Elite Proxies, Algorithmic Bottlenecks, and Multi-Platform Information Cascades in the 2026 Iran War

arXiv:2607.09676v1 Announce Type: new Abstract: As high-stakes executive crisis communication shifts into fragmented digital ecosystems, unmediated narratives increasingly originate within closed-broadcast enclaves before diffusing into structurally distinct open networks. Yet the mechanisms governing cross-platform information migration remain under-theorized, underscoring the need for systematic computational analysis. This study tracks the transformation of President Trump's statements during the 2026 Iran War as they move from Truth Social into X-Twitter and Bluesky. Using a Metric-to-Semantic-Linkage framework and a high-throughput dataset of 9,891 temporally aligned public responses, we examine how divergent platform architectures shape the downstream vernacular of geopolitical conflict. The results reveal pronounced behavioral divergence driven by sociotechnical affordances. On X-Twitter, discourse is dominated by structural information bottlenecks: viral retweet cascades from e

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Gender Gap Analysis in News and Talk Online Radio Broadcast

arXiv:2607.09675v1 Announce Type: new Abstract: Radio broadcasting remains a dominant medium of communication, reaching 82% of Americans ages 12 and older weekly. Given its broad media impact, gender representation on radio news and talk stations may play an important role in shaping social and cultural perceptions. In this study, we examined patterns of gender representation in radio broadcasts, focusing on gender-based differences in total speaking time, air-time allocation across the day, and participation across broadcast topics. The dataset comprises filtered recordings from 74 US news and talk radio stations, collected over a 24-hour period and yielding more than 1,400 hours of content. We analyzed the data using VANPY, an in-house voice-analysis framework that combines multi-channel radio recording with AI-based speaker diarization, gender classification, speech-to-text transcription, and topic analysis. The results revealed consistent gender differences in allocated broadcast t

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

From Tools to Teacher-Built Teammates: No-Code Pedagogical Plugin Authoring with LearnAdapt Agentic Studio and PedOS 1.1 Lumina

arXiv:2607.09674v1 Announce Type: new Abstract: Teachers and researchers need to adapt educational AI to local goals, but most systems remain difficult to customize or study without coding expertise. We present LearnAdapt Agentic Studio on PedOS 1.1 Lumina, a no-code authoring and governed runtime environment for educational AI plugins. A non-coder describes a desired learning interaction in plain English; the system prepares a previewable plugin artifact, runs safety checks, and supports submission for review. PedOS then deploys approved plugins into a directory for installation. Crucially, telemetry is strictly gated to authenticated users running approved plugins. The demo shows the complete lifecycle from prompt to governed evidence capture, shifting from fixed tools to teacher-built teammates.

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.CY

Critsly: An Artefact-Aware AI Critique Teammate for Design Education and Project-Based Learning

arXiv:2607.09673v1 Announce Type: new Abstract: Critique is central to design education and project-based learning, yet high-quality critique is often scarce, uneven, hard to document, and disconnected from evolving artefacts. We present Critsly, an artefact-aware AI critique workspace that turns AI from a detached feedback tool into a critique teammate in the learner's board context. Critsly combines a visual design canvas with structured AI-supported reflection, multi-perspective critique, action planning, optional peer/jury settings, and educator evidence traces. Unlike chatbot feedback tools that rely on isolated text prompts, Critsly grounds critique in a structured board state containing design intentions, board elements, annotations, links, and prior critique history. Reflecture, Critsly's guided reflection flow, works with Six Thinking Hats-inspired personas, board synthesis, generated action plans, exportable critique records, and educator evidence views in one workflow. The d

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Structuring license permissiveness from pairwise comparisons

arXiv:2606.31032v2 Announce Type: replace-cross Abstract: Licenses are legal instruments that inventors rely upon to protect the technologies they build and regulate how they are used---however, the nature of their authorship and selection implies that how they are interpreted, chosen, and enforced is largely unstructured. In practice, this makes it difficult to compare licenses at scale---when is one license considered more permissive than the other, and when are their terms incomparable to each other? Currently, there is a growing list of licenses that are introduced and used, yet no systematic way to study their relationships. This matters for platforms such as Hugging Face, GitHub, and the Python Package Index, where developers publish or build upon technologies that each have their own licenses. Using large language models (LLMs), we introduce methods for comparing licenses at scale: first, in a pairwise fashion to construct and validate a partial ordering based on permissiveness;

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

arXiv:2605.12530v2 Announce Type: replace-cross Abstract: LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally unreliable: surface-level prompt construction choices, although entirely orthogonal to the fairness question being tested, account for the majority of score variance, shift fairness conclusions in both the direction and the magnitude, and result in severe discordance in model rankings. We develop MAC-Fairness, a framework that embeds controlled variation factors into in-situ behavioral evaluation, examining how models' disparate-treatment behaviors shift when identity is varied as part of natural multi-agent conversation. Repurposing standardized-test questions as conversation seeds rather than as the evaluation instrument, we evaluate within-model differences in position persistence (how they hold positions, from the self-perspective) and peer receptive

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Culturally Situated AI Safety for Youth: Saudi Arabian Perspectives of Youth, Parents and Teachers

arXiv:2604.26494v2 Announce Type: replace-cross Abstract: Generative AI tools are widely used by youth and have introduced new privacy and safety challenges. While prior research has explored youths safety in GenAI within a Western context, it often overlooks the cultural, religious, and social dimensions of technology use that strongly shape youths digital experiences in countries like Saudi Arabia. To address this gap, this study explores youths (aged 7-17), parents and teachers interactions with GenAI tools and risk perceptions through a non Western lens. We analyzed 736 Reddit psots, and 1,262 X (Twitter) posts, and conducted interviews with 31 Saudi Arabian participants (8 youth, 13 parents, 10 teachers). Our findings highlight context dependent and relational privacy and safety needs of GenAI use from non-Western context, which are often shaped by communal structure and prescribed norms. We found significant risks tied to youths disclosure of personal and family information, whic

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

You can just review things: A digital ethnography of informal peer review

arXiv:2604.16764v2 Announce Type: replace-cross Abstract: Across scholarly communities, manuscripts face similar evaluative rituals: editors invite experts to privately assess submissions through formal peer reviews. This closed, loosely structured, and publisher-mediated process is now being supplemented by critiques on open, distributed platforms. We call this practice, a blend of three open peer review variants, informal peer review as it is accessible to outsiders, unmediated by publishers, and conducted across public platforms. Informal peer reviewers range from occasional error detectors to experienced sleuths who identify plagiarism, fraud, errors, conflicts of interest, and conceptual flaws. They may interpret methods, clarify jargon, assess value, and connect to related work. Here, we asked four questions: (1) Who are informal peer reviewers? (2) Where do they work? (3) How do they evaluate research? and (4) What are their impacts? To answer these questions, we conducted a cro

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Theory of Strategic Evolution: Games with Endogenous Players and the Seven Laws of Strategic Replicators

arXiv:2512.07901v4 Announce Type: replace-cross Abstract: Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. Rational players do not control their replication, and replicators do not choose strategically. Contemporary AI systems expose this gap: they optimize objectives, yet the population of AI systems is not fixed but expands and contracts based on performance. When capital can spawn capital, we need a theory that captures both rationality and replication. The Theory of Strategic Evolution analyzes strategic replicators: entities that optimize under resource constraints and spawn copies of themselves. The framework is organized around Seven Laws: 1. Strategic Selection: Mean fitness serves as a Lyapunov function; dominated types are eliminated. 2. ESDI Characterization: Equilibria exist, are generically finite, and satisfy Nash-KKT-LP equivalence. 3. H-$\gamma$ Stability: Multi-level systems are stable iff the spectral

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Multilingual Agent-Based World Modeling for Social Science

arXiv:2512.07195v2 Announce Type: replace-cross Abstract: Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly monolingual without cross-lingual interaction, an essential property of real societies. We introduce MAWM, the first Multilingual Agent-based World Modeling framework that supports multi-turn multilingual interactions among generative agents with diverse sociolinguistic profiles. MAWM enables two modes of analysis: (i) global public opinion modeling, which tracks how attitudes toward open-domain survey questions evolve across languages and cultures, and (ii) media influence and information diffusion, via autonomous news agents that dynamically generate content and shape user behavior. To ground the simulation in realistic population distributions, we construct the MAPS benchmark, which combines survey questions and demographic personas drawn from global population distributions. Experiments o

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

Has ACL Lost Its Crown? A Decade-Long Quantitative Analysis of Scale and Impact Across Leading AI Conferences

arXiv:2512.04448v3 Announce Type: replace-cross Abstract: The recent surge of language models (LMs) has rapidly expanded NLP/AI research, driving an exponential rise in submissions and acceptances at major conferences. Yet this growth has been shadowed by escalating concerns over conference quality, such as plagiarism, reviewer inexperience, and collusive bidding. However, existing studies rely largely on qualitative accounts, for example expert interviews and social media discussions, lacking longitudinal empirical evidence. To fill this gap, we conduct a ten-year empirical study (2014-2024) spanning seven leading conferences. We build a four-dimensional bibliometric framework covering conference scale, core citation statistics, impact dispersion, and cross-venue and journal influence. Notably, we further propose a metric called Quality-Quantity Elasticity (QQE), which measures the elasticity of citation growth relative to acceptance growth. We highlight two key findings. First, confe

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CY

flowengineR: A Modular and Extensible Framework for Fair and Reproducible Workflow Design in R

arXiv:2511.00079v2 Announce Type: replace-cross Abstract: flowengineR is an R package designed to provide a modular and extensible framework for building reproducible algorithmic workflows for general-purpose machine learning pipelines. It is motivated by the rapidly evolving field of algorithmic fairness, where new metrics, mitigation strategies, and methods continuously emerge. A central challenge in fairness, but also far beyond, is that existing toolkits either focus narrowly on single interventions or treat reproducibility and extensibility as secondary considerations rather than core design principles. flowengineR addresses this by introducing a unified architecture of standardized engines for data splitting, execution, preprocessing, training, inprocessing, postprocessing, evaluation, and reporting. Each engine encapsulates one methodological task yet communicates via a lightweight interface, ensuring workflows remain transparent, auditable, and easily extensible. Although imple

Source ↗
Showing 601–650 of 1593 signals
← Prev Page 13 of 32 Next →