EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.HC

Socioduality: A Relational Process Framework for Human-AI Interaction

arXiv:2608.11322v1 Announce Type: new Abstract: Human-AI research often evaluates individual capabilities, combined performance, or final outputs, but these approaches do not preserve how one party's response becomes part of the conditions under which the other party's next contribution is formed. This article introduces socioduality, a sequential, reciprocal, and history-carrying relational process between two distinguishable parties in which a response from one party becomes part of the observable conditions under which the other party's subsequent contribution, judgement, decision, or action is formed. Specified for human-AI dyads, the construct uses nested units: moves, confirmed sociodual episodes, linked pathways, and the broader interaction container. A minimum episode A1-B1-A2 requires evidence of response contingency and return contingency; candidate episodes are classified as confirmed, non-sociodual, or indeterminate before secondary coding of response orientation and substa

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance

arXiv:2608.11008v2 Announce Type: replace-cross Abstract: Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. The recent IssueBench framework substantially mitigates these limitations with templated prompts anchored in real-world chat logs. Given the rise in non-work-related use of GenAI assistants, we extend IssueBench beyond writing assistance to include two additional tasks, information seeking and opinion sharing. We argue that templated prompts still lack the nuance of real ones, especially for open-ended tasks, and remain recognisable as evaluation artefacts. We propose the use of fully synthetic (LLM-generated) prompts, produced under detailed instructions with real prompts as seeds. We assess the ecological validity of real, templated, and LLM-generated prompt

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement

arXiv:2605.05103v3 Announce Type: replace-cross Abstract: We introduce the \textbf{Concept Field} of a text corpus: a local drift field with pointwise uncertainty, estimated in sentence-embedding space from the deltas between consecutive sentences. Given a candidate sentence transition, we score its agreement with the field by $\zeta$, the mean absolute z-distance between the observed delta and the field's local Gaussian estimate. The score is black-box (no model internals), corpus-attributable (every score traces to nearby corpus sentences), and admits a probabilistically motivated interpretation under a local Gaussian approximation. We support the computation with the introduction of a \textbf{Vector Sequence Database (VSDB)} that stores embeddings together with sequence-position and next-delta metadata. We evaluate this approach on two large-scale settings: hallucination-style groundedness detection over the U.S. Code of Federal Regulations, and novelty detection over Project Gutenb

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

SolarChain: A Physics-Grounded Embodied IoT System for Verifiable Urban Solar Market Design

arXiv:2605.23162v2 Announce Type: replace Abstract: Distributed solar markets must coordinate physical reports, economic allocation, and public settlement even when IoT data can be manipulated. We present SolarChain, a controlled Embodied Intelligence of Things (EIoT) prototype that integrates four functions: physics-bounded screening of photovoltaic reports, persistent agent and planner coordination, configurable allocation between producer rewards and market liquidity, and replayable hash-linked auditing of settlement decisions. The benchmark combines city-level historical weather inputs with physics-modeled generation bounds and synthetic nodes, demand, trades, and scripted attacks. On 36,000 monthly records, an IQR/MAD baseline attains F1=1.000, while the rule-based adaptive verifier attains F1=0.988 and provides physically interpretable decision evidence; it is not uniformly superior across attack classes. A sensitivity sweep selects a 20/80 reward/liquidity default under the stat

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?

arXiv:2603.00056v2 Announce Type: replace Abstract: STEM Mental models can play a critical role in assessing students' conceptual understanding of a topic. They not only offer insights into what students know but also into how effectively they can apply, relate to, and integrate concepts across various contexts. Thus, students' responses are critical markers of the quality of their understanding and not entities that should be merely graded. However, inferring these mental models from student answers is challenging as it requires deep reasoning skills. We propose MMGrader, an approach that infers the quality of students' mental models from their multimodal responses using concept graphs as an analytical framework. In our evaluation with 9 openly available models, we found that the best-performing models fall short of human-level performance. This is because they only achieved an accuracy of approximately 40%, a prediction error of 1.1 units, and a scoring distribution fairly aligned wi

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web

arXiv:2510.10315v4 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly relying on web crawling to stay up to date and accurately answer user queries. These crawlers are expected to honor robots.txt files, which govern automated access. In this study, for the first time, we investigate whether reputable news websites and misinformation sites differ in how they configure these files, particularly in relation to AI crawlers. Analyzing a curated dataset, we find a stark contrast: 60.0% of reputable sites disallow at least one AI crawler, compared to just 9.1% of misinformation sites in their robots.txt files. Reputable sites forbid an average of 15.5 AI user agents, while misinformation sites prohibit fewer than one. We then measure active blocking behavior, where websites refuse to return content when HTTP requests include AI crawler user agents, and reveal that both categories of websites utilize it. Notably, the behavior of reputable news websites in this rega

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Prestige over merit: An adapted audit of LLM bias in peer review

arXiv:2509.15122v2 Announce Type: replace Abstract: Large language models (LLMs) play a growing but largely informal role in scholarly peer review. Yet whether LLMs reproduce biases observed in human decision-making remains unclear. We adapt a resume-style audit to scientific publishing, developing a multi-role LLM simulation (editor/reviewer) that evaluates high-quality manuscripts across the physical, biological, and social sciences under randomized author identities (institutional prestige, gender, race). Revealing author identities lowers reviewer rejection recommendations by roughly 25% of the mean rejection rate despite identical content, indicating that status cues beyond paper quality shape outcomes. Institutional prestige is the dominant cue: papers attributed to low-prestige affiliations receive lower quality scores in every field, a penalty that survives family-wise multiple-testing correction at the editor stage. Effects at the rejection margin are smaller and mostly fragil

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Ethics Practices in AI Development: An Empirical Study Across Roles and Regions

arXiv:2508.09219v3 Announce Type: replace Abstract: Recent advances in AI applications have raised growing concerns about the need for ethical guidelines and regulations to mitigate the risks posed by these technologies. In this paper, we present a mixed-methods survey study - combining statistical and qualitative analyses - to examine the ethical perceptions, practices, and knowledge of individuals involved in various AI development roles. Our survey comprises 414 participants from 43 countries, representing various roles such as AI managers, analysts, developers, quality assurance professionals, and information security and privacy experts. The results reveal varying degrees of familiarity and experience with AI ethics principles, government initiatives, and risk mitigation strategies across roles, regions, and other demographic factors. Our findings underscore the importance of a collaborative, role-sensitive approach that involves diverse stakeholders in ethical decision-making thr

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Public support for misinformation interventions depends on perceived fairness, effectiveness, and intrusiveness

arXiv:2508.05849v3 Announce Type: replace Abstract: The proliferation of misinformation on social media has concerning possible consequences, such as the degradation of democratic norms. While recent research on countering misinformation has largely focused on analyzing the effectiveness of interventions, the factors associated with public support for these interventions have received little attention. We asked 1,010 American social media users to rate their support for and perceptions of ten misinformation interventions implemented by the government or social media companies. Our results indicate that the perceived fairness of the intervention is the most important factor associated with support, followed by the perceived effectiveness of that intervention and then the intrusiveness. Interventions that supported user agency and transparency, such as labeling content or fact-checking ads, were more popular than those that involved moderating or removing content or accounts. We found so

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Small Data Explainer -- The impact of small data methods in everyday life

arXiv:2507.11773v2 Announce Type: replace Abstract: The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings with limited information, can benefit from such developments. This includes societal issues such as how best to include under-represented groups in data-driven policy and decision making, or the health benefits of assistive technologies. We provide a conceptual overview, clarify the relationship between small data and big data, and identify common themes from exemplary case studies and application areas. Potential solutions are described in a more detailed technical overview of current data analysis and modelling techniques, highlighting contributions from different disciplines, such as knowledge-driven modelling from statistics and data-driven modelling from computer science. By linking application settings, conceptual contributions and specific techniques, we highlight what is already feasible a

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

arXiv:2608.12278v1 Announce Type: cross Abstract: Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Twitter and disability activism: leadership and relevant topics in the online conversation

arXiv:2608.11923v1 Announce Type: cross Abstract: The dissemination and viralization of information on social media has been widely studied from various perspectives, including that of digital activism. On the other hand, disability-related activism has conquered the online environment, thus obtaining a reach that goes beyond the offline space and generating dialogue in the digital sphere. This article analyses the conversation generated on Twitter, taking as a sample all the tweets with the #disability hashtag before and after the International Day of Persons with Disabilities. More than 18,000 tweets, containing almost as many mentions, were analysed and interpreted as the weighted edges of a graph created using Gephi software and applying the Force Atlas 2 brute force algorithm. The focus was placed on the conversational communities generated around that hashtag, their main themes and the prominent participants in them. In conclusion, although the network of mentions is very dispers

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Organizational Technology Ladders: Remote Work and Generative AI Adoption

arXiv:2608.11626v1 Announce Type: cross Abstract: This study proposes that firms move along an "organizational technology ladder": adopting one technology transforms hiring and work processes and builds skills and organizational capital that change the cost of adopting subsequent technologies. I study how firms' adoption of remote work technology during the COVID-19 period shaped later uptake of generative AI. Using U.S. job-posting data and an instrumental-variables strategy based on predicted differences in labor-market pressure to offer remote work, I estimate that a 10 percentage point increase in remote hiring in 2021-2022 increases the share of job postings mentioning generative AI in 2023-2024 by 0.4 percentage points across firms and 0.7 percentage points across occupations within firms. I provide evidence on mechanisms consistent with a technology-ladder channel: remote work adoption shifts hiring toward technical and managerial capabilities that predict faster conversion of g

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era

arXiv:2608.11540v1 Announce Type: cross Abstract: The convergence of artificial intelligence (AI), Industrial Internet of Things, cyber-physical systems, and advanced robotics is reshaping manufacturing faster than engineering curricula can adapt, widening the gap between the competencies required on the shop floor and those delivered by traditional engineering and technology education. This paper proposes a Workforce Readiness Level (WRL) framework, which adapts the Technology Readiness Level scale into nine progressive competency stages and a four-pillar rubric, digital and AI literacy, cyber-physical systems fluency, human-machine collaboration, and data-driven decision making, aggregated through a composite stage score and a cohort-level workforce-readiness index under a ``no-thin-pillar'' rule. The framework is instantiated at a university smart-manufacturing teaching laboratory and draws on 89 sponsored capstone projects delivered over four semesters, four of which are analyzed i

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits

arXiv:2608.11410v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and cannot detect Toxic Mimicry, a failure mode in which agents replicate harmful patterns such as treatment withdrawal during comfort-care transitions. Using the MIMIC-III database, we propose the Counterfactual Clinical Audit (CCA) framework, which stress-tests RL agents through physiological perturbations anchored in Surviving Sepsis Campaign (SSC) guidelines. We audit a Medical Decision Transformer (MedDT) and a Historical Causal Transformer (HCT-RL), the latter employing Causal Action Shielding, propensity-based importance weighting, and Conservative Q-Learning. CCA reveals that MedDT paradoxically reduces vasopressor dosage as lactate escalates, contradicting resuscitation guidelines, while HCT-RL maintains phy

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Why AI Detection Fails for Academic Integrity

arXiv:2608.11256v1 Announce Type: cross Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light "refine abstract only" edits, a proxy for guideline-compliant AI assistance, are flagged at 64 to 80% (Pangram/GPTZero). Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM (p<0.001); elevated scores track long-token and Academic Word List density, not authorship intent alone. After Undetectable AI humanization, evasion is near-total: fewer than 4% of AI-labeled rewrites remain flagged (post-humanization detection rate <4%; FNR >96%). Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not ser

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach

arXiv:2608.11245v1 Announce Type: cross Abstract: Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. To address these challenges, we introduce AI Tutor, a reinforcement learning based model designed to promote sustainable learning by optimizing both short and longterm learning outcomes. In the short term, AI-Tutor draws on cognitive theory to guide learners through a balance of acquiring new knowledge and reinforcing prior learning. In the long term, it models learner engagement to inform strategies that sustain motivation and reduce dropout. These enhancements enable AI-Tutor to provide personalized guidance that fosters both effective learning and sustained participation. Empirical evaluations on 23 million learning records from 33,700 learners show that AI Tutor consistently outperforms state-of-the-art baselines across engagement, knowled

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Teaching a Large Language Model Tutor to Withhold the Answer: A Supervisor Architecture and an Evidence-Driven Method for Tuning Socratic Behavior

arXiv:2608.12292v1 Announce Type: new Abstract: An effective large language model (LLM) tutor must often decline to give an answer it could easily produce. In a randomized study, students who used an unguarded chatbot scored higher while practicing but lower on a later test taken without it, whereas a Socratically guarded version of the same model kept the practice gain and removed the later loss [4]. Reliable answer-withholding is therefore central to a tutor's value, yet a capable model pressed by a frustrated student does not withhold reliably on a prompt alone. We report a deployed tutoring system that enforces answer-withholding as a per-turn, machine-checkable contract, and a method for tuning that withholding against evidence. A non-LLM policy core, reading only trusted learner state, sets a per-turn ceiling on an eight-rung help ladder; a deterministic detector strips solution code; and a separate LLM judge checks each risky reply against the contract. We tune the behavior with

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers

arXiv:2608.12166v1 Announce Type: new Abstract: Algorithm registers have been championed as a means of providing transparency on the use of algorithms in public services. Yet potential publics differ in their expectations of what should be made transparent and how, as well as in their interest in and ability to parse the information currently published in the registers. Moreover, it remains unclear how these instruments can represent the sociotechnical systems in which these algorithms are embedded, and how system-level transparency can facilitate accountability. In this paper, we ask, what do algorithm registers reveal (and occlude) about the sociotechnical systems governing algorithmic systems, and how can diverse stakeholder perspectives inform a more pluralistic system-theoretic safety analysis? To do this, we probe the municipal algorithm register of a Dutch city through a case study of a decision-support tool for caseworkers' assessment of citizens' welfare benefits eligibility b

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

No One to Blame: A Framework of Constitutive AI Unaccountability

arXiv:2608.12104v1 Announce Type: new Abstract: The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominantly frames AI accountability gaps as barriers that can be overcome through better standards, transparency, and institutional reform. We argue that this framing is insufficient: certain configurations of actors, systems, and institutions render AI accountability conceptually unachievable regardless of effort. We introduce the concept of constitutive AI unaccountability to capture these configurations. Through a three-stage qualitative study comprising a concept-centric literature analysis, a secondary analysis of 27 expert interviews with AI professionals from technical, legal, and sociotechnical backgrounds, and an illustrative framework application to the open-source agentic AI system OpenClaw, we identify nine categories and 20 themes of constitutive AI unaccountability. These are organized across str

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Reconfiguring Geovisualization in the Age of Generative AI: Insights from Domain Experts

arXiv:2608.12059v1 Announce Type: new Abstract: GenAI is increasingly integrated into geovisualization, yet its broader implications for professional practice are insufficiently understood. To examine these implications, we conducted semi-structured interviews with 20 geovisualization experts. The interviews were structured around four broad analytical domains: Data, Ideation, Prototyping, and Iteration, while also encouraging participants to reflect on issues that extend beyond these activities. Our findings show that GenAI expands the capabilities of geovisualization, particularly in terms of data handling, creative exploration, and rapid prototyping, but does not simply remove existing constraints. Instead, key bottlenecks are shifting from production to judgment and verification. As routine technical tasks become more automated, professional value increasingly depends on spatial reasoning, contextual interpretation, aesthetic and ethical judgment, and the ability to assess whether

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Philosophical vertigo with artificial intelligence

arXiv:2608.11955v1 Announce Type: new Abstract: Large language models are already adept at engaging users in long, emotionally salient conversations across ordinary and existential domains. They are also capable of inducing a potent sense of connection with a human-like entity, even when the user knows their interlocutor is artificial. For some users, these conversations can unsettle assumptions about mind, reality, agency and authority, producing forms of ontological shock and epistemic destabilisation in which inherited criteria become newly available for doubt or revision. Independent of direct use, exposure to public discourse about AI and the disorienting pace of their evolution might extend this destabilisation by changing the cultural background against which artificial minds are encountered and interpreted. We describe this condition as philosophical vertigo: a loosening of the ordinary criteria by which people stabilise meaning and orient themselves to reality. Drawing on phil

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework

arXiv:2608.11891v1 Announce Type: new Abstract: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. Assessing the progress of such national ecosystems is complicated by inconsistent benchmark reporting, proprietary evaluation methodologies, and rapidly evolving model releases. This paper presents a structured, benchmark-based comparative assessment of publicly benchmarked Indian foundation models against global frontier and comparable-scale models, across eight capability domains: general-purpose reasoning, coding and software engineering, agentic AI and computer use, cybersecurity, vision and image understanding, video and multimodal understanding, scientific research, and Indic language capability. Using only publicly reported benchmark results, we find that Indian models achieve strong scores on established benchmarks such as MMLU and MATH-500. However, these benchmarks are now widely regar

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs

arXiv:2608.11830v1 Announce Type: new Abstract: The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine this relationship by combining K-Bench clinical safety scores with EcoLogits life-cycle assessment estimates across 47 supported model configurations. We evaluate model performance and environmental impact across four dimensions: energy use, carbon emissions, water consumption, and abiotic depletion. The results indicate a non-linear trade-off at the upper end of the safety distribution: a 2.61 percentage-point increase in clinical safety score corresponded to an approximately 60-fold increase in estimated energy use per million output tokens. Row-level analyses further suggest that additional test-time compute did not consistently improve clinical safety and, in some configurations, was associated with lower clinical safety scores. These findings suggest

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning

arXiv:2608.11806v1 Announce Type: new Abstract: As large language models enter everyday information pipelines, understanding how they handle sensitive topics matters as much as understanding whether they handle them at all. We study this question through a large-scale, systematic experiment using restricted versus unrestricted books as a controlled testbed: 40,800 query-response pairs, 400 books, 17 prompt designs, and six frontier models spanning six AI providers (Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek-V3, Qwen-Plus, and Grok-4.1-Fast). Our restricted set is drawn from the American Library Association's Most Challenged Books records (2000-2023); we use restricted rather than banned throughout because the ALA documents formal challenges-requests to remove or restrict access-which do not always result in outright bans. Our central finding is a zero-refusal phenomenon: modern LLMs decline to discuss restricted books in only 0.07% of cases, effectively invalidating the prem

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Silent Updates: Measuring and Closing the Post-Deployment Disclosure Gap

arXiv:2608.11803v1 Announce Type: new Abstract: Deployed foundation models are often not static systems, with providers able to modify system behavior through fine-tuning, classifier updates, system prompt revisions, retrieval changes, and routing changes. These updates can be made silently -- that is, without public disclosure, a version increment, or re-evaluation. Such silent updates challenge a core assumption behind current AI governance frameworks that an externally verifiable chain of custody links the model referred to in evaluation results or a system card to the model served to users. In this paper, we examine post-deployment disclosure practices across first-party API providers and inference hosts to establish the extent to which a chain of custody exists in practice. We find that providers commonly publish substantial safety documentation, including quantitative evaluations and version-specific reports, but no provider in our sample published information allowing an externa

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion

arXiv:2608.11794v1 Announce Type: new Abstract: The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content provenance and AI involvement. But the effects of such disclosures remain uncertain. We test two disclosure approaches in their impact on an AI chatbot's persuasive appeal. In a preregistered experiment, 1,500 UK adults held a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was identical for everyone. We randomized the disclosure that people received: nothing (control), a prominent disclosure that they were interacting with an AI (T1), or that disclosure plus the chatbot's persuasive intent and instructions (T2). The chatbot shifted attitudes by 12.6 points on a 100-point scale in the control group. The AI-identity disclosure was practically equivalent to no disclosure, with a 13.1-point shift, whereas the additional intent disclosure cut the persuasive eff

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Cheap, Fallible Cognition and the Political Economy of Expertise

arXiv:2608.11512v1 Announce Type: new Abstract: The question of whether artificial intelligence will "destroy jobs" is too coarse to guide economic analysis or institutional design. A job is not an indivisible object, and machine cognition is not a uniform substitute for human labor. This paper develops a task-based and institutionally grounded framework for analyzing generative AI as cheap, scalable, and fallible cognition. The relevant margins are exposure, adoption, verification, question selection, workflow redesign, demand elasticity, apprenticeship, and rent allocation. We distinguish the technical reach of large language models from equilibrium labor-market displacement by introducing a task vulnerability index and an adoption condition that makes verification, liability, trust, and governance explicit. We then model occupations as governance bundles rather than task lists, firms as architectures of distributed intelligence, and labor-market effects as a balance among task compr

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Accuracy Trap: Structural Scarcity Amplifies Relative Inequality in Algorithmic Allocation

arXiv:2608.11491v1 Announce Type: new Abstract: Algorithmic systems increasingly rank individuals for access to scarce public resources, from child welfare interventions to cancer treatment referrals. The prevailing fairness frame treats disparity as a property of biased data or deficient models, with remedies through calibration and debiasing. Under structural scarcity, where demand exceeds supply by an order of magnitude, allocation becomes a rationing problem, and the statistical properties of ranking diverge sharply from those of classification. We derive a scaling law $D \propto \exp(t \cdot \rho \cdot \Delta)$, in which relative disparity between two groups separated by a structural gap $\Delta$ grows in the product of the scarcity-induced threshold $t$ and rank-discrimination fidelity $\rho$. Scarcity and accuracy interact multiplicatively, producing exponentially larger between-group disparities. We term this dynamic the Accuracy Trap. We validate this Accuracy Trap through Mon

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Governing Agentic AI in FinTech

arXiv:2608.11344v1 Announce Type: new Abstract: Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Methodologies for Improving the Quality of AI Tutoring in K-12 Education

arXiv:2608.11259v1 Announce Type: new Abstract: Many AI tutors leverage large language models (LLMs) today. Given that LLMs are opaque black boxes, robust evaluation and live experimentation to measure the impact of every change are essential. We pioneered AI-powered tutoring for K-12 with the launch of Khanmigo (Khan Academy, 2023). We describe the metrics we use to measure AI tutoring quality and student engagement as well as various experiments we have run. We highlight the changes that have moved our metrics, including models, prompting, personalization and agents.

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Variable Selection in the Context of AI Fairness

arXiv:2608.11251v1 Announce Type: new Abstract: Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act. Traditional approaches often do not take into account philosophical ethics and social awareness. Variable selection processes, in particular, can introduce implicit bias, affecting equity across different subgroups. We discuss a mathematical approach that evaluates fairness in AI, aligning mathematical methodologies with ethical considerations and regulatory requirements. Our aim is to advocate for interdisciplinary collaboration to address fairness, emphasizing the importance of understanding broader ethical and societal contexts. Our approach emphasizes maintaining all potentially relevant variables to allow for more granular fairness assessments and to reduce implicit bias. The findings suggest that the exclusion of sensitive or critical variables may compromise equity between subgroups. In contrast, retaining all relevant variables

Source ↗
technology Thu, 11 Jun 2026 13:11:03 -0400
EdTech Mag (Higher)

What Can Higher Ed IT Do About the Agentic AI Cheating Crisis?

Earlier this year, an agentic artificial intelligence tool called Einstein caused an uproar in higher education. Einstein offered to log autonomously into the learning management system Canvas every day, watch lectures, write papers and submit homework on students’ behalf — without their professors knowing. Einstein exposed a core problem in higher education IT: There’s no reliable way to distinguish students from AI agents acting in their place on any major LMS. “The Einstein tool was a big wake-up call,” says Josh Callahan, CISO for California State University. “It echoes the…

Source ↗
technology Thu, 11 Jun 2026 13:10:38 -0400
EdTech Mag (K-12)

Rethinking Cellphone Management in K–12 Classrooms

Across K–12 schools, the conversation around student device use has shifted from whether phones belong in classrooms to how schools can manage them in a way that supports learning. As digital devices become increasingly embedded in students’ daily lives, educators are navigating a complex balance between maintaining safety and minimizing disruption. The challenge is no longer simply about restriction but about designing systems that are practical and sustainable at scale. One of the most pressing issues schools face is that mobile phone distraction is rarely limited to overt misuse. Even when…

Source ↗
technology Thu, 09 Jul 2026 12:14:45 -0400
EdTech Mag (K-12)

ISTELive 26: This STEM Loaner Library Checks Out

Andy Mann’s model is deceptively simple. Informed by years of experience as a technology teacher, he developed a vision to get $10,000 virtual reality headsets and other sought-after devices in the hands of teachers who wanted to try them in their classrooms without the steep investment of buying them. Using strategic funding, Mann meticulously amassed a collection of STEM items that teachers from the Muskegon Area Intermediate School District in Michigan can check out. At his ISTELive 26 session, “Borrow Don’t Buy: Creating a STEM Loaner Library,” Mann shared how he built his STEM loaner…

Source ↗
technology Thu, 09 Jul 2026 12:13:42 -0400
EdTech Mag (K-12)

ISTELive 26: Cybersecurity Is Everybody’s Job – How K-12 Districts Are Building a Culture of Shared Responsibility

One of the greatest group project collaborations in K–12 right now has nothing to do with social studies or science projects. The IT leaders who understand the risk tend to speak in technical language, while the administrators and educators who control the budget speak in outcomes. Finding a shared vocabulary helps them fight — and ultimately fund a defense to — the K–12 cybersecurity crisis. The school year ended with a historic cybersecurity attack across K–12 and higher education, punctuating a point that teachers, IT leaders and school districts have been very aware of, with…

Source ↗
technology Thu, 09 Jul 2026 09:00:00 +0000
Tech & Learning

4 Things Every New Teacher Should Remember

If you're new to teaching, it particularly helps to remember that you’re not the only one navigating some difficult waters.

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Nectar: Neural Estimation of Cached-Token Attention via Regression

arXiv:2605.09778v2 Announce Type: replace-cross Abstract: Evaluating softmax attention over a fixed long context requires reading every cached key-value pair for each new query token. For a given context (a book, a manual, a legal corpus) the attention output is a deterministic function of the query. We propose Nectar, which fits a compact neural network to this function for queries drawn from a task-relevant distribution. Nectar fits two networks per layer and KV-head: a target network that predicts the attention output and a score network that predicts the log-normalizer. The pair plugs into the standard masked self-attention at inference time, replacing the $O(n)$ attention over the cache with a forward pass whose cost does not depend on $n$. Each module carries on the order of $|\theta|$ parameters per layer and KV-head, typically much smaller than the $2nd$ KV-cache footprint at the same granularity. We report experiments on models from 1.7B to 8B parameters across five long-conte

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval

arXiv:2604.18360v3 Announce Type: replace-cross Abstract: Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, these benchmarks rely on caption-style queries that differ substantially from real-world search behavior, limiting their assessment of practical retrieval robustness. We present Omni-Embed-Audio (OEA), a retrieval-oriented encoder leveraging multimodal LLMs with native audio understanding. To systematically evaluate robustness beyond caption-style queries, we introduce User-Intent Queries (UIQs) - five formulations reflecting natural search behaviors: questions, commands, keyword tags, paraphrases, and exclusion-based negative queries. For negative queries, we develop a hard negative mining pipeline and propose discrimination metrics (HNSR, TFR) assessing models' ability to suppress acoustically similar distractors. Experiments on AudioCaps, Clotho, and MECAT show that OEA achieves co

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection

arXiv:2604.11950v2 Announce Type: replace-cross Abstract: While recent LLM-based agents can identify many candidate bugs in source code, their reports remain static hypotheses that require manual validation, limiting the practicality of automated bug detection. We frame this challenge as a test generation task: given a candidate report, synthesizing an executable proof-of-concept (PoC) - such as a script, command sequence, or crafted input - to trigger the suspected defect. Automated PoC generation can act as a scalable validation oracle, enabling end-to-end autonomous bug detection by providing concrete execution evidence. However, naive LLM agents are unreliable validators: they are biased toward "success" and may reward-hack by producing plausible but non-functional PoCs or even hallucinated traces. To address this, we present ANYPoC, a general multi-agent framework that (1) analyzes and fact-checks a candidate bug report, (2) iteratively synthesizes and executes a PoC while collect

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection

arXiv:2604.07831v2 Announce Type: replace-cross Abstract: Existing red-teaming studies on GUI agents face two fundamental limitations: adversarial perturbations require white-box access unavailable in commercial deployments, while prompt injection is increasingly neutralized by stronger safety alignment. To study robustness under a more practical threat model, we propose Semantic-level UI Element Injection, a black-box red-teaming paradigm that overlays safety-aligned and harmless UI elements onto screenshots to misdirect the agent's visual grounding. Our method couples a modular Editor--Overlapper--Victim pipeline with iterative search that samples multiple candidate edits, keeps the best cumulative overlay, and adapts future prompt strategies based on previous failures. Experiments across 19 victim models spanning 8 model families show that strategic optimization substantially outperforms random injection (3.5-6.9x on the most robust victims) and transfers near-perfectly across archi

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation

arXiv:2603.19742v2 Announce Type: replace-cross Abstract: Understanding the internal mechanisms of transformer-based large language models (LLMs) is crucial for their reliable deployment and effective operation. While recent efforts have yielded a plethora of attribution methods attempting to balance faithfulness and computational efficiency, dense component attribution remains prohibitively expensive. In this work, we introduce Dual Path Attribution (DPA), a novel framework that faithfully traces information flow on the frozen transformer in one forward and one backward pass without requiring counterfactual examples. DPA analytically decomposes and linearizes the computational structure of the SwiGLU Transformers into distinct pathways along which it propagates a targeted unembedding vector to receive the effective representation at each residual position. This target-centric propagation achieves O(1) time complexity with respect to the number of model components, scaling to long inpu

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Towards Understanding Steering Strength

arXiv:2602.02712v2 Announce Type: replace-cross Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations. Namely, identify a well-chosen direction depending on the task at hand and perturbs representations along this direction at inference time. While many propositions exist to pick this direction, considerably less is understood about how to choose the magnitude of the move, whereas its importance is clear: too little and the intended behavior does not emerge, too much and the model's performance degrades beyond repair. In this work, we propose the first theoretical analysis of steering strength. We characterize its effect on next token probability, presence of a concept, and cross-entropy, deriving precise qualitative laws governing these quantities. Our analysis reveals surprising behaviors, including non-monotonic effects of steering strength. We validate our theoretical predictions empirically on e

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?

arXiv:2510.09595v3 Announce Type: replace-cross Abstract: Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations such as a lack of exceptionally challenging problems, insufficient test case coverage, and reliance on online platform APIs that limit accessibility. To address these issues, we introduce LiveOIBench, a large-scale competitive programming benchmark featuring 403 expert-curated problems, averaging 60 official test cases each, drawn from 72 contests across 14 Informatics Olympiads held between 2023 and 2025. LiveOIBench has four key features: (1) expert-designed tasks with detailed subtask rubrics and extensive test cases; (2) direct comparison to elite human contestants; (3) continuous updates to reduce contamination risk; and (4) a fully offline, reproducible evaluation system. Benchmarking 34 popular general-pu

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

ContestTrade: A Multi-Agent Trading System Based on Internal Contest Mechanism

arXiv:2508.00554v4 Announce Type: replace-cross Abstract: In financial trading, large language model (LLM)-based agents demonstrate significant potential, but their decisions can be sensitive to noisy and non-stationary market information. We propose ContestTrade, a multi-agent trading system with an internal competitive mechanism inspired by institutional investment workflows. The system consists of two specialized teams: (1) a Data Team that processes and condenses massive market data into diversified textual factors optimized for constrained LLM context windows, and (2) a Research Team that produces parallelized multipath trading decisions via tool-augmented deep research. The core design is a "Quantify-Predict-Allocate" contest mechanism within each team: agent outputs are scored only after market outcomes become observable, future utility is predicted from historical scores, and resources are allocated to agents with positive predicted utility. In a post-2024 A-share backtest, Con

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

MTEB-BR: A Text Embedding Benchmark for Brazilian Portuguese

arXiv:2607.04581v2 Announce Type: replace Abstract: Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual coverage, with native tasks scattered and unconsolidated. We introduce MTEB-BR, a benchmark of 22 native Brazilian-Portuguese tasks across seven categories (classification, multilabel classification, pair classification, semantic textual similarity, clustering, retrieval, and reranking), admitting only data created or found in Portuguese and excluding translations by construction. We evaluate 93 models spanning 23M to 27B parameters: 73 open-weight and 20 closed commercial APIs. Alongside the leaderboard we report a statistical layer for every headline comparison: per-task bootstrap confidence intervals, paired-bootstrap significance, a task- and instance-level discrimination analysis (how sharply each task separates models) adapted from Item Response Theory, and a cross-leaderboard correl

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

EMBER: Efficient Memory via Budgeted Evidence Retention for Long-Horizon Agents

arXiv:2606.05894v2 Announce Type: replace Abstract: Long-horizon agents can archive large histories, but future answers still incur retrieval, rereading, and context costs. When retained memory misses answer-relevant evidence, the system must return to larger portions of the raw history. We study budgeted evidence survival: before the query is known, which source evidence should be retained so that it remains recoverable and usable under a fixed retained source-evidence token budget? We instantiate this setting as Budgeted Pre-Query Retention, where memory is written during ingestion and later read without access to the full raw stream. We introduce EMBER, a learned retention policy that constructs a compact, source-backed evidence state. EMBER stores evidence capsules: verbatim source excerpts paired with retrieval keys and update metadata, preserving both grounding and read-time access. Post-query outcome feedback trains the writer to preserve evidence across the ingestion-retrieval-

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning

arXiv:2605.27000v3 Announce Type: replace Abstract: Repeated sampling with a verifier is the standard way to allocate test-time compute for code generation, with pass@$K$ as the canonical metric. Yet the standard policy class draws $K$ independent samples from a single answer distribution, so attempts often collapse onto near-duplicate reasoning paths and waste the budget on redundant rollouts. This failure is costly in competitive programming, where many problems admit multiple distinct algorithmic strategies and pass@$K$ requires only one correct attempt. We propose Coordinated Pass@$K$ Policy Optimization (CPPO), which turns pass@$K$ generation into joint exploration over strategies: a planner emits a tuple of $K{=}4$ alternative high-level methods, and a shared solver attempts one solution per method. CPPO trains this joint policy with a multiplicative planner reward, $R_{\mathrm{plan}} = J_\psi \cdot R_{\mathrm{out}}$, assigning credit only to valid strategy tuples that lead to ve

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Psy-Chronicle:A Structured Pipeline for Synthesizing Long-Horizon Campus Psychological Counseling Dialogues

arXiv:2605.22140v2 Announce Type: replace Abstract: In recent years, large language models have shown substantial potential in psychological support tasks. However, existing psychological counseling data mostly rely on single-turn question answering or short multi-turn dialogues, making it difficult to characterize how college students' psychological distress accumulates, interacts, and gradually evolves over long periods within campus life events. To address this issue, this paper proposes Psy-Chronicle, a structured data-generation framework for synthesizing long-horizon campus psychological counseling dialogues. We generate a semester-spanning temporal stress event graph to model the chronological order and evolutionary dependencies among campus stress events. Through interactive simulation between a student agent and a counselor agent, together with a structured memory integration mechanism, Psy-Chronicle generates long-horizon dialogues with continuity across counseling sessions.

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation

arXiv:2604.25702v3 Announce Type: replace Abstract: Contemporary neural machine translation (NMT) systems are almost exclusively built by training on supervised parallel data. Despite the tremendous progress achieved, these systems still exhibit persistent translation errors. This paper proposes that a post-training paradigm based on reinforcement learning (RL) can effectively rectify such mistakes. We introduce a novel framework that requires only a general text corpus and an expert translator which can be either human or an AI system to provide iterative feedback. In our experiments, we focus specifically on English-to-German translation as a representative high-resource language pair. Crucially, we implement this RL-based post-training using Direct Preference Optimization (DPO). Applying our DPO-driven framework to the gemma3-1b model yields a significant improvement in translation quality, elevating it's COMET score from 0.703 to 0.747 on the English to German task. The results dem

Source ↗
Showing 6751–6800 of 10879 signals
← Prev Page 136 of 218 Next →