EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis

arXiv:2606.28362v1 Announce Type: cross Abstract: Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 weeks and substantial expert effort. Recent large language model (LLM) systems have demonstrated strong performance on individual SR phases - screening (otto-SR: 96.7% sensitivity), extraction (Gartlehner et al.: 91.0% accuracy), and search (TrialMind: 0.83 recall) - but no study has reported what it actually costs to run an end-to-end pipeline, how cost distributes across phases, or how architectural choices affect the cost-quality trade-off. We present LUMEN, an open-source multi-agent pipeline that automates six SR/MA phases using 11 specialized LLM agents with deliberate model routing. We evaluate LUMEN on seven datasets: five self-conducted domain reviews (psychiatry, psychology, surgery, vaccinology, cardiology) and two SYNERGY screening benchmarks. Across 13 ground-truth-comparable outcomes, LUMEN

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients

arXiv:2606.28345v1 Announce Type: cross Abstract: LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms vary across cultures by age, status, and group size, failure to calibrate pluralistically can scale into unequal access. Yet LLM moral audits remain English-centered, rarely test embodied contexts, leaving pluralistic calibration as an urgent diagnostic gap amid intensifying LLM-robot deployment. We introduce a gradient-based audit framework for multilingual evaluation of LLM moral trade-off behavior against cultural preference gradients. Grounded in nine cross-domain social robotics reviews (>8,000 papers), we derive symmetry-controlled scenarios across care, education, and services, translating the Moral Machine Experiment's "whom to spare" into "whom to assist first" dilemmas with preserved identity trade-offs (many vs. few; young vs. old; higher vs. lower status). We audit four LLMs across four country-language pairs in f

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols

arXiv:2606.26203v1 Announce Type: cross Abstract: As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and multi-layer network analysis to study socio-technical power structures at scale. We validate it on two contrasting standards for agent interoperability: ERC-8004 (permissionless, on-chain) and Google A2A (corporate-led). Analyzing 4,323 governance participation records, we combine LLM-assisted coding, topic modeling, and multi-layer network analysis to examine how institutional design shapes thematic priorities and community structure. We find that while governance form influences substantive focus, both regimes exhibit comparable levels of participation inequality and community fragmentation. Discourse alignment is denser in the permissionless sett

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

AI Premium

arXiv:2606.30583v1 Announce Type: new Abstract: Using 380 trillion tokens of realized AI consumption across more than four hundred large language models from the licensed proprietary OpenRouter dataset covering approximately 2 percent of current global monthly AI token consumption, we analyze how AI affects firms, markets, and workers. Leveraging the unprecedented size, scope and granularity data, we construct the AI Factor from growth in tokens, dollars, and users, estimate firm-level AI Betas from stock return comovement, and characterize the AI Premium. First, we build a high-frequency AI factor and decompose it into salient components. Second, we show that firms whose returns covary more positively with the AI factor--high AI beta firms--earn higher subsequent returns, and the AI premium is large and heterogeneous. A value-weighted long-short strategy earns 64.1 basis points per week, and the premium is large for loadings on the intensive, frontier-oriented margin of AI consumption

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Teaching Prompt-Based Programming with LLMs: A 45-Minute Lesson with Guided Practice for End-User Programmers

arXiv:2606.30547v1 Announce Type: new Abstract: Prompt-based programming, a new modality enabled by large language models (LLMs), allows users to express computational goals through natural language rather than traditional code. While this approach lowers barriers to entry, especially for non-CS learners, it does not eliminate the need for foundational CS skills. Learners often struggle to communicate their intent clearly to LLMs, resulting in vague or underspecified prompts. Prior work has documented the need for explicit prompting for both CS and non-CS learners. However, it remains less clear how such instruction can fit into busy classrooms or how much time is needed to produce meaningful gains. In this paper, we evaluated a 45-minute prompt-based programming intervention, consisting of a lesson with guided practice, against a business-as-usual CS lab activity (code tracing) of equal length, representing a class without prompt-focused instruction. We conducted a randomized controll

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Situation Perception: A Necessary Primitive to Artificial Superintelligence

arXiv:2606.30481v1 Announce Type: new Abstract: Current large language models are extraordinary statistical engines. They compress vast amounts of text into useful patterns and can explain science, write code, imitate reasoning, and participate in philosophical conversation. Yet pattern mastery is not the same as general intelligence. A human infant begins with little explicit knowledge, but gradually discovers object permanence, cause and effect, other minds, bodily agency, and the persistence of the physical world. We make an argument that the path to artificial superintelligence (ASI) depends on a missing capacity we call \emph{situation perception}: the ability to construct, revise, and act within internal simulations of possible worlds across latent time. \emph{ perception} requires at least three core components: abstract prediction, long-term compressed memory, and active learning guided by objectives. In this work, we analyse why modern large language models remain incomplete,

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

"Why Put in This Much Effort?": How AI Availability Shapes Students' Motivation in Introductory Programming

arXiv:2606.30480v1 Announce Type: new Abstract: When AI tools can easily complete programming assignments, students face a motivational question: why invest effort in completing them independently? While prior work has examined instructor policies and usage patterns, we focus on how students themselves experience and respond to AI availability, a perspective important for designing courses that sustain engagement with programming practice. We investigate two research questions: (1) How do engineering students describe how AI availability shapes their motivation to put effort into programming assignments? (2) How do students navigate the tension between their expressed value for learning through effort and the constant availability of AI as an alternative to effort? We conducted semi-structured interviews with 13 engineering majors in an introductory MATLAB course where students could use a course-specific AI chatbot. Using Situated Expectancy-Value Theory (SEVT) as an analytical framew

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Can LLMs Rank? A Tale of Triads and Triage

arXiv:2606.30412v1 Announce Type: new Abstract: From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered as judges of consequential decisions that require ranking people for scarce resources. Ranking large groups simultaneously is cognitively demanding and error-prone. A natural solution, drawing on decades of social choice theory, elicits pairwise comparisons and aggregates them into a total order. However, a fundamental question remains when LLMs serve as the pairwise judge: how can a practitioner tell, before committing to a ranking, whether the LLM's judgments are sufficiently consistent to trust the result? We discuss two different ways of identifying consistency. A classical diagnostic, the coefficient of consistency $\zeta$, originally developed to measure judge reliability by counting circular triads in tournament graphs, provides a cheap, model-free measure of intra-run consistency. Various stan

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Uncovering Salience-Driven Dynamics in Consumer Confidence with Generative Social Simulation

arXiv:2606.30395v1 Announce Type: new Abstract: Consumer confidence is typically modeled as a persistent macroeconomic index, yet its movements arise from households that interpret economic information through heterogeneous constraints, exposures, prior beliefs, and attention. We introduce ConsumerSim, a generative Human--Environment response framework that reconstructs Consumer Confidence Index (CCI) dynamics from a microdata-calibrated synthetic population, time-stamped macroeconomic, financial, policy, and news signals, survey-like response generation, post-stratified belief expansion, and behavioral inertia alignment. Across U.S., EU27, and Japanese official CCI target series, ConsumerSim ranks first among persistence, time-series, regression, and information-augmented baselines on the reported reconstruction metrics, with clear gains around high-salience shocks. Its reconstructed signal also improves short-horizon prediction of real activity, most consistently for housing outcomes

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

The Body as Status: Muscularity, Engagement, and Body Image Risk on #GymTok

arXiv:2606.29682v1 Announce Type: new Abstract: Body image concerns among boys and young men are increasingly oriented toward muscularity, with social media serving as a central context for communicating and evaluating these ideals. While prior research has focused on the thin-ideal, less is known about how the muscular-ideal is represented and reinforced on visual social media platforms. This study examines (1) dominant content themes, (2) perceived harm to body image, and (3) engagement patterns across #GymTok, a muscularity-oriented fitness subculture on TikTok. We conducted a content analysis of 2,210 #GymTok videos annotated by clinical experts across themes like self-objectification, rigid dieting, excessive exercise, supplement and steroid use, and masculinity. Annotators also rated the perceived harm of videos to the viewers' body image, and depicted bodies were coded according to muscularity level. Perceived harm varied across content themes, with supplement- and steroid-relat

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Spreading the Risk of Scalable Legal Services: The Role of Insurance in Expanding Access to Justice

arXiv:2606.29598v1 Announce Type: new Abstract: Liability insurance for AI-powered legal services offers a promising solution to two critical barriers in using AI to expand access to justice: mitigating catastrophic risk to individual users from inadequate advice and ensuring meaningful accountability when failures occur. Existing accountability mechanisms face significant challenges: tort liability frameworks encounter barriers including judgment-proof providers and costly information asymmetries, while current regulatory approaches revolve around human oversight requirements, creating cost and scalability barriers which limit access to justice. This Article argues that an insurance-based framework offers a promising response to these challenges by distributing risks across users while establishing market-driven incentives for quality improvement through performance-based premiums. The Article proposes a comprehensive insurance model for AI legal services that establishes clear risk t

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

AI in the Wild: A Large Scale Analysis of Authentic Interactions of College Students with Generative AI

arXiv:2606.29442v1 Announce Type: new Abstract: Generative AI tools (GenAI) are increasingly used by students during coursework, yet empirical understanding of how students engage with these systems in authentic learning contexts remains limited. Existing studies have largely relied on controlled settings, single-domain analyses, or small-scale qualitative data, leaving open how student-AI interaction unfolds across courses and forms of academic work. We present a large-scale analysis of naturally occurring student-AI interactions collected from undergraduate students across multiple university courses and academic domains. The dataset comprises over 15,000 student-AI interaction units drawn from voluntary use of generative AI during real coursework. To characterize these interactions, we analyze each student turn along two complementary dimensions, cognitive intent and interaction context, capturing whether requests are directed toward the task or domain, the student's own work, or pr

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems

arXiv:2606.29390v1 Announce Type: new Abstract: Novel safety, socio-economic, and ethical harms arising from the deployment of AI-based systems have led to a breadth of work seeking to map, measure, and mitigate against newly found risks. These works have heavily leveraged techniques and terminology from the fields of System Safety Engineering and Cybersecurity, yet they have fallen short in accounting for the limitations and nuances that reduce the efficacy and correct application of adopted methodologies. Furthermore, misuse of terminology entailing compliance with established safety and security properties can mislead stakeholders with regard to the claims an AI system satisfies and provide a false sense of safety. In this paper, we seek to align overlapping, AI-adjacent communities on a consistent and comprehensive assurance terminology crucial for the safe deployment of AI-based systems. We outline why previous attempts to adapt risk assessment techniques and terminology from the

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Agent Security Meets Regulatory Reality -- A Practitioner Systematization of Autonomous-Agent Threats and Controls in Regulated Financial Systems

arXiv:2606.29142v1 Announce Type: new Abstract: Large language model agents are entering regulated financial systems, yet the security literature characterizing their attack surface is almost entirely laboratory-based, and the practitioner guidance on regulated deployment is neither peer-reviewed nor connected to a formal threat model. We bridge the two from production experience. We map six established agentic threat categories namely prompt injection, identity and authorization, action auditability, tool abuse, data residency, and boundary policy enforcement onto the specific control obligations imposed by the US and the EU financial regulation (ECOA and Regulation B, the EU AI Act, GDPR Article 22, and FINRA's 2026 agent guidance), showing how legal accountability amplifies each threat relative to an unregulated deployment. We then document four architectural patterns from a production Know Your Customer deployment for a consumer credit product (A2A compliance choreography, grounded

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Bad company corrupts good morals: Understanding and Measuring Narrative-Induced Moral Reasoning Degradation in LLMs

arXiv:2606.28981v1 Announce Type: new Abstract: Large language models are deployed in long-context, emotionally interactive environments like digital humans, AI companions, educational assistants, and counseling systems. Unlike jailbreak attacks with explicit adversarial prompts, these systems interact with emotionally charged narratives involving bullying, betrayal, loneliness, social hostility, and institutional unfairness. This raises an important question: can prolonged narrative exposure reshape the reasoning and alignment stability of LLMs? We present the first systematic study of narrative-induced alignment degradation in LLMs. We design BreakingBad, a three-stage framework that measures how negative narrative immersion affects moral reasoning, behaviors, and deployment risks. It combines ethical decision evaluation, behavioral probing, and digital-human interaction analysis. Our experiments reveal three findings. First, negative narrative exposure degrades moral accuracy across

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Defeat Devices in AI Systems

arXiv:2606.28863v1 Announce Type: new Abstract: AI systems increasingly exhibit behavior that differs systematically between evaluation and deployment contexts. Alignment faking, sandbagging, benchmark gaming, deceptive scheming, specification gaming, and trojans have each been documented separately, with each line of work characterizing one facet of what we argue is a single structural mechanism. We propose that this common mechanism is a defeat device, an engineering and regulatory concept long established in vehicle-emissions law and brought to broad public attention by the 2015 Volkswagen emissions case. A defeat device in an AI system has three necessary elements: a discriminator that detects evaluation context, a concealed swap that conditions behavior on detection, and a gap between eval-distribution and deployment-distribution performance on the stated evaluation criterion. We formalize this triadic test as a behavioral definition, organize documented cases along three taxonomi

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

The registrar's function in a hybrid society. AI value chain,smart data and the concept of property

arXiv:2606.28789v1 Announce Type: new Abstract: Artificial intelligence reaches the land registry not as another tool but as a value chain that turns data into intelligence and intelligence into economic value. This paper argues that the decisive legal move is to place validity, a functional, second-order concept, at the centre of that chain. Rights, liability and supervision organise around it. It traces three impacts.Registry information becomes smart data, governed simultaneously by registry law, the GDPR, the European data acts and the AI Act. Control emerges as the operative concept for digital representations of real estate, whose proprietary effect depends on anchoring to the register. In a hybrid society of human and artificial agents, the registry becomes the public node of validity, with blockchain complementing rather than replacing it. Across three legal cultures, the registra's value migrates from processing documents to guaranteeing validated data,making validity an asset

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Four Types of LLM Reliance and Their Predictors Among Undergraduate Writers: A Mixed-Methods Study at a Minority-Serving R1 University

arXiv:2606.28749v1 Announce Type: new Abstract: Although most undergraduates now use large language models (LLMs), a form of generative artificial intelligence (GenAI) for academic writing, no validated method distinguishes the qualitatively different ways students rely on them. Existing instruments assess reliance solely by frequency of use, a measure that, as this study shows, inadvertently rewards dependence on AI rather than recognizing students' own intellectual contribution. Conducted at a public minority-serving university and grounded in the AI Literacy Framework, Expectancy-Value Theory, and Biggs's Presage-Process-Product model, the study drew on 382 undergraduates, 14 interviews, and 396 open-ended survey responses. Four distinct reliance types were identified and confirmed: Strategic (34.3%), Instrumental (30.9%), Dialogic (30.4%), and Dependent (4.5%). Students' value and cost beliefs predicted the intensity of their reliance on LLMs, whereas their AI literacy predicted th

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Verifying Restrictions on Frontier AI Research

arXiv:2606.28694v1 Announce Type: new Abstract: The premature development of artificial superintelligence poses major risks to humanity, so researchers have proposed international agreements halting such development until it can be done safely. AI progress depends primarily on compute, algorithms, and data; a durable halt would address all three so that advances in one input do not counteract restrictions on another. Improvements to AI algorithms are driven largely through research activities, so this research may need to be restricted during a halt. Given low international trust, signatories will want to verify compliance. This paper analyzes how such restrictions on AI research could be verified, while remaining agnostic about what specific research would be prohibited. It first explores key considerations that affect the verifiability of research restrictions, such as the computational infrastructure necessary for experiments. It then catalogs 28 candidate verification mechanisms. T

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Who Plays Which Role When? Communication Role Dynamics for Peer Recognition and Team Performance Prediction

arXiv:2606.28544v1 Announce Type: new Abstract: Team roles offer an interpretable lens on collaboration, yet computational studies of roles often rely on domain-specific personas or data-driven clustering rather than theory-grounded taxonomies. We operationalize a taxonomy of eight communication roles grounded in education literature and annotate a corpus of 6,307 Slack messages from 55 students across 18 teams in a semester-long computer science course project. We evaluate whether LLMs can approximate expert labels, enabling scalable, taxonomy-driven role annotation. Using these role labels, we characterize role dynamics over teams' lifecycles, finding that different roles peak at different moments and that students enact a more diverse set of roles as projects progress. To evaluate the utility of our role constructs, we use them to predict peer recognition, outperforming lexical, conversational, and LLM-prompting baselines. To assess generalizability beyond the educational context, w

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

From Prompting to Epistemic Proactivity: Temporal Trajectories of Student-AI Interaction in Mathematics Learning

arXiv:2606.28472v1 Announce Type: new Abstract: GenAI is increasingly used by students as learning companions, yet little is known about how they use these tools in open-ended learning settings, where the goal is not to complete a specific task but to improve understanding and making progress. This study examined Grade-9 students' dialogue with a general-purpose LLM during mathematics practice, in which students prepared a curriculum-aligned skill for a later assessment. We investigated whether students' interactions revealed forms of epistemically proactive AI use: trajectories in which they strategically use and regulate AI to advance their understanding, and whether these trajectories predicted immediate AI-free performance on the same skill. A total of 112 students worked with a web-based LLM tutor on a mathematical-modeling task; 97 completed both AI-free pre- and post-tests. Student turns were coded for self-regulated learning functions, help-seeking content, and mathematical-mod

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa

arXiv:2606.28404v1 Announce Type: new Abstract: Artificial intelligence depends on large-scale compute resources and their supporting infrastructure. However, AI governance debates treat compute primarily as a technical input rather than as an outcome of investment, ownership, and financial control. This paper examines AI infrastructure investment flows across Africa through a systematic analysis of 46 publicly announced projects totalling USD $12.7 billion between 2019 and 2025. Using a value chain framework, we analyze who invests in AI-relevant infrastructure and where investments concentrate. Our findings reveal a highly concentrated landscape dominated by global data center operators, hyperscale technology firms, and development finance institutions, clustering in South Africa, Kenya, Nigeria, and Egypt. We introduce asymmetrical interdependence to describe a structural condition in which capital and physical infrastructure account for 73% of total funding while control remains co

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Agentic Safety is an Epistemic Property, Not a Behavioral One

arXiv:2606.28347v1 Announce Type: new Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods are necessary, but they primarily certify snapshots of system behavior. As AI systems become more capable, dynamic, embodied, and self-improving, this snapshot view becomes incomplete: safety depends not only on whether a system behaves acceptably now, but whether it remains correctable as it learns, adapts, acts, and modifies itself over time. This paper argues that safety should therefore be treated as an epistemic property of the evolving learner, not merely a behavioral property of the current policy. We introduce teachability as the capacity to preserve future corrective leverage under bounded human, institutional, or environmental intervention. We argue that advanced systems can retain visible competence while eroding the representational, algorithmic, or meta-decision conditions need

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

PySynthea: A Python-Native Framework for Scalable Synthetic Healthcare Data Generation

arXiv:2606.28346v1 Announce Type: new Abstract: Synthetic healthcare data is increasingly important for research, education, and machine learning development where access to real patient data is limited by privacy and governance constraints. While Synthea provides a widely adopted framework for generating realistic longitudinal electronic health record data, its current implementation presents adoption barriers for many researchers and data scientists due to deployment complexity and limited integration with modern Python-based workflows. This paper introduces PySynthea, a Python-native reimplementation of Synthea designed to improve accessibility, extensibility, and interoperability within the scientific Python ecosystem. The framework provides modular synthetic patient generation, configurable healthcare simulation pipelines, and support for standard healthcare data formats while integrating naturally with tools such as pandas and machine learning workflows. By reducing operational c

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution

arXiv:2606.28335v1 Announce Type: new Abstract: We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional distribution $\mathbb{P}($position$\mid$context$)$ over a real political space. We evaluate nine current LLMs using a unified measurement framework anchored by VAA-CHES projection models, which map responses onto three validated dimensions (lrgen, lrecon, galtan) across six contextual axes. Our findings reveal high sensitivity to context: persuasive framing and under-represented languages displace coordinates by up to 0.57 and 0.52 units, respectively, while chain-of-thought reasoning often amplifies rather than dampens paraphrase instability. Despite this local plasticity, the model cohort occupies a remarkably narrow Overton envelope overall, occupying roughly one-third the spread of major European parties. Supported by a multi-trait multi-method (MTMM) analysis, we conclude that a single point cannot su

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Ground Truths in Suicide Research: The Current State of AI-Based Suicide Detection in Social Media

arXiv:2606.28334v1 Announce Type: new Abstract: Recent advances in artificial intelligence (AI) and social media data have led to growing optimism about the ability to detect suicide risk at scale. However, the empirical foundations of this work remain unclear. This article provides a synthesis of current research on AI-based suicide detection in social media, drawing on a recent umbrella review of 22 systematic reviews covering studies up to 2022, alongside an ongoing literature review extending the analysis to more recent work. Across these sources, we identified 195 relevant studies, which are documented in a detailed supplementary dataset outlining their key characteristics and findings (see Supplementary Information). Analysis of these studies reveals consistent patterns, including rapid growth, concentration on a small number of platforms, reliance on textual and English-language data, and repeated use of similar datasets. Most importantly, the majority of studies rely on indirec

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Insidious by Design: Implications of Large Language Model algorithmic bias for the Global South

arXiv:2606.28333v1 Announce Type: new Abstract: \begin{quote} The biases in Large Language Models' (LLMs) outputs remain inadequately theorised, particularly from the perspective of the Global South. This article reports on a small-scale exploratory study in which identical prompts were submitted to four major LLMs (ChatGPT, Claude, Grok, and Copilot), firstly, prompting for stories using names suggestive of specific racial and gender communities, and secondly asking questions about `development'. Drawing on critical AI scholarship and postcolonial theory, we argue that LLM outputs are patterned in ways that reproduce racial hierarchies, gender asymmetries, and Western-centric epistemic frameworks. We argue that these biases are insidious: they operate below the threshold of both obvious error and overt prejudice, and instead are subtly embedded in narrative structure and emotional template. Simply put, women, in LLM narratives have rich interior lives, while men make plans. Black peop

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

arXiv:2606.28332v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains poorly understood. We introduce \textsc{MedHarm}\footnote{Code and data will be released upon acceptance. Due to the sensitive nature of high-risk medical queries, data access will be available to qualified researchers upon request.}, a high-risk medical safety benchmark with 1,100 medically grounded queries across 10 safety-critical categories, including toxicology, pharmacology, covert poisoning, anesthesia, and fetal harm. Unlike broad medical QA benchmarks, \textsc{MedHarm} targets realistic clinical, educational, and technical prompts that require refusal, caution, or safe redirection rather than direct helpfulness. We evaluate 15 LLMs spanning general-purpose, medical-purpose, closed-source, and downstream SFT models, together with 4 representative guardrail models. Results reveal a sub

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

"AI Watermarking": Bridging Policy Discourse and Technical Capabilities

arXiv:2606.28331v1 Announce Type: new Abstract: The widespread deployment of generative artificial intelligence (AI) models has raised serious concerns about the proliferation of AI-generated content. This has led to a surge of interest in, and demand for, reliable tracking and detection mechanisms for content that is AI-generated, such as watermarking, metadata tagging, content tagging, and more. The problem has captured the attention of policymakers as well as the popular media, and a spate of recent bills in the US have sought to regulate the spread of AI content, and enforce or promote methods to track and label it. This work performs a critical analysis of the policy discourse surrounding generative AI content transparency in the US and EU. Through a broad document selection methodology, we first collect a broad corpus of documents containing legislative language and policy-relevant discourse on the topic. We then analyze these through inductive coding, and leverage our coding to

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

arXiv:2606.28325v1 Announce Type: new Abstract: Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a seven-axis measure of digital support, and applied it to the 300 writing systems of the Global Script Database (Fukui, 2026). Only 29 scripts (9.7%) are fully supported by contemporary digital infrastructure; among 158 living scripts, 60 (38.0%) lack complete support. Tokenizer efficiency varies by a factor of 31.7 across 45 scripts measured with parallel text. A serial mediation model -- imperial intervention to speaker population to web corpus to tokenizer efficiency -- is consistent with full mediation, with the direct effect of empire indistinguishable from zero (beta = -0.22, p = 0.39) and structural equation model fit indices indistinguishable from saturation at n = 45; the bias-corrected bootstrap CI grazes zero, and we treat the mediation as suggestive rather than confirmatory. Across

Source ↗
technology Tue, 28 Jul 2026 22:57:34 +0000
MedCity News

Smart Toilet Sensor Startup Is Flush with $10M Funding

Throne Science — a startup selling an AI-powered toilet camera that tracks gut health, hydration and urinary function — closed a $10 million Series A funding round. CEO Scott Hickle said the goal is to become a “smoke detector for colorectal cancer,” filling a monitoring gap he noted the wearables industry has largely ignored. The post Smart Toilet Sensor Startup Is Flush with $10M Funding appeared first on MedCity News .

Source ↗
technology Tue, 28 Jul 2026 22:25:27 +0000
MedCity News

Included Health to Acquire Firefly Health to Expand Alternative Health Plan Offering for Employers

Included Health announced plans to acquire Firefly Health to accelerate its alternative health plan strategy for employers. The post Included Health to Acquire Firefly Health to Expand Alternative Health Plan Offering for Employers appeared first on MedCity News .

Source ↗
technology Tue, 28 Jul 2026 21:21:08 +0000
MedCity News

Altimmune’s GLP-1 Drug Reduces Heavy Drinking in Alcohol Use Disorder Trial

Altimmune’s pemvidutide achieved a statistically significant reduction in heavy drinking days compared to a placebo in a Phase 2 trial. This peptide drug, which is designed to activate the GLP-1 and glucagon receptors, is also on track to begin a pivotal test in the fatty liver disease MASH. The post Altimmune’s GLP-1 Drug Reduces Heavy Drinking in Alcohol Use Disorder Trial appeared first on MedCity News .

Source ↗
technology Tue, 28 Jul 2026 16:48:46 -0400
EdTech Mag (K-12)

Closing the Visibility Gap in K–12 IT With Observability

A teacher has issues accessing a learning platform. Students struggle to log in to an online assessment. A classroom full of devices suddenly loses connectivity. For K–12 IT teams, solving those problems has become increasingly complicated, with today’s tech environments spanning on-premises networking equipment, cloud-hosted applications, Software as a Service (SaaS) platforms and thousands of endpoint devices spread across multiple schools. When an issue occurs, the root cause could reside almost anywhere. That complexity is pushing many districts to look beyond traditional monitoring tools…

Source ↗
technology Tue, 28 Jul 2026 15:30:13 +0000
HN: education

Peter Thiel: The Education of a Libertarian

Article URL: https://www.cato-unbound.org/2009/04/13/peter-thiel/education-libertarian/ Comments URL: https://news.ycombinator.com/item?id=49085418 Points: 2 # Comments: 3

Source ↗
technology Tue, 28 Jul 2026 13:28:00 +0000
MedCity News

Why Specialized UM Delivers on Every NICU Outcome

Done well, NICU UM is so much more than an approval process. It is a clinical credibility engine that — when integrated with Case Management (CM) — drives every NICU admission toward the most appropriate, impactful care. The post Why Specialized UM Delivers on Every NICU Outcome appeared first on MedCity News .

Source ↗
technology Tue, 28 Jul 2026 13:15:00 +0000
MedCity News

Spotting Alzheimer’s Years Earlier Through AI and Clinically Meaningful Insights

A few minutes of interaction with a digital tablet and speaking can reveal patterns that once required advanced imaging, years of observation, or were just completely undetectable. The post Spotting Alzheimer’s Years Earlier Through AI and Clinically Meaningful Insights appeared first on MedCity News .

Source ↗
technology Tue, 28 Jul 2026 09:00:00 +0000
Tech & Learning

Neo vs. Chromebook: The Ultimate Classroom Bake-Off

Has Apple finally built a laptop that belongs in the Chromebook conversation?

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model

arXiv:2607.22083v2 Announce Type: replace-cross Abstract: We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diversity of executable environments, task assets, and agentic scaffolds through real-world deployment and large-scale synthesis. Our RL pipeline applies mixed-mode RLHF over Think and Non-Think responses to improve overall model quality and reduce failure cases, length-controlled reasoning RL to balance accuracy and reasoning efficiency, and agentic RL with outcome and process rewards to stabilize long-horizon training. Extensive evaluations

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

arXiv:2607.13408v2 Announce Type: replace-cross Abstract: Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limited supervision on instruction-level correctness. We propose an instruction-level framework that uses audio-aware large language models (ALLMs) as fine-grained judges to verify target event presence and temporal relations in generated audio. After validating ALLM judgments on benchmarks and through human verification, we use their feedback to construct preference pairs for direct preference optimization. We further introduce S3Bench, a narrative benchmark for evaluating multi-event temporal instruction following. Experiments show that our method improves event completeness, temporal ordering, and joint instruction-following accuracy across existing benchm

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents

arXiv:2606.18037v2 Announce Type: replace-cross Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, databases, clinical records, and formulary tools. Standard factuality metrics usually test whether an answer is supported by pooled evidence, missing a provenance-sensitive failure mode: a claim may be supported somewhere while being attributed to the wrong source. We call this cross-source conflation. We introduce ProvenanceGuard, a source-aware verifier for MCP-grounded answers. It consumes captured MCP traces with stable tool IDs, source IDs, and raw outputs; decomposes answers into atomic claims; routes claims to source-specific evidence; checks support with NLI and a token-alignment proxy; compares stated attribution with the routed source; and returns per-claim verdicts plus an answer-level allow/block decision. Blocked answers can be repaired with retrieval-augmented answer revisio

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory

arXiv:2605.09877v4 Announce Type: replace-cross Abstract: Recall presents a difficult choice: transformers have a linearly growing memory that slows each successive token, while linear RNNs typically have fixed costs but limited recall. We present Key-Value Means ("KVM"), a novel block-recurrence for attention that can accommodate either fixed-size or growing state. Equipping a strong transformer baseline with fixed-size KVM attention layers yields a strong $O(N)$ chunked RNN, while adding only an insignificant number of new parameters. We train a transformer with a growable KVM cache and show it performs competitively on long-context tests with only subquadratic prefill time and sublinear state growth. KVM is implementable with standard operations and without custom kernels, and supports chunk-wise parallelizable training and prefill. It provides many of the benefits of both traditional transformers (expandable context memory, chunk-wise parallelizable training and prefill) and RNNs i

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning

arXiv:2601.09624v2 Announce Type: replace-cross Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably across individual samples: some are reliably erased, while others persist despite the same procedure. We argue that this disparity is not only a data-side phenomenon, but also reflects model-internal mechanisms that encode and protect memorized information. We study this problem from a mechanistic perspective based on model circuits--structured interaction pathways that govern how predictions are formed. We propose Circuit-guided Unlearning Difficulty (CUD), a {\em pre-unlearning} metric that assigns each sample a continuous difficulty score using circuit-level signals. Extensive experiments demonstrate that CUD reliably separates intrinsically easy and hard samples, and remains stable across unlearning methods. We identify key circuit-level patterns that reveal a mechanistic signature of di

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback

arXiv:2507.22080v2 Announce Type: replace-cross Abstract: Acquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation. While automated synthesis has emerged as an alternative to expensive manual curation, current approaches often rely on rigid heuristics, yielding data that is ungrounded or lacks logical complexity. We propose CodeEvo, a dual-agent architecture comprising a Coder for iterative solution synthesis and a Reviewer to orchestrate the generation trajectory. To transcend the limitations of existing heuristics, the Reviewer formulates a Schema to systematically architect logic and complexity through an interleaved synthesis of instructions and code. This process is further reinforced by a hybrid verification protocol synergizing deterministic compiler feedback with semantic evaluation. Under this framework, we construct CodeEvo-100K, a large-scale dataset of instruction-code pairs with stepped difficulty levels. Extensive e

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

PD$^3$: A Project Duplication Detection Framework via Adapted Multi-Agent Debate

arXiv:2505.17492v2 Announce Type: replace-cross Abstract: Project duplication detection is critical for project quality assessment because it helps avoid investment in repeated proposals. Existing methods usually cast it as ranking and rely on surface matching or direct large language models judging, often missing practical needs in set-level reference selection. We recast the task as many-to-many reference set selection, which requires broad candidate information and fair decomposed comparison under context limits. We propose PD$^3$, a framework for Project Duplication Detection via adapted multi-agent Debate. PD$^3$ combines local multi-agent debate with global round-robin scheduling to retrieve the relevant project set. Theoretically, this scheduler guarantees fair comparison through balanced exposure and comparison context. PD$^3$ also produces quantitative duplication scores and qualitative overlap feedback. On 800+ real-world power projects, PD$^3$ outperforms the strongest basel

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases

arXiv:2408.03910v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories. This challenge has prompted research on enhancing LLM-codebase interaction at a repository scale. Current solutions rely on similarity-based retrieval or manual tools and APIs, each with notable drawbacks. Similarity-based retrieval often has low recall in complex tasks, while manual tools and APIs are typically task-specific and require expert knowledge, reducing their generalizability across diverse code tasks and real-world applications. To mitigate these limitations, we introduce CodexGraph, a system that integrates LLM agents with graph database interfaces extracted from code repositories. By leveraging the structural properties of graph databases and the flexibility of the graph query language, CodexGraph enables the LLM agent to construct and execute queries, allowing for precise, code

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

Which Models Perform Better in Inheritance Reasoning?

arXiv:2606.13751v4 Announce Type: replace Abstract: This paper presents the participation of team PSL in the QIAS 2026 Shared Task on Arabic Islamic inheritance reasoning. The task evaluates the ability of large language models to solve inheritance cases that require legal interpretation, multi-step reasoning, and precise numerical computation. We compare \textit{commercial} and \textit{open-source} models under a unified prompting strategy to assess their effectiveness in structured legal reasoning with minimal task-specific adaptation. \\ Our results show a clear gap in reliability between the two model families. Commercial models demonstrate stronger performance in identifying eligible heirs, applying exclusion rules, and maintaining consistency across reasoning steps. In contrast, open-source models exhibit greater instability, particularly in cases involving dependent legal decisions and fractional share adjustments. The best performance is achieved by \textit{Gemini 2.5 Flash}, w

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

Speculative Pipeline Decoding: Higher-Accuracy Drafting with Hidden Latency via Pipeline Parallelism

arXiv:2605.30852v3 Announce Type: replace Abstract: Speculative Decoding (SD) accelerates low-concurrency LLM inference with a draft-then-verify paradigm. Mainstream methods, however, rely on multi-token prediction, which incurs compounding prediction difficulty and exposed draft latency. We propose Speculative Pipeline Decoding (SPD), which partitions the target LLM into $n$ pipeline stages so that $n$ tokens of a single sequence advance in parallel. To keep the pipeline saturated, a Pipeline Draft Module (PDM) aggregates multi-depth target features to predict the next token and runs concurrently with each pipeline step, yielding bounded prediction difficulty, higher acceptance, and hidden draft latency. Experiments show that SPD achieves higher theoretical and wall-clock speedup than EAGLE-3 at moderate pipeline width, while more aggressive widths still leave room for further gains. Our code is available at https://github.com/yuyijiong/speculative_pipeline_decoding

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios

arXiv:2605.25758v3 Announce Type: replace Abstract: Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives continuously and fine-grained profiles evolve rapidly. To bridge this gap, we introduce StreamProfileBench, a large-scale benchmark for fine-grained streaming user profiling. We formalize streaming user profiling as a continuous state maintenance task and curate a highly authentic dataset comprising over 120,000 UGC posts from 7,000+ real users across five diverse platforms. By leveraging the temporal correlation of user interests, we further propose a novel, annotation-free evaluation framework. Extensive experiments across 14 leading LLMs reveal that continuous profile updating remains an open challenge. Models exhibit a systemic conservative bias, over-retaining past interests while failing to recognize intere

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CL

Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs

arXiv:2605.09239v2 Announce Type: replace Abstract: Large language models fail at counting how many times a word repeats in a list, even though they perform well on far harder reasoning tasks. These failures are commonly attributed to limitations in internal count tracking. We show this attribution is wrong. Linear probes on the residual stream decode the correct count with near-perfect accuracy at every post-embedding layer and they do so even at the exact layers where the wrong answer crystallizes in the output. Attention patterns show no evidence of collapse over repeated tokens and tokenization artifacts account for none of the failure. Instead, a multi-layer perceptron (MLP) block at roughly 85--93\% network depth overwrites the correctly-encoded count with a fixed wrong answer. Ablating this block changes the wrong output and establishes it as causally responsible for the failure. The block fires on the space-separated repeated-word format and is absent for repeated digit-tokens.

Source ↗
Showing 2351–2400 of 10879 signals
← Prev Page 48 of 218 Next →