EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Two-Phase Simulated Annealing for Equitable Team Formation: Eliminating Complaints in Large Engineering Cohorts

arXiv:2606.07270v2 Announce Type: replace Abstract: Contribution: This paper presents a novel two-phase algorithmic approach that decouples preference satisfaction from fairness optimization in student team formation, achieving both objectives without compromise. The method applies simulated annealing -- a core materials science technique -- to an educational challenge, demonstrating pedagogical integration of administrative processes. Background: Forming effective teams in large engineering cohorts (100+ students) requires balancing student preferences, academic fairness, and demographic diversity. Existing tools either optimize for fairness while ignoring preferences (CATME, Team-Anneal) or accommodate preferences while compromising balance (self-selection), leaving complaint rates at 5--35%. Intended Outcomes: Eliminate formal complaints, achieve near-zero GPA variance between teams, prevent gender isolation, and maintain high preference satisfaction while creating a scalable, repro

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Access Timing as Scaffolding: A Reinforcement Learning Approach to GenAI in Education

arXiv:2605.15850v3 Announce Type: replace Abstract: In recent years, generative AI (GenAI) in educational settings has become ubiquitous in university students' daily lives, despite its potential to induce over-reliance, metacognitive disengagement, and diminished learning when used unrestrictedly. While most prior research has focused on how to pedagogically scaffold its usage, the question of when to allow off-the-shelf GenAI remains understudied and lacks pedagogically grounded empirical investigation. We treat access timing itself as a form of implicit scaffolding and operationalize it through a reinforcement learning (RL) agent that decides when students should access GenAI, with a reward function grounded in metacognitive theory, cognitive load theory, and productive failure. In a mixed-methods controlled lab study with N=105 higher education students, we compared the agent's effect on learning gains and metacognitive engagement to unrestricted and fully restricted use. Results s

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Faster Results from a Smarter Schedule: Reframing Collegiate Cross Country through Analysis of the National Running Club Database

arXiv:2509.10600v5 Announce Type: replace Abstract: Collegiate cross country teams often build their season schedules on intuition rather than evidence, partly because large-scale performance datasets were not publicly accessible prior to the National Running Club Database (NRCD). We analyze the comprehensive-era Cross Country subset of NRCD, 23,355 results from 7,083 athletes (2023-2025; >97% course/weather coverage). Under leakage control and temporal validation, race-result features do not support out-of-year forecasting of individual improvement (best men's R^2=0.043; women's -0.029), capturing only a small fraction of the outcome's reliability ceiling (~0.23-0.28). Against this null, team race frequency associates with nationals placement (pooled RR =2.09; GEE OR =2.56/SD). Program-wide opportunity (roster depth; Effective Racing Opportunity) outranks a single workhorse's max race count cross-sectionally, but overall team depth for race count is controlled. `Converted Only' times

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

arXiv:2608.11171v1 Announce Type: cross Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021-2022, comprising 37% of papers by 2025-2026), fairness remains the most consistent theme, and exp

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Templated or fully Synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance

arXiv:2608.11008v1 Announce Type: cross Abstract: Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. The recent IssueBench framework substantially mitigates these limitations with templated prompts anchored in real-world chat logs. Given the rise in non-work-related use of GenAI assistants, we extend IssueBench beyond writing assistance to include two additional tasks, information seeking and opinion sharing. We argue that templated prompts still lack the nuance of real ones, especially for open-ended tasks, and remain recognisable as evaluation artefacts. We propose the use of fully synthetic (LLM-generated) prompts, produced under detailed instructions with real prompts as seeds. We assess the ecological validity of real, templated, and LLM-generated prompts in a s

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Auditable AI-Assisted Research Writing: An Engineering Discipline with Pre-Registered Process Observation

arXiv:2608.10858v1 Announce Type: cross Abstract: Language models now draft, classify and criticise inside research production, yet the artifacts they help produce carry little accountable history. Rather than detecting machine involvement afterwards, we specify an auditability discipline built at production time: git sealing with an anchor lineage, hash-bound provenance, red-line gates that refuse non-compliant artifacts and log every refusal, cross-model role separation, and programmatic assembly from registered sources. Adherence is instrumented by metric cards, each carrying a pre-registered blind spot and evidential standing, frozen before the prospective case it observes. In that case the observed project's pre-registered confirmatory test was executed under seal and returned No-Go, and that project's frozen stopping rule halted the work, against its own operators. A lower-graded retrospective case covers families whose machinery predates the protocol. Current observations are pr

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI-Generated Interactive Fiction for Educational Use: A Pilot Study of Perceived Comprehensibility, Coherence, and Engagement

arXiv:2608.10818v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) can produce educational content at scale, including interactive and narrative learning experiences, but technical generation alone is not sufficient: scenarios that are confusing, narratively inconsistent, or unengaging are unlikely to be useful in practice. This paper presents a pilot user-centred evaluation of AI-generated interactive fiction (IF) for educational use in higher education. Using a previously described domain-agnostic pipeline and a shared STEM content base, we generated a controlled pool of scenarios and asked participants (N = 22, STEM higher-education) to play one generated episode and rate it on narrative clarity, story-content coherence, engagement, and length acceptance. A free-text prompt captured open feedback. Narrative clarity and length acceptance were rated positively, engagement sat near the neutral mid-point of the scale, and story-content coherence was the weakest di

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Unveiling the Predators: Contemporary Approaches to Identifying Illegitimate Open Access Journals in the Academic Publishing Ecosystem

arXiv:2608.10739v1 Announce Type: cross Abstract: Predatory journals pose a significant challenge to the integrity of the Open Access (OA) publishing model by exploiting its framework for financial gain while bypassing essential editorial and peer-review standards. This study critically evaluates existing methodologies for identifying such journals, ranging from manual blacklist checks to advanced automated approaches utilizing machine learning. The analysis highlights critical limitations, including the lack of a universally accepted definition of predatory journals, over-reliance on binary classification systems (e.g., blacklists and whitelists), and issues with scalability, reliability and interpretability. To address these shortcomings, this paper introduces a novel methodology based on multivariate graph analysis. By modeling the academic publishing ecosystem as a network of interconnected entities (such as authors, articles, journals, and publishers), this approach provides broad

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Most biomedical publications show signs of LLM-assisted writing

arXiv:2608.10715v1 Announce Type: cross Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but e

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators

arXiv:2608.10492v1 Announce Type: cross Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students but also to think like them. INSIDE generates internal dialogue grounded in Bloom's Taxonomy across cognitive, affective, and action dimensions, and fine-tunes models on paired think traces and actions. We baseline against different prompting frameworks and evaluate on two axes: fidelity of simulated actions and quality of generated internal dialogue. Our evaluations show that INSIDE improves simulation fidelity in both action fidelity, matching c

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

When the Interviewer Is a Bot: Behavior, Breakdowns, and Trust in MLLM-Led Interviews

arXiv:2608.10412v1 Announce Type: cross Abstract: Semi-structured interviews are a cornerstone of qualitative research but remain labor-intensive. We report an empirical study of what actually happens when the interviewer is an off-the-shelf real-time multimodal LLM (MLLM). We built InterviewBot, a voice-based interviewing system that wraps a real-time MLLM with a researcher-authored outline, and deployed it not as a novel architecture but as a research instrument for observing default MLLM interviewing behavior. In a practice study (N=15), participants completed a bot-led semi-structured interview and then a human-led reflection session about that experience. We contribute (i) a turn-level behavioral analysis of an MLLM interviewer (N_turns=428) showing that it is acknowledgment-heavy but probe-light (deepening probes account for 4.9% of all turns), and that 28.7% of question-bearing turns pack multiple questions into one turn despite an explicit one-question-at-a-time instruction; (i

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Fine-Tuning Large Language Models for Codebook-Guided Coding of Students' Mathematics Metaphor Responses

arXiv:2608.10276v1 Announce Type: cross Abstract: Student-generated metaphors about mathematics can reveal students' attitudes, beliefs, identities, and experiences, but human expert coding of these thematically and semantically complex open-ended responses is time-intensive and difficult to scale. This study examines whether LoRA-based supervised fine-tuning of large language models (LLMs) can improve their performance on codebook-guided coding tasks for student mathematics metaphors. We used a human-coded corpus of 2,265 Grade 6-8 responses to food- and animal-based metaphor prompts and instructed the LLMs to perform two coding tasks: valence-intensity coding to capture the direction and strength of students' affective orientations toward mathematics, and thematic coding to capture students' framings of mathematics as expressed through their metaphors. We compared two proprietary models, GPT-4o mini and GPT-5 mini, under prompt-only conditions with two open-weight models, DeepSeek-R1

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Toward Human Rights Benchmarking for LLMs: A Pilot Methodology

arXiv:2608.10268v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this end, we report our efforts to develop a robust and scalable methodology for creating HumRightsBench: the first expert-validated, scenario-based benchmark for evaluating reasoning grounded in the obligation structure of international human rights law. We adapt the IRAC framework for legal reasoning to better suit the unique reasoning patterns of human rights work (substituting P, "proposing remedies," for C, "legal conclusion," yielding IRAP) to structure our evaluation heuristics. We also produce a pilot series of authentic scenarios designed to implicate the many dimensions of real-world human rights issues and annotated by human rights lawyers and professionals across the world. Ultimately, we find that model

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse

arXiv:2608.10186v1 Announce Type: cross Abstract: LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer exists and where decision quality instead depends on integrating pluralistic perspectives to find mutually acceptable solutions. We argue that LLM reasoning capacity on this class of problems cannot be fully inferred from verifiable-task benchmarks, and that procedural evaluations of LLM discourse (respectfulness, justification, engagement) are systematically insufficient. We apply the Deliberative Reason Index (DRI), a measure developed in political science and validated across citizen assemblies, as a tool for evaluating reliable group-level reasoning on pluralistic, non-verifiable problems. Synthesizing recent evidence a

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Human versus Computer Vision

arXiv:2608.10181v1 Announce Type: cross Abstract: Computer vision saliency models predict where people will look, one map per image, and a billion-dollar predicted-attention industry sells those maps in place of measuring real viewers. I test the leading models from the audience side, against 11.4 million webcam gaze points from 3,023 US adults recruited to national quotas, viewing circulating news photographs. I show that an untrained central marker outperforms every trained network, because the content the networks add on top of the center falls where these audiences never look. What accuracy remains is systematically biased, favoring younger, White, and moderate viewers over older, Black, and ideologically extreme ones. I propose a way forward and build on what a group's own gaze reveals about whether a model can learn that group at all, and I apply it across every demographic axis this sample supports. Ultimately, I show how systems that decide what people see can learn to see ever

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Status Association Does Not Reliably Predict Decision Leakage

arXiv:2608.10089v1 Announce Type: cross Abstract: Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Detecting Soft Skills in ML Engineering Roles CVs

arXiv:2608.10046v1 Announce Type: cross Abstract: Soft skills shape collaboration among ML engineers, data scientists, and software engineers building ML-enabled systems, yet what we know about them comes almost entirely from the demand side. Job advertisements, surveys, and hiring manager interviews capture what employers ask for. How candidates themselves articulate these competencies has not been studied, and existing CV-mining work is both keyword-based, so it cannot see skills conveyed through narrative, and descriptive, reporting frequency rankings without testing whether group differences exceed sampling variation. We close both gaps. Using a balanced corpus of 300 curated CVs spanning the three roles, we extract explicitly listed and implicitly narrated soft skills with an LLM-based pipeline validated against a human-annotated ground truth, a distinction that existing extractors were not designed to make. We then convert the demand-side literature's claims into 13 falsifiable h

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models' Carbon Footprint

arXiv:2608.09998v1 Announce Type: cross Abstract: Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks. Despite their benefits, growing attention has been directed toward their environmental implications, primarily due to their high energy demands and associated carbon emissions. This concern is particularly relevant in light of the increasing deployment of large-scale models, especially Deep Learning (DL) architectures, which provide advanced predictive capabilities but require substantial computational resources. This paper presents a systematic review of research on Green AI, Green DL, and optimization techniques aimed at reducing the environmental impact of AI models. In addition, we examine and compare several carbon measurement tools for estimating emissions generated by AI algorithms. To complement the review, we conducted an empirical evaluation using a CPU-based experimental setup, in which six DL m

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Co-Lecturing With the DED: Explaining Circuit Design via the Draw Encode Display Loop

arXiv:2608.09945v1 Announce Type: cross Abstract: When representing digital circuits, 2 dimensional hand drawings free us from the linear structure of hardware description languages, enabling intuitive reasoning and making structure explicit. However these drawings are imprecise and inert: they do not enforce that the circuits are well defined and cannot be tested. We want both intuitive visual representations and well defined testable ones but students can struggle to link one to the other. To bridge this gap we present the Draw Encode Display Loop (DED), a Co-Lecturing dynamic which equips students with a systematic method to tackle natural language specifications: 1. Draw: visually informative intermediate representations (truth tables, characteristic tables) to generate a structured circuit diagram. 2. Encode: the diagram by labelling inputs, outputs and intermediate values which can be directly converted to code. 3. Display: the code using an in-house diagrammatic renderer. This i

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

The impact of design factors of virtual and augmented reality on tertiary students user experience in Metaverse

arXiv:2608.09940v1 Announce Type: cross Abstract: The Metaverse is a convergent space integrating virtual reality (VR) and augmented reality (AR) technologies, with market projections rising from \$65.5 billion in 2022 to \$1.3 trillion by 2030. Despite rapid adoption in education, the specific contributions of visual elements, environmental design, and communication features to user experience (UX) remain underexplored, limiting evidence-based design and resource allocation. This study examined how these design factors in VR and AR environments influence UX in Metaverse platforms. Using a correlational research design, data were collected from 321 purposively sampled tertiary students from engineering and computer science departments across four higher education institutions, all familiar with Metaverse platforms. A structured questionnaire with validated 5-point Likert scales measured UX; visual elements (field of view, resolution, color, complexity, and style); environmental design

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory

arXiv:2608.09937v1 Announce Type: cross Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural environments within a country. In this work, we leverage cultural consensus theory (CCT) from cultural anthropology to model such multidimensional nuance. Applying CCT to the World Values Survey (WVS) across 10 countries and 12 domains, we demonstrate that models frequently misrepresent cultural structures by either failing to form cohesive consensus or severely over-regularizing consensus. Through explicit representation of intra-group variance, CCT provides actionable diagnostics to evaluate when models reflect true human diversity versus algorithmic homogenization.

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Mediatised Participation: Citizen Journalism and the Decline in User-Generated Content in Online News Media

arXiv:2608.11159v1 Announce Type: new Abstract: The second generation of web tools shook the journalist profession approximately two decades ago with the proactive incorporation of audiences into the media. Citizen journalism and user-generated content arose as an object of interest due to the democratising value of participation attributed to them, with empowered citizens who could emulate the professional and institutional practises of journalists. However, difficulties soon came to the surface, and audience participation in news media began to be limited. Within this context, this article conducts a critical review of studies on audience participation in news media based on a systematic literature review. The results indicate that, in general, audiences showed low interest in the creation of informative content and that their participation has grown increasingly problematic. In addition, journalists are reticent as they defend their professional role above all else, while company st

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Who Uses Open-Weight Models? China and the Shifting Geography of AI in Science

arXiv:2608.11090v1 Announce Type: new Abstract: As LLMs have become a flashpoint for scientific research, computer scientists and STS scholars have advocated the use of open-weight models. Since LLM research has matured and more high-quality model families are available, have researchers adopted open-weight models? We present the first systematic study of model selection in scientific research, analyzing 21 million full-text articles through June 2026 from the Semantic Scholar Open Research Corpus (S2ORC). We employ a mixed NLP pipeline to extract model occurrences in article full text and determine whether they are used or merely mentioned by researchers. We divide our corpus into single- and multi-model family studies, which we take as a proxy for applied and foundational AI research. We find GPT-family models dominate both single- and multi-family research, but that both areas are becoming more diverse over time. In single-family papers, open-weight model use rises steadily, reachin

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis

arXiv:2608.11006v1 Announce Type: new Abstract: Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies. Comparing national and regional AI strategies to identify their convergences and divergences can uncover their common practices, understand regional variations, and provide policy designers a comprehensive set of policy design elements for their ongoing AI strategy developments. Yet, existing work has not examined their underlying policy design elements or assessed whether those elements are horizontally (country-to-country) or vertical (region-to-country) converging or diverging over time. This paper addresses that gap by coding and analyzing 74 national and 3 regional AI strategies drawn from a global scan of all 205 UN member and non-member states. The coding used a latent-inductive approach organized around three functional policy design elements: goals, approaches, and principles. Two research questions guided the anal

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Technology, education and critical media literacy: potential, challenges, and opportunities

arXiv:2608.10778v1 Announce Type: new Abstract: This study examines the impact of technology within media education, media literacy, and educommunication, and explores how these fields are perceived and understood by students and academic experts, focusing on the development of critical competencies and critical media literacy. Based on semi-structured in-depth interviews with leading experts in the field of critical media literacy, and a survey conducted with 141 university students in Communication and Education programs, this study explores how recent technological advances are linked to challenges in information consumption-such as disinformation, fake news, incidental exposure to information, and deepfakes-as well as the challenges and opportunities these issues present within educational contexts. The results reveal that, although such technologies provide opportunities to improve teaching-learning processes, their inclusion in the curriculum is limited and often superficial. In

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms During the 2025 Parliamentary Election

arXiv:2608.10773v1 Announce Type: new Abstract: The increasing use of Generative Artificial Intelligence (GenAI) in journalism raises concerns about possible detrimental effects both on journalism and its democratic function. We explore these risks through a case study of GenAI in Norwegian Newsrooms during the 2025 parliamentary election campaign. Based on interviews with managers and journalists over a ten-month period, we analyse how ambitious visions fared in the face of technological and practical challenges. We highlight the risk of an internal threat stemming from the journalists' own use of AI, contrasting the dominant focus on external disinformation threats. We show how newsroom managers shared sociotechnical imaginaries resulting in unrealistically optimistic beliefs about the capabilities of the technology and the pace of development, leading to plans for audience-facing GenAI services collapsing and giving way to more mundane uses of GenAI tools internally in the newsrooms

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGI

arXiv:2608.10730v1 Announce Type: new Abstract: The pursuit of artificial general intelligence (AGI) rests on a seemingly self-evident premise: that general intelligence, the kind of flexible, domain-general cognitive capacity exemplified by Homo sapiens, is extraordinarily valuable. This paper subjects this premise to critical scrutiny. We first present the intuitive case for the value of general intelligence before mounting an evolutionary challenge. We argue that, on evolutionary timescales, its adaptive value is far from empirically established. Numerous taxa, from cyanobacteria to horseshoe crabs, have persisted for hundreds of millions or even billions of years without anything resembling general intelligence, while Homo sapiens has existed for roughly 300,000 years and already faces self-generated existential risks. Mass extinction events do not preferentially favour cognitively sophisticated species. We argue that general intelligence may be the only biological strategy that ge

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Inferential Capability Does Not Determine Legal Scope

arXiv:2608.10601v1 Announce Type: new Abstract: Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capability to infer constitutively: it is the central feature separating the regulated category from conventional software. The GDPR never defines inference, yet governs it protectively: the consequences follow from the processing of personal data and from what the inference says about, or does to, a person, whether or not the technology that produced it qualifies as an AI system. The two perimeters are not concentric. Their non-coincidence remained invisible in single-shot systems; agentic architectures make it operationally acute. The thesis: inferential capability does not determine legal scope, and its absence does not create immunity. The framework is two-level. Inference performs two legal functions, constitutive and protective; the protective function operates through three pathways - identificatory

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking

arXiv:2608.10329v1 Announce Type: new Abstract: Notice-and-comment rulemaking gives any affected party the same formal right to influence federal regulation, but formal access is not substantive capacity to shape rule text. Existing strategies operate at the rule or aggregate-corpus level, too coarse to capture the discrete regulatory obligations where commenters seek change. We introduce obligation-level responsiveness auditing, an auditable, AI-assisted framework for measuring whether public-comment engagement co-occurs with changes to specific regulatory duties. The framework extracts proposed and final-rule obligations, matches comments to the obligations they address, and classifies proposed-final outcomes; each load-bearing component is evaluated against blind human judgment. We apply the framework to 70,075 comments across 36 EPA anchor rulemakings, drawn from a corpus of 786,197 comments across 6,145 dockets from 2010-2022. Three descriptive findings emerge. First, engagement i

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Context and Symmetry in Auditing: A Case Study of Skeleton Inference in Motion Capture

arXiv:2608.10194v1 Announce Type: new Abstract: Humans are increasingly expected to interact with AI systems that observe and make inferences about them - but do these systems actually work? A standard approach to answering this question is AI auditing. Conducting an AI audit requires identifying how a system behaves (i.e., determining what types of inputs to audit it with and then observing and documenting actual system behavior) and contrasting that with how a system should behave (i.e., determining what the nominal outputs of a system should look like). We argue that this is best done through a contextual audit, which we introduce as a method for auditing measurements within the context of the practices that produce them. We show how contextual auditing enables the interrogation of assumptions implicit in the audit process and allows auditors to be explicit about what serves as ground truth, which we define as verifiable measurements about the real world against which systems are ev

Source ↗
technology Wed, 10 Jun 2026 12:45:00 -0400
EdTech Mag (K-12)

4 Tips for Upgrading K–12 Physical Security Systems

Weapon detection solutions have evolved from locked doors and metal detectors to AI-powered systems. What were previously one-off purchases are now integrated with a broader, layered approach to safety that balances security with a more welcoming environment. Here are tips CIOs and CTOs should keep in mind as they consider physical security and related solutions. Click the link below to discover the benefits of modernizing your physical security infrastructure.

Source ↗
technology Wed, 10 Jun 2026 09:00:00 +0000
Tech & Learning

Customer Service Matters in Educational IT Support

Technology support in education is ultimately a service profession

Source ↗
technology Wed, 10 Jun 2026 09:00:00 +0000
eCampus News

Invisible translators: What funky techs do for higher ed

You don't see 'funky tech' in scholarly literature. But if you've spent time at higher-ed tech conferences, you've heard it. Someone introduces themselves during a networking break: "I'm a funky tech over in Academic Advising." The post Invisible translators: What funky techs do for higher ed appeared first on eCampus News .

Source ↗
technology Wed, 10 Jun 2026 06:31:26 +0000
HN: education

Computer Lessons - A history of computers in education

Article URL: https://technicshistory.com/2026/06/06/computer-lessons/ Comments URL: https://news.ycombinator.com/item?id=48472275 Points: 4 # Comments: 0

Source ↗
technology Wed, 10 Jul 2024 03:17:07 +0000
HN: medical education

Rethinking NEET, and Medical Education

Article URL: https://www.mayin.org/ajayshah/MEDIA/2024/medical_education.html Comments URL: https://news.ycombinator.com/item?id=40923378 Points: 1 # Comments: 0

Source ↗
technology Wed, 08 Jul 2026 17:06:15 -0400
EdTech Mag (Higher)

Rethinking Managed Security Services in Higher Ed: From Ticket Closers to Strategic Partners

Managed Security Service Providers used to be seen as a way to “outsource the pain” of monitoring and incident response. For many universities — where small teams are tasked with protecting sprawling, open and highly connected environments — that was reason enough to consider them. But in today’s threat landscape, simply shipping logs to a vendor and measuring them by ticket closure rates is no longer enough. The question for higher ed leaders is not, “Do we have an MSSP?” It’s “Is our MSSP helping us achieve our institutional outcomes?” Why Managed Security Service Providers Matter More Than…

Source ↗
technology Wed, 08 Jul 2026 16:34:17 -0400
EdTech Mag (K-12)

ISTELive 26: Accessible by Design: How K–12 Districts Are Building AI and UDL Into Every Classroom

Accessibility has always been a legal obligation. But leaders in the K–12 space are expanding that conversation to include how it could be more than that: a design principle that creates an improved learning experience for every student, not just the ones with IEPs. At ISTELive 26, Tara Nattrass, chief innovation strategist for education at Lenovo, broke down what hybrid AI means at the device level and why it matters for students who rely on accessibility tools. Nattrass helped lead seven sessions at the conference, including “Strengthening Accessibility and Inclusivity with Hybrid AI.” The…

Source ↗
technology Wed, 08 Jul 2026 13:57:47 -0400
EdTech Mag (Higher)

4 Critical Security Considerations for AI in Higher Education

Generative artificial intelligence is revolutionizing how colleges and universities operate, streamlining workflows, supporting research and enhancing learning. But as adoption grows, so do the risks. Higher education institutions manage sensitive student data, proprietary research and intellectual property. Without the right guardrails in place, AI systems can expose this information or violate compliance standards, putting institutions at risk for reputational damage. For CISOs and CIOs, securing AI environments must be a strategic priority. Here are four key security considerations when…

Source ↗
technology Wed, 08 Jul 2026 09:00:00 +0000
Tech & Learning

Handling Student Personal Relationships With AI

What to be aware of and do if your student is in a toxic or troubling relationship with an AI chatbot and how teaching AI literacy can help

Source ↗
technology Wed, 08 Jul 2026 09:00:00 +0000
eCampus News

Nobody’s a loser: What genuine education leaders realize

My father, Jake M. Schrum, would take me with him to cattle shows and sales. As an impressionable teenager, I treasured these outings mostly to have one-on-one time with my dad, but also, to learn about Hereford cattle and how to judge them. The post Nobody’s a loser: What genuine education leaders realize appeared first on eCampus News .

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CL

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

arXiv:2606.20023v2 Announce Type: replace-cross Abstract: As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, prior tool-selection studies focus on safety-agnostic metadata preferences, leaving privilege-sensitive choices underexplored. To address this gap, we study over-privileged tool selection, in which an agent selects or escalates to a higher-privilege tool despite a sufficient lower-privilege alternative. We introduce ToolPrivBench to evaluate whether agents choose higher-privilege tools despite sufficient lower-privilege alternatives, measuring both initial selection and escalation after transient tool failures. Across eight domains and five recurring risk patterns, we find that over-privileged tool selection is common among mainstream LLM agents and is further amplified by transient failures. We further find that general safety alignment does not reliably transfer to least-privilege tool choi

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CL

Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval

arXiv:2605.23826v2 Announce Type: replace-cross Abstract: Keyframe selection is a direct way to provide verifiable visual evidence for long-video question answering (QA). Queries differ in what they require, and finding the right frames depends on knowing what to look for. Existing keyframe selectors either score every frame against a single query, or decompose the query into a fixed schema evaluated by a single visual tool. We propose ToolMerge, a keyframe retrieval method based on decomposition and merging: an Large Language Model (LLM) based planner decomposes the query into tool calls and specifies how their per-tool rankings are merged using boolean operators. To evaluate retrieval directly, we construct Molmo-2 Moments (M2M), a benchmark in which every question is anchored to a specific time interval by construction. Across QA, question retrieval, and caption retrieval, ToolMerge is competitive with prior keyframe selectors, most notably on caption retrieval, outperforming other

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CL

Learning from Execution: Self-Evolving Memory for Private-Library Code Generation

arXiv:2604.24222v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved strong performance on general code generation, but their effectiveness drops sharply in enterprise settings where software development relies on internal private libraries absent from public pre-training corpora. Existing Retrieval-Augmented Generation (RAG) methods provide a training-free solution by retrieving static API documentation, but our analysis shows that documentation mainly helps models identify what APIs to use and remains insufficient for teaching how to use them correctly. Even with oracle API-document retrieval, LLMs still make recurring errors at the API, cross-API, and task levels, including API misuse or hallucination, flawed API composition, and incorrect solution strategies. To address this limitation, we propose MEMCoder, a training-free self-evolving memory framework for private-library code generation. MEMCoder augments existing RAG pipelines with a Multi-level E

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CL

EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving

arXiv:2604.22851v2 Announce Type: replace-cross Abstract: While Vision-Language Models (VLMs) have advanced high-level reasoning in autonomous driving, their ability to ground this reasoning in the underlying physics of ego-motion remains poorly understood. We introduce EgoDyn-Bench [Project page: (https://tum-avs.github.io/EgoDyn-Bench-Website/), Code: (https://github.com/TUM-AVS/EgoDyn-Bench), Dataset: (https://huggingface.co/datasets/fnc1901/EgoDyn-Bench)], a diagnostic benchmark for evaluating the semantic ego-motion understanding of vision-centric foundation models. By mapping continuous vehicle kinematics to discrete motion concepts via a deterministic oracle, we decouple a model's internal physical logic from its visual perception. Our large-scale empirical audit spanning 20$+$ models, including closed-source MLLMs, open-source VLMs across multiple scales, and specialized VLAs, identifies a significant Perception Bottleneck: while models exhibit logical physical concepts, they c

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CL

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation

arXiv:2603.15600v2 Announce Type: replace-cross Abstract: Accurate process supervision remains a critical challenge for long-horizon robotic manipulation. A primary bottleneck is that current video MLLMs, trained primarily under a Supervised Fine-Tuning (SFT) paradigm, function as passive "Observers" that recognize ongoing events rather than evaluating the current state relative to the final task goal. In this paper, we introduce PRIMO R1 (Process Reasoning Induced Monitoring), a 7B framework that transforms video MLLMs into active "Critics". We leverage outcome-based Reinforcement Learning to incentivize explicit Chain-of-Thought generation for progress estimation. Furthermore, our architecture constructs a structured temporal input by explicitly anchoring the video sequence between initial and current state images. Supported by the proposed PRIMO Dataset and Benchmark, extensive experiments across diverse in-domain environments and out-of-domain real-world humanoid scenarios demonstr

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CL

Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs

arXiv:2601.12494v3 Announce Type: replace-cross Abstract: Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dialect-rich settings such as Arabic-English remains challenging. We present a controlled study of multi-task instruction tuning for an Arabic-centric audio LLM across generative tasks, including automatic speech recognition (ASR) and speech and text summarization, as well as discriminative tasks, including dialect identification (DID) and speech emotion recognition (SER), in a resource-constrained setting. To support end-to-end Arabic speech summarization, we introduce AraMega-SSum, the first Arabic speech summarization dataset designed for training and benchmarking Arabic-centric audio LLMs. We compare four training strategies: (i) Uniform Mixing (UM), (ii) Task-Progressive Curriculum (TPC), (iii) Aligner-Based Diverse Sampling (ADS) for training-time batch construction, and (iv) a two-stage TP

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CL

Geometric Stability: The Missing Axis of Representations

arXiv:2601.09173v5 Announce Type: replace-cross Abstract: Representational similarity analysis and related methods compare the internal geometries of neural networks, but they measure only alignment between spaces, leaving a blind spot -- whether a representation's structure is reliably recoverable, not merely similar. We introduce geometric stability, a distinct axis, and \textit{Shesha}, a metric that quantifies it from a single representation by correlating dissimilarity matrices built from complementary random halves of the feature dimensions. Unlike CKA and Procrustes distance, Shesha is provably non-invariant to orthogonal rotations of the feature basis. This is by design: the basis is privileged for learned models, since probes, patching, and steering act on coordinates, and a rotation-invariant metric cannot see whether the targeted structure survives them. A double dissociation isolates the mechanism -- removing the top principal component collapses CKA while Shesha holds, whe

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CL

BabyVision: Visual Reasoning Beyond Language

arXiv:2601.06521v2 Announce Type: replace-cross Abstract: While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs s

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CL

SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models

arXiv:2512.18542v3 Announce Type: replace-cross Abstract: AI coding assistants produce vulnerable code in 45\% of security-relevant scenarios~\cite{veracode2025}, yet no public training dataset teaches both traditional web security and AI/ML-specific defenses in a format suitable for instruction tuning. We present SecureCode, a production-grade dataset of 2,185 multi-turn security training examples spanning two domains: web application security (1,435 examples covering the OWASP Top 10 2021 across 11 languages and 9 frameworks, 100\% grounded in documented CVEs and security incidents) and AI/ML security (750 examples covering all 10 OWASP LLM Top 10 2025 categories across more than 40 frameworks, including LangChain, OpenAI, and Hugging Face). Every example follows a 4-turn conversational structure -- feature request; vulnerable and secure implementations with attack demonstrations; advanced probing; and defense-in-depth operational guidance -- designed for direct use in instruction tu

Source ↗
technology Wed, 08 Jul 2026 00:00:00 -0400
arXiv cs.CL

LLM4Delay: Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation

arXiv:2510.23636v4 Announce Type: replace-cross Abstract: Flight delay prediction has become a key focus in air traffic management (ATM), as delays reflect inefficiencies in the system. This paper proposes LLM4Delay, a large language model (LLM)-based framework for predicting flight delays from the perspective of air traffic controllers monitoring aircraft after they enter the terminal maneuvering area (TMA). LLM4Delay is designed to integrate textual aeronautical information, including flight data, weather reports, and aerodrome notices, together with multiple trajectories that model airspace conditions, forming a comprehensive delay-relevant context. By jointly leveraging comprehensive textual and trajectory contexts via instance-level projection, an effective cross-modality adaptation strategy that maps multiple instance-level trajectory representations into the language modality, the framework improves delay prediction accuracy. LLM4Delay demonstrates superior performance compared

Source ↗
Showing 1101–1150 of 10876 signals
← Prev Page 23 of 218 Next →