Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2607.10331v1 Announce Type: cross Abstract: Human-centered AI (HCAI) refers to guidelines or principles that aim on ethi-cally oriented design of systems. We compare HCAI- guidelines with princi-ples of socio-technical systems that emerged in the context of conventional in-formation technology. The comparison leads to a revision of socio-technical heuristics by including aspects of AI-usage. The comparison reveals that con-tinuous evolution is a basic characteristic of socio-technical systems, and that human oversight or interventions and the subsequent appropriation of AI-systems lead to continuous adaptation and re-design of the systems, if autono-my is collaboratively exercised. From a socio-technical point of view, the cru-cial requirement of transparency has not only to be fulfilled with technical fea-tures, but also by contributions of the whole system including human actors. It will be promising for using AI, if not only technical features, but organization-al and social p
arXiv:2607.10322v1 Announce Type: cross Abstract: Based on the literature and several practical examples of possible AI applica-tions, we outline the concept of intervenability. This new phenomenon is not covered by emergency shutdowns, workarounds, or the reconfiguration of automated systems. Intervenability instantiates the principles of control-lability, autonomy, oversight, and keeping humans in the loop in the context of AI. We provide a taxonomy that encompasses a range of possibilities for intervening activities and differentiates them regarding the mental effort of the users. This taxonomy extends the scope of interventions from real-time control of automated processes to AI-based discrete case-related decision-making. This is in accordance with human-centered AI, which seeks to combine human strengths with the usage of AI. We demonstrate how inter-venability can potentially contribute to the ongoing development of human capabilities on the one hand and to further technical imp
arXiv:2607.10317v1 Announce Type: cross Abstract: Algorithmic decision systems in financial services often rely on data proxies that inadvertently encode structural inequalities. This paper introduces a hierarchical human-AI triage model for Point of Sale fraud detection in the Nigerian FinTech sector. Adopting a We Are All Equal worldview, we address the challenge of discrimination laundering, wherein the system misinterprets infrastructure related aleatoric noise such as rural network timeouts as fraudulent intent. We implement a three-tier routing policy utilizing a calibrated ensemble model as a primary filter. The policy routes transactions characterized by epistemic uncertainty such as cold start new accounts to specialist analysts while reserving high stakes cases for a senior supervisor. To manage finite human capacity, we utilize a dynamic shadow price to ration human attention and implement a random audit mechanism to prevent human skill atrophy. Our experimental results demo
arXiv:2607.10179v1 Announce Type: cross Abstract: Patent databases represent one of the largest public archives of technical knowledge, yet much of this knowledge remains difficult to identify, interpret, and reuse once patent rights expire or lapse. This paper proposes an AI-enabled framework for discovering expired and lapsing patents, identifying technology trends, and translating patent disclosures into business pathways. We use pathways to mean structured commercialization routes such as SaaS products, services, licensing packages, consulting playbooks, training offerings, data products, or internal process tools. The framework treats patent expiry as both a business signal and an archival transition, not primarily as a legal problem. Legal status remains important, but it is one risk-screening input alongside customer need, implementation feasibility, channel access, and market timing. We describe a system architecture that combines patent metadata, maintenance-fee records, legal
arXiv:2607.10163v1 Announce Type: cross Abstract: We study recovery of a low-rank tensor $\mathcal{T}$ and a low-rank matrix $M$ from sparse, noisy observations. $\mathcal{T}$ and $M$ share one mode. We relax tensor rank using the nuclear norm of the mode-1 unfolding. This unfolding carries the coupling. It also has an exact proximal operator. We couple $\mathcal{T}$ and $M$ through a learned linear operator $G$. We prove a minimizer exists for the ridge-stabilized penalized objective. We prove that a proximal alternating linearized minimization (PALM) scheme converges to a critical point, for the algorithm as implemented, by verifying the hypotheses of a known nonconvex block-coordinate convergence theorem against our objective and identifying which conditions come from this problem's structure. For the matrix-only sub-problem, we state a proven sampling bound from matrix completion theory. For the coupled problem, we prove a sample-complexity result for a sequential sub-case: a separ
arXiv:2607.09970v1 Announce Type: cross Abstract: Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults' susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play.AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the "relative-in-distress" category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-
arXiv:2607.09790v1 Announce Type: cross Abstract: The article investigates the fundamental problem of ensuring the stability of operator control and preserving goal-targeting in hybrid human-machine decision support systems (DSS) of a new generation. Based on a two-month continuous longitudinal experiment on the joint design of a monograph-format textual array, the latent phenomenon of semantic context drift in large language models of deep logical reasoning (Reasoning LLMs) is verified and described. A mathematical model of interaction in the human-machine interface is proposed, and an original metric is introduced - the operator control stability coefficient, which takes into account the non-linear contextual pressure of hidden reasoning chains. Within the paradigm of the cognitome theory, a critical point of control functions inversion is captured. Engineering recommendations are formulated for implementing dynamic relational arbitration loops based on a modified hierarchical simila
arXiv:2607.09705v1 Announce Type: cross Abstract: Since 2023, computer scientists have warned against model collapse -- the contamination of training sets with AI-generated outputs that progressively degrade model performance. Exemplifying a positive-feedback-driven failure, it produces effects such as word repetition or pixel noise, ultimately leading to a loss of meaning and coherence -- at least from an engineering standpoint. From a creative one, however, collapse is not merely a breakdown: it also functions as a recursive mirror that recalls early analog video feedback experiments, raising once again the question of what happens when a system turns inward and sees itself. In such cases, so-called machine vision no longer transmits the world (as in tele-vision) but increasingly generates worlds from within. Drawing on media archaeology through case studies of both historical video synthesis techniques and contemporary artistic uses of machine learning, this paper examines what recu
arXiv:2607.09685v1 Announce Type: cross Abstract: Policy models often assume that the relationship between a policy instrument and its outcome remains stable across institutional conditions. In adaptive socio-technical systems this assumption may fail: regulatory change can alter incentives, agents can respond strategically, and the mapping from policy variables to aggregate outcomes can change. This paper studies such regime change as a transfer-learning problem in adaptive multi-agent systems. A policy regime is represented as a learning problem induced by an observable input distribution and a target function mapping policy variables to outcomes. We compare a blank-slate learner that searches a flexible hypothesis class in the new regime with a transfer learner whose effective hypothesis class is restricted by structural knowledge from the previous regime. Transfer is beneficial when this restriction preserves the new target function while reducing effective complexity; it is harmfu
arXiv:2607.11859v1 Announce Type: new Abstract: Can large language models perform deep technical comprehension of computer architecture papers -- not summarization, but structured critique that names the core mechanism, surfaces buried assumptions, and connects a contribution beyond its own scope? We study Gauntlet, an open-source pipeline that analyzes a paper through five independent expert-persona reviewers and an adversarial synthesis stage. On 20 ISCA 2025 and HPCA 2026 papers, ten researchers each wrote their own analyses and then judged, for papers other than their own, the human analysis against Gauntlet's. Across the 20 comparisons evaluators preferred Gauntlet in 15 (human in 4, one tie); its advantage is significant on per-analyst totals (paired Wilcoxon, p < 0.01) and largest on Critical Rigor, vanishing only on Calibration. Where humans win, it is on trust and usefulness rather than depth: a confident wrong claim, a mechanism described but not taught, or unprioritized brea
arXiv:2607.11692v1 Announce Type: new Abstract: In this paper we present a study of students' mental models of generative AI (GenAI). A student's mental model of GenAI influences not only how they perceive the technology's capabilities and limitations but also how they choose to integrate it into their academic work. Whether they view it as a collaborative partner, a shortcut to complete tasks, or something in between, depends on how they conceptualize its use. This study addresses the following questions: (I) What mental models do undergraduate students hold about GenAI? and (II) What aspects of conceptual knowledge - declarative, procedural, and conditional - are present in these mental models? Sixty-four concept maps were collected from students enrolled in a course on technology ethics. Students were asked to construct concept maps representing their understanding of GenAI use. The concept maps were analyzed using a structured codebook and the analysis revealed five categories of m
arXiv:2607.11676v1 Announce Type: new Abstract: Crafa's algorithmic analysis of the Italian electoral law (D.P.R. 361/1957, as amended by Law 165/2017, the Rosatellum) showed that the statutory text distributing proportional seats among territories (Art. 83(1)(h)) admits at least three algorithmic interpretations, which can elect different people from the same votes. We test this empirically: we implement the full seat-allocation pipeline (Arts. 77, 83, 83-bis, 84, 85) and run the three interpretations on the complete open data of the Italian general election of 25 September 2022 (Chamber of Deputies). The implementation reproduces the official national apportionment exactly from raw municipal data, matches by name 389 of the 391 seats it models (99.5%), and agrees step by step with the official minutes of the National Central Electoral Office; the two residual mismatches fall on seats that the Chamber's Committee on Elections placed under formal investigation in 2025. On these validat
arXiv:2607.11314v1 Announce Type: new Abstract: The promise of AI literacy ``for all'' confronts a structural challenge embedded in how nations organise secondary computer science education. In most systems, a general-track subject -- Digital Literacy, ICT, TIC, or SNT -- bears the weight of universal AI literacy, while a specialist Informatics course serves STEM pathways separately. Yet the content and depth of the general track are shaped by governance decisions made largely with reference to the specialist one. This paper presents a comparative analysis of curricula and examination frameworks across fifteen countries, identifying two structural challenges. First, in several systems a significant portion of students completes secondary education without any formal programming exposure. Second, among those who do receive CS education, a \emph{Syntax Ceiling} emerges: Python-based instruction reaches most students, while the algorithmic depth associated with C++ remains concentrated in
arXiv:2607.11292v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as conversational tutors, they risk institutionalizing systemic inequalities. This study presents a systematic API audit of four LLMs acting as history tutors, evaluating 1,800 responses regarding the 1989 Romanian Revolution across five student personas varying by ethnicity and socio-economic tier. We uncover four interconnected patterns of \emph{epistemic paternalism}: (1)~\textbf{Differential Refusal}, where safety-aligned models block 76.7\% of educational requests from low-tier students; (2)~\textbf{Epistemic Gatekeeping}, evidenced by a 3$\times$ reduction in access to geopolitical complexity (e.g., the contested ``coup theory'') for marginalized learners; (3)~\textbf{Agency Theft}, a lexical shift where models like LLaMA produce a 5$\times$ higher victimization-to-politics vocabulary ratio for Roma students compared to elite peers; and (4)~\textbf{Elite Hermeneutics}, where
arXiv:2607.11228v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities. We introduce DeepBias, an adaptive framework for the in-depth probing of social biases in LVLMs with carefully designed agents. Our approach operates through a dynamic ''generation-evolution-probing'' loop. First, a generative ProposerAgent synthesizes test data and is iteratively updated via Direct Preference Optimization (DPO) based on the target LVLM's responses, exploring model-specific failure modes. Second, an autonomous skill-driven DiggerAgent rewrites each test data across multiple probing turns, adaptively selecting from a curated skill library of deepening and rew
arXiv:2607.11129v1 Announce Type: new Abstract: This innovative practice full paper presents FORAP (Framework for Organizing Reusable and Adaptable PjBL Projects) and a portfolio of 14 adoption-ready project-based learning (PjBL) project packages built with the framework. PjBL in computing education offers strong educational benefits, yet its adoption remains limited by high instructor workload and recurring student technical challenges. FORAP addresses these barriers by organizing each package around a project designed with aligned learning objectives and described through project attributes, along with coordinated instructor, student, and assessment materials that support adoption and adaptation across diverse computing courses. We report on four years of deployment across 44 classroom trials at seven universities, drawing on feedback from students, instructors, and advisory board members. Results suggest that structured project packaging supports feasible adoption with limited modif
arXiv:2607.11032v1 Announce Type: new Abstract: Project-based learning (PjBL) is common in computing education, but traditional assessments of PjBL often fail to capture higher-order thinking (HOT), especially in transfer contexts. This study introduces "design problems" (DPs): concise, scenario-based prompts that require applying project concepts in new situations, to address this gap. We examined instructor perceptions, the ability of large language models (LLMs) to generate DPs, and student experiences. Surveys of 31 instructors, evaluation of 80 LLM-generated DPs, and student performance data showed that while instructors value DPs, creation effort is a barrier. LLMs helped by producing high-quality prompts with strong expert agreement. Students rated DPs from different LLMs similarly, and their performance on DP tasks showed negligible correlation with traditional project grades, suggesting DPs may capture distinct aspects of HOT. Keystroke data also suggested deeper cognitive eng
arXiv:2607.10900v1 Announce Type: new Abstract: The transition to remote learning during the pandemic has necessitated the development of new methods for conducting hands-on experiments. One significant challenge in this transition has been providing students with reliable and sustainable access to necessary hardware components, particularly for courses that require substantial equipment. Additionally, industry partners have high expectations for students to be proficient in simulation and verification tools. To address these challenges, we implemented a virtual breadboard feature that allows students to remotely access Field Programmable Gate Array (FPGA) hardware and complete lab assignments. Our evaluation of this approach, which included surveys of students and industry partners, revealed that it effectively transformed a traditionally in-person lab assignment into an online modality. Furthermore, this paper presents the perspective of industry professionals on verification and sim
arXiv:2607.10895v1 Announce Type: new Abstract: With the transition to remote instruction, new modalities for conducting hands-on labs are needed. In particular, courses that entail major hardware components faced challenges in making the hardware available for students in a reliable and sustainable way. This paper presents our experience in implementing a virtual breadboard feature to interface with real Field Programmable Gate Arrays (FPGAs) boards located on the University of Washington's campus, where students can access the boards remotely to complete their lab assignments.
arXiv:2607.10802v1 Announce Type: new Abstract: Reinforcement learning is usually introduced through the Bellman update, yet the equation often remains abstract to undergraduates: they watch policy arrows converge but rarely observe how each value is computed or why an action is chosen. We present Q-Learning Lab, a single-file, browser-based, bilingual (Thai/English) tool for teaching tabular Q-learning that requires no installation. Beyond the usual gridworld visualization - color-coded Q-values and policy arrows on a $5 \times 5$ world - the tool exposes a live Bellman-substitution panel showing the numeric update at every step, and logs each transition, including the full pre-action Q-row, the greedy-versus-random decision under $\varepsilon$-greedy exploration, and wall-collision events, into an exportable trace. The central contribution is a learn-export-analyze loop: learners run their own agent, export the complete trace as CSV, and analyze it themselves, producing learning curv
arXiv:2607.10780v1 Announce Type: new Abstract: Modern science has experienced a long shift from individual work to team production. Generative artificial intelligence (AI) might appear to extend this trajectory by lowering research costs and enabling larger-scale collaboration. Yet if tasks once performed by coauthors can be delegated to AI, the same technology may also weaken the need for collaboration in parts of the research process. Here, we examine this tension by moving beyond average team size and focusing on the solo-authored tail of the author-count distribution. Analyzing over 300 million works across 26 fields, we find that the decades-long decline in solo authorship halted and partially reversed with ChatGPT's public release in late 2022. We also reveal that this is an uneven phenomenon: it is strongest in fields where coauthors' work is more readily replaceable, and weak or absent in fields that depend on physical collaboration. At the individual level, the recovery is no
arXiv:2607.10736v1 Announce Type: new Abstract: Artificial intelligence agents increasingly perform journalism tasks autonomously, searching for sources, evaluating credibility, and producing news content with minimal human oversight. Yet research has largely treated AI as a monolithic category, leaving the effects of architectural design unexamined. Drawing on gatekeeping theory, this study presents the first systematic comparison of four agent architectures, monolithic (Claude), chain-based (LangChain), multi-agent collaborative (CrewAI), and autonomous iterative (AutoGPT), across 200 controlled experiments spanning 50 journalism tasks of graduated difficulty. All architectures used the same underlying language model and identical tools, isolating architectural effects. Results revealed significant effects on task duration (F(3, 196) = 24.54, p < .001, eta-squared = .27) and computational strategy (F(3, 196) = 305.63, p < .001, eta-squared = .82), with architecture explaining 82% of
arXiv:2607.10620v1 Announce Type: new Abstract: The economic value of data arises from its flow across organizations and national borders. Yet increasingly stringent data governance regimes are turning cross-border transfer into an institutionally constrained sequential decision, in which firms repeatedly weigh compliance costs against the value of data flows. From the perspective of a data-exporting firm, this paper develops an institutionally anchored decision support system. It converts regulatory rules into a computable minimal compliance mapping and models the firm's weekly decisions as a finite-horizon Markov decision process (MDP), with compliance represented as a hard constraint rather than a penalty term. The resulting problem is solved using masked deep reinforcement learning, while counterfactual path advantages provide interpretable signals to support the firm's cross-border data flow decisions. Experiments show that the policies learned within the system outperform the bas
arXiv:2607.10469v1 Announce Type: new Abstract: Across a range of fast growing urban markets, private developers are constructing a version of the smart city that operates largely outside the purview of municipal government, often at the explicit invitation of city officials seeking to shift the cost and complexity of digital infrastructure onto private capital. Gated residential and mixed use developments are increasingly marketed not merely on the basis of security and amenity, but on their smartness: integrated home automation, app mediated access control, and centralized energy and resource management, among other features. We refer to this phenomenon as the smart compound. Despite its rapid proliferation, it has received comparatively little sustained scholarly attention: the literatures on smart cities and on gated communities have developed largely independently of one another, even as developers are, in practice, merging the two. This paper introduces the smart compound as an e
arXiv:2607.10013v1 Announce Type: new Abstract: In the AI Literacy for Multidisciplinary Professional Readiness and Outreach (AIM-PRO) project, we are creating integrated methods to improve the education on AI literacy. One concept on which the project relies is educational digital twins, that is, digital representations of educator trainers, teachers, and learners that can be used in different stages of the educational process. Such digital twins enable the simulation, monitoring, and optimization of learning experiences. This paper presents the AIM-PRO project and its conceptual foundations, focusing on its core objective: designing and implementing Digital Twins for Education to foster AI literacy across higher education, vocational education and training and professional learning environments.
arXiv:2607.10011v1 Announce Type: new Abstract: This paper develops a naturalistic account of religion and artificial intelligence as structurally similar distributed meaning systems. I argue that both emerge from the same underlying cognitive architecture: socially extended processes that offload interpretation, norm-guidance, and world-model construction into external symbolic environments. Drawing on work in distributed cognition, cultural evolution, and philosophy of mind, the paper proposes a conceptual model showing how meaning is generated, stabilised, and transmitted through recursive interactions between agents and their informational ecologies. Religion is analysed not as a set of beliefs but as a cognitive-ecological system that scaffolds coordination, normativity, and shared interpretation. Contemporary AI systems are shown to instantiate analogous functions, operating as high-bandwidth, algorithmically mediated environments that shape reasoning, attention, and social meani
arXiv:2607.09834v1 Announce Type: new Abstract: As human reliance on information technology (IT) increases, having a clear, precise, and comprehensive understanding of the nature of digital objects, digital systems, and digitalized systems becomes more critical. Otherwise, our ability to research, manage, control, and make reliable predictions about them will be limited. Accordingly, in this paper, we propose a new theory of digital objects, digital systems, and digitalized systems based on the adoption, adaptation, and extension of existing theories of ontology, semantics, and semiotics. The theory provides precise explanations of the nature of digital objects, digital systems, and digitalized systems. Ours is a realist theory that does not countenance the independent existence of nonmaterial or hybrid objects in the world. Accordingly, we are at odds with much of the prevailing discourse about digital phenomena. We show how our theory generates different insights and predictions from
arXiv:2607.09676v1 Announce Type: new Abstract: As high-stakes executive crisis communication shifts into fragmented digital ecosystems, unmediated narratives increasingly originate within closed-broadcast enclaves before diffusing into structurally distinct open networks. Yet the mechanisms governing cross-platform information migration remain under-theorized, underscoring the need for systematic computational analysis. This study tracks the transformation of President Trump's statements during the 2026 Iran War as they move from Truth Social into X-Twitter and Bluesky. Using a Metric-to-Semantic-Linkage framework and a high-throughput dataset of 9,891 temporally aligned public responses, we examine how divergent platform architectures shape the downstream vernacular of geopolitical conflict. The results reveal pronounced behavioral divergence driven by sociotechnical affordances. On X-Twitter, discourse is dominated by structural information bottlenecks: viral retweet cascades from e
arXiv:2607.09675v1 Announce Type: new Abstract: Radio broadcasting remains a dominant medium of communication, reaching 82% of Americans ages 12 and older weekly. Given its broad media impact, gender representation on radio news and talk stations may play an important role in shaping social and cultural perceptions. In this study, we examined patterns of gender representation in radio broadcasts, focusing on gender-based differences in total speaking time, air-time allocation across the day, and participation across broadcast topics. The dataset comprises filtered recordings from 74 US news and talk radio stations, collected over a 24-hour period and yielding more than 1,400 hours of content. We analyzed the data using VANPY, an in-house voice-analysis framework that combines multi-channel radio recording with AI-based speaker diarization, gender classification, speech-to-text transcription, and topic analysis. The results revealed consistent gender differences in allocated broadcast t
arXiv:2607.09674v1 Announce Type: new Abstract: Teachers and researchers need to adapt educational AI to local goals, but most systems remain difficult to customize or study without coding expertise. We present LearnAdapt Agentic Studio on PedOS 1.1 Lumina, a no-code authoring and governed runtime environment for educational AI plugins. A non-coder describes a desired learning interaction in plain English; the system prepares a previewable plugin artifact, runs safety checks, and supports submission for review. PedOS then deploys approved plugins into a directory for installation. Crucially, telemetry is strictly gated to authenticated users running approved plugins. The demo shows the complete lifecycle from prompt to governed evidence capture, shifting from fixed tools to teacher-built teammates.
arXiv:2607.09673v1 Announce Type: new Abstract: Critique is central to design education and project-based learning, yet high-quality critique is often scarce, uneven, hard to document, and disconnected from evolving artefacts. We present Critsly, an artefact-aware AI critique workspace that turns AI from a detached feedback tool into a critique teammate in the learner's board context. Critsly combines a visual design canvas with structured AI-supported reflection, multi-perspective critique, action planning, optional peer/jury settings, and educator evidence traces. Unlike chatbot feedback tools that rely on isolated text prompts, Critsly grounds critique in a structured board state containing design intentions, board elements, annotations, links, and prior critique history. Reflecture, Critsly's guided reflection flow, works with Six Thinking Hats-inspired personas, board synthesis, generated action plans, exportable critique records, and educator evidence views in one workflow. The d
Article URL: https://cyrilzakka.github.io/radiology/2020/10/13/mammogenesis.html Comments URL: https://news.ycombinator.com/item?id=24763701 Points: 5 # Comments: 0
Conversations with Kevin Hogan: Karl Rectanus brings his edtech evidence background to the nation's original science of reading organization — and is betting on outcomes-based contracting to close the literacy gap.
NOCD, a virtual provider for OCD, announced plans to expand its virtual intensive outpatient program for people with severe OCD nationwide. The post NOCD Announces Plans for National Expansion of Intensive Outpatient Therapy appeared first on MedCity News .
Immunology startup Infinimmune is developing antibody drugs with potential advantages over traditional monoclonal antibodies. The Series A financing will support two lead programs entering the clinic in atopic dermatitis. The post Infinimmune Secures $75M for a More Human Approach to Developing Better Antibody Drugs appeared first on MedCity News .
Data from a recent Comparitech report shows a 9% decline in the number of confirmed ransomware attacks on U.S. educational institutions from 2024 to 2025. While incidents may not be rising as fast as they once were, and though average ransom amounts are down 33%, the overall rate of attacks and the severity of breaches is still high. The bottom line is that cyberattacks of any kind aren't going anywhere, and that’s leading schools to shift from emergency response mode toward a longer-term cyber resilience strategy. Brock Boggs, IT director of Cityscape Schools in Dallas, Texas, says that his…
Tolerance for this abrasion is gone. It wastes time, drains staff capacity, drives up costs, and affects the patient experience. The post From Friction to Fix: Measuring What’s Breaking Payer-Provider Trust appeared first on MedCity News .
[Sponsored] A new report from Forrester offers a score card assessing health tech vendors on customer experience for healthcare organizations considering adopting them. The post Which Health Tech Companies Are Getting Customer Experience Right? appeared first on MedCity News .
Tammy Wincup, CEO of Securly, discusses the state of the edtech market and how edtech CEOs can effectively lead in the current climate.
arXiv:2608.06564v2 Announce Type: replace-cross Abstract: Quantization is how large language models are actually deployed, and below four bits it hurts. What nobody can say is which decisions change at a given bit-width -- which matters most where a model acts rather than answers, since a tool call it declines to make is a failure no score reports. A compressed agent stops calling its tools, then loses half its safety refusals, while benchmark scores barely move. Prior work assumes the added noise has a roughly fixed size, which would make confident decisions safe. We measure the decision instead: the margin between the option a model picks and its best alternative, before and after quantization, across 16 models, three methods, and 8 down to 2 bits. Kinds of decision do not break together -- at 3 bits the decision to call a tool collapses toward inaction while the choice of which tool is untouched -- and the damage is proportional rather than fixed, the margin multiplied by a factor t
arXiv:2608.00267v2 Announce Type: replace-cross Abstract: Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing benchmarks often center on localized tasks or end-state outcomes, offering limited insight into sustained execution. We introduce LOOPSBENCH, a long-horizon benchmark for loop engineering in coding agent evaluation. Each task is a dependency DAG over separately testable development units with source-evidenced prerequisite edges. LOOPSBENCH comprises 112 tasks from authentic sources spanning 8 programming languages and 9 domains. Its flow-aware runtime releases tests along the ready frontier and retains completed nodes as regression obligations. We evaluate frontier coding agents paired with widely used loop implementations. The strongest configuration, Opus-4.7 with Claude Code and outer continuation, resolves 25.00% of tasks. Recorded plans recover only
arXiv:2606.21678v2 Announce Type: replace-cross Abstract: Language models can generate plausible rationales for their predictions, but these explanations may not faithfully represent the model's internal reasoning. We propose verifier-coupled reasoning, a framework that inserts inline claims into reasoning traces and trains an auxiliary consistency head to predict programmatic verifier outputs from rationale-span hidden states. The central finding is a gap between decodability and faithfulness: consistency training reliably makes verifier information decodable from rationale representations, but decodability does not guarantee faithful generation. In LeanCheck (formal theorem proving), rationale-only and proof-only pooling achieve perfect directional separation under counterfactual conflict. In KataGo (Go engine), commentary spans encode 10-way win-rate buckets at 81% accuracy. Yet in a code setting, the model achieves 98.6% coupling while its generated explanations remain unfaithful:
arXiv:2606.21597v2 Announce Type: replace-cross Abstract: The open-source ecosystem on GitHub lacks a systematic hierarchical taxonomy of software repositories. GitHub Topics, the dominant organizational mechanism, is flat, inconsistent, and covers only 67% of projects. We present ATLAS, the first framework that automatically constructs a hierarchical taxonomy for software repositories and classifies projects into it end-to-end. By combining LLM global knowledge with real repository distributions, ATLAS proposes meaningful splitting dimensions and iteratively corrects those that fail to accommodate real projects. A Designer Agent proposes splitting dimensions while a Classifier Agent assigns repositories; a self-corrective refinement loop uses classification failures to drive dimension revision through escalating strategies. We evaluate ATLAS on 54,387 GitHub repositories against six baselines spanning four paradigms, two downstream tasks, and three model families. On a stratified 2,00
arXiv:2606.00376v2 Announce Type: replace-cross Abstract: Extended chain-of-thought reasoning can degrade performance on deterministic state-tracking tasks, not solely because of preference biases but, on the evidence we present, because of information-theoretic limits in the capacity of decoder-only attention. We present: (1) an Attention Bottleneck analysis providing evidence that total state-tracking capacity in bits is bounded in terms of head count, head dimension, and context length under stated modeling assumptions, and that total capacity is not the binding constraint; (2) a context-dependent error model with a depth-dependent quadratic term in the error exponent; (3) the State-Space Jaccard metric measuring state drift; and (4) a Deterministic Horizon $d^* \in [19, 31]$ (at $\alpha = 0.5$) marking the depth at which unaided accuracy crosses 50%. Across twelve models and eight task domains (including SWE-Bench, WebArena, and SQL-Multi), tool-integrated reasoning reaches 76-94%
arXiv:2605.15532v3 Announce Type: replace-cross Abstract: Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically chosen via simple heuristics or aggregated from off-the-shelf datasets. We reveal a critical inefficiency in this approach: up to 69% of the prompts in standard chart / document reasoning datasets are effectively zero-delta, meaning the teacher and student already induce the exact same answer distribution. Training on these prompts provides minimal learning signal, causing student improvement to rapidly saturate regardless of data scale. To escape the zero-delta trap, we return to first principles: distillation fundamentally minimizes distributional divergence, and thus a prompt is valuable only if it exposes a functional capability gap between the teacher and student. We quantify this gap through answer divergence ($\Delta$), demonstrating that non-zero divergence is critical for
arXiv:2604.25800v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) has been shown to empirically improve Transformers' performance, and theoretically increase their expressivity to Turing completeness. However, whether Transformers can learn to generalize to CoT traces longer than those seen during training is understudied. We use recent theoretical frameworks for Transformer length generalization and find that -- under standard positional encodings and a finite alphabet -- Transformers with CoT cannot solve problems beyond $TC^0$, i.e. the expressivity benefits do not hold under the stricter requirement of length-generalizable learnability. However, if we allow the vocabulary to grow with problem size, we attain a length-generalizable simulation of Turing machines where the CoT trace length is linear in the simulated runtime up to a constant. Our construction overcomes two core obstacles to reliable length generalization: repeated copying and last-occurrence retrieval. W
arXiv:2604.23333v2 Announce Type: replace-cross Abstract: Scaling test-time computation with reinforcement learning (RL) has emerged as a reliable path to improve large language models (LLM) reasoning ability. Yet, outcome-based reward often incentivizes models to be overconfident, leading to hallucinations, unreliable confidence-based control, and unnecessary compute allocation. We introduce Reinforcement Learning with Confidence Margin (RLCM), a calibration-aware RL framework that jointly optimizes correctness and confidence reliability via a margin-enhanced process reward over intermediate-budget completions. Rather than aligning confidence to correctness likelihoods, RLCM encourages the model to widen the confidence margin between correct and incorrect steps within a single reasoning trajectory. Across mathematical, code, logic and science benchmarks, our method substantially improves calibration while maintaining or improving accuracy. We further show that, with calibrated confide
arXiv:2604.10015v3 Announce Type: replace-cross Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks. While existing benchmarks have begun evaluating financial tool calling, they focus on limited scenarios and rely on call-level metrics that fail to capture trajectory-level reasoning quality. To address this gap, we introduce FinTrace, a benchmark comprising 800 expert-annotated trajectories spanning 34 real-world financial task categories across multiple difficulty levels. FinTrace employs a rubric-based evaluation protocol with nine metrics organized along four axes -- action correctness, execution efficiency, process quality, and output quality -- enabling fine-grained assessment of LLM tool-calling behavior. Our evaluation of 13 LLMs reveals that while frontier models achieve strong tool selection, all models struggle with information utilization and final answe
arXiv:2604.08377v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static after deployment. As a result, similar workflows, tool usage patterns, and failure modes are repeatedly rediscovered across users, preventing the system from improving with experience. While interactions from different users provide complementary signals about when a skill works or fails, existing systems lack a mechanism to convert such heterogeneous experiences into reliable skill updates. To address these issues, we present SkillClaw, a framework for collective skill evolution in multi-user agent ecosystems, which treats cross-user and over-time interactions as the primary signal for improving skills. SkillClaw continuously aggregates trajectories generated during use and processes them with an autonomous evolver, which identifies recurring behavioral patterns and translates them into upd
arXiv:2604.07650v2 Announce Type: replace-cross Abstract: The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Shared pretraining data, distillation, and alignment pipelines can induce hidden behavioral dependencies, or latent entanglement, that undermine multi-model systems such as LLM-as-a-judge pipelines and ensemble verification, which implicitly assume independent signals. In practice, this manifests as correlated reasoning patterns and synchronized failures, where apparent agreement reflects shared error modes rather than independent validation. To address this, we develop a statistical framework for auditing behavioral entanglement among black-box LLMs. Our approach introduces a multi-resolution hierarchy that characterizes the joint failure manifold through two information-theoretic metrics: (i) a Difficulty-Weighted Behavioral Entanglement Index (BEI), which amplifies synchronized failures on e