Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2508.05952v2 Announce Type: replace Abstract: Large language model (LLM) tutors are increasingly used to generate educational feedback, but existing research has focused mainly on feedback generation rather than feedback evaluation. As a result, LLM-generated feedback may offer limited pedagogical value and carry risks of hallucination. The current study introduces DeanLLM, an automated review framework for comprehensively evaluating feedback generated by LLM tutors before it is shared with students. We developed a 16-dimension evaluation framework covering feedback content, educational effectiveness, and hallucination risks, and validated it using using human-expert annotations of LLM-generated tutor feedback on synthetic computer science assignment submissions derived from real coursework. We then examined whether LLMs could serve as automated LLM-generated tutor feedback reviewers, and used the best-performing reviewer to benchmark tutor feedback generated by 10 commercial LLM
arXiv:2607.02432v1 Announce Type: cross Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation. This paper evaluates whether four frontier Large Language Models (GPT, Claude Opus, Gemini, and GLM) can approximate expert judgment when grading short Linux/bash command responses. The study adopts a four-level cognitive taxonomy that combines cognitive complexity and operational impact, ranging from information retrieval (L1) and basic file manipulation (L2) to structural operations (L3) and advanced system management (L4). The models were tested with two prompt variants, a minimal baseline and a rubric-enhanced version, on 1200 real responses from second-year Computer Engineering students independently graded by three expert instructors. Gemini~3.0 Pro with rubric-guided prompti
arXiv:2607.02245v1 Announce Type: cross Abstract: Mental health disorders affect nearly one billion people globally, yet 75% of individuals in low- and middle-income countries receive no treatment due to workforce shortages, cost barriers, and stigma. Current AI-powered wellness solutions predominantly rely on single-mode conversational interfaces that suffer high abandonment rates and fail to provide measurable, immediate relief calibrated to users' dynamic emotional states. This paper presents Copewell, a novel multi-agent swarm system designed to expand access to mental wellness support through human-centered AI principles. Our architecture introduces three technical innovations: (1) a multi-source assessment framework integrating self-reported, physiological, and contextual data to mitigate algorithmic bias; (2) valence-arousal emotion mapping using Russell's Circumplex Model of Affect to route users to specialized AI agents; and (3) dual-mode intervention delivery combining conver
arXiv:2607.02181v1 Announce Type: cross Abstract: Americans' warmth toward members of the opposing political party has fallen sharply over the past three decades -- yet meaningful cross-partisan contact remains scarce, in part because people actively avoid it. Across five preregistered studies (total N = 3,960 U.S. partisans), we test whether brief conversations with AI chatbots representing the political outgroup can substitute for the contact people shun. Synthetic contact first lowers the barrier to entry: partisans would endure almost twice as long contemplating their own mortality to avoid a human outgroup partner as an AI one. These conversations then correct the misperceptions that fuel division. At baseline, Democrats placed Republicans more than a standard deviation past their actual position on environmental consumption attitudes -- enough to flip the average Republican from supportive to opposed -- and a single ten-minute conversation with an outgroup chatbot corrected those
arXiv:2607.02049v1 Announce Type: cross Abstract: Large Language Models are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-lingual abilities in these circumstances remain underexplored. Existing benchmarks emphasize multilingual performance but rarely examine crisis-related empathy and cultural grounding in low-to-mid-resource languages. We introduce SPLIT, a 500-prompt benchmark designed to evaluate LLM consistency in generating emotionally grounded responses across five categories: Stress, Panic, Loneliness, Internal Displacement, and Tension. We evaluate three technically diverse LLMs across three dimensions: Empathetic Accuracy, Linguistic Naturalness, and Contextual & Cultural Grounding. The framework aims to assess and compare the quality of LLM responses in both English and Ukrainian languages, as well as to explore the reliability of the LLM-as-a-jury paradigm. Our findings reveal that Gemini-2.5-Flash and LLaMA-3.3-
arXiv:2607.01833v1 Announce Type: cross Abstract: The global development of Library and Information Science (LIS) is influenced by various factors such as the economy, society, culture, discipline, tradition, and more. Consequently, the research methods of LIS vary greatly among countries. To better understand these differences, we conducted a study of 5,281 research papers from 81 countries published in internationally representative journals over the past thirty years. We manually annotated the research methods used in some articles through content analysis, and subsequently developed and trained a deep learning model for automatic classification of research methods. Using this method, we conducted a comparative analysis of the usage of research methods in different countries. Our findings reveal that there are differences in the research methods used across countries, with each country having its unique research profile and distribution of research methods. Even when investigating t
arXiv:2607.01828v1 Announce Type: cross Abstract: Research in the social sciences has shown that there are gender differences in the selection of research methods, with women often opting for qualitative methods while men prefer quantitative methods. However, it is important to consider that research methods are generally chosen based on the research topic. To figure out the influence of gender on research method selection, a study was conducted in the field of Library and Information Science, using a more fine-grained method classification system and an automatic classification model called CogFT, which is based on full-text cognition. The findings showed that women tend to use Interview while men prefer Theoretical approach, across a range of topics. The study offers insights into the specific research design processes that contribute to gender differences in method selection and suggests ways to promoting gender inclusivity and equality in academia by considering research method use
arXiv:2607.01730v1 Announce Type: cross Abstract: Liquid democracy promises to improve collective decision-making by allowing voters to vote directly, delegate their voting power to trusted participants, or combine both approaches through fallback mechanisms. However, existing deployments typically rely on transparent delegation, which exposes voters to popularity-driven herding, makes coercion verifiable, and introduces systemic fragility when highly-backed delegates abstain. In this paper, we propose a secure liquid democracy mechanism that resolves the tension between informed expertise routing and systemic robustness. We introduce a sealed delegation regime using decentralized timed-release encryption, which cryptographically hides delegation choices during the formation phase to prevent herding and coercion, while restoring full public auditability for the final tally. To address delegate failures, we extend the protocol with ranked multi-delegation and personal fallback ballots.
arXiv:2607.01245v1 Announce Type: cross Abstract: We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.docx, .xlsx, .pptx) and their variants. OCB consists of two tracks. File Fidelity Q&A tests structural and visual perception of office artifacts - tables, charts, embedded images, formulas, and app-specific elements such as headers, speaker notes, and named ranges. Domain Q&A tests expert-level reasoning grounded in real-world industry documents across 12 professional domains, with queries requiring multi-step analysis and synthesis across documents. Each reference answer is decomposed into atomic, binary-gradable claims, and an ensemble of LLM judges scores responses against each claim independently. Even the strongest frontier system in its default reasoning mode reaches only about 59.3% on Domain Q&A increasing thinking depth within a tier does not move perfo
arXiv:2607.01244v1 Announce Type: cross Abstract: The growing number and complexity of technical regulations represent an important challenge for all professionals in regulated industries. This paper describes a case study, from design to deployment, of building a Retrieval-Augmented Generation system for the consultation of complex technical regulations in the railway domain. Although developed for the railway sector, this testimony of an industrial experience is of particular value for technical domains where regulatory compliance and accurate information retrieval from complex documentation are essential requirements. It also constitutes a human-centered approach for implementing LLM-powered technical documentation consultation across various regulated industries, balancing technological capabilities with domain expertise.
arXiv:2607.02467v1 Announce Type: new Abstract: Whether pairing people with AI helps or hurts is usually reported as a single average effect. Using a real-money prediction market (Polymarket) as an objective, externally resolved benchmark, this pilot shows that the value of human-AI collaboration depends on a specific, measurable form of human capital. Analyzed at the level of the individual forecaster, hybrid performance is trimodal: most people either deferred to the model (matching it) or used it to rubber-stamp a prior guess (performing worse than the model alone), while a minority engaged in genuine complementary reasoning and reached accuracy matching or even exceeding (i.e., lower error than) the market itself. Collaborative traits (perspective-taking, intellectual humility, and curiosity) rather than raw cognitive ability or model benchmarks, distinguished who reached that mode. The results are preliminary but statistically robust, and motivate a pre-registered replication now
arXiv:2607.02313v1 Announce Type: new Abstract: As conversational AI systems become more deeply integrated into daily life, the implications for human agency are increasingly urgent to understand. AI's potential to amplify capability sits alongside risks of individual and collective disempowerment, yet empirical, ecologically-valid evidence about cumulative usage is scarce. We analyze deep ethnographic data from a study of daily AI chatbot users (n = 51) in the United States, Germany, and Singapore to illuminate conversational AI usage in situated context as a sociotechnical practice. We show that people consistently link sustained AI usage to perceived gains in individual agency. Crucially, these perceived gains often outweigh concerns about accuracy, reliability, and consistency to shape usage patterns. Our findings challenge prevailing assumptions about how and why humans use AI systems over time, suggesting that traditional trust-based models are not sufficient for explaining human
arXiv:2607.02201v1 Announce Type: new Abstract: The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains fragmented across competing risk taxonomies that catalog risks without showing how an audit is executed. At least 74 AI risk taxonomies exist, and almost all stop at the catalog. The hard part of auditing is not naming a risk but operationalizing it: turning it into a test run against a real system, a measured value, a calibrated severity, and a defensible grade. This paper leads with that bridge. We present the operationalization layer Eticas has built and run, shown end to end on a single risk (PII leakage) against a public benchmark, and then the open taxonomy that makes the method scale. On GPT-4-0314, a disclosure risk that seven external frameworks require be controlled is measured at 0%, 51%, and 84% disclosure as adversarial conditioning increases, mapping through calibrated severity bands to a
arXiv:2607.02197v1 Announce Type: new Abstract: The society and emerging risk-based regulatory frameworks for AI underscore the need for rigorous risk assessment to ensure safe and reliable AI systems. In response to this imperative, this paper presents an overview of AI risk assessment (identification and analysis) and management methodologies. It begins by reviewing the worldwide regulatory landscape that drives the need for systematic AI risk assessment. Then we characterize the spectrum of AI-related risks identified in the literature, from technical failures to ethical and social impacts. Subsequently, it reviews key risk assessment methodologies proposed for AI systems, focusing on general frameworks. The paper highlights best practices and illuminates methodological gaps, highlighting areas for further research on AI risk assessment.
arXiv:2607.02144v1 Announce Type: new Abstract: While AI promises major benefits, its development and deployment can shift costs onto others, including environmental pressures on local communities, labor and creative displacement, and systemic risks from rapid frontier development. Taxation is an integral part of policy design, and recent academic, industry, and policy debates have begun to consider whether tax instruments can help address these harms. In this paper, we explore the viability of AI taxation. More broadly, AI taxation should not be understood only as Pigouvian correction. In the AI context, taxation can also correct harmful activity, redistribute unevenly borne costs and gains, and fund regulatory capacity. We discuss the main externalities associated with AI and survey possible tax instruments, including corporate income and rent-based taxes, consumption taxes on AI-related services, and excise instruments tied to specific AI activities. We further assess the benefits a
arXiv:2607.01913v1 Announce Type: new Abstract: Organizations increasingly make strategic decisions about AI systems whose behaviour, failure modes, and institutional effects cannot be fully known at design time. This technical report reframes strategic red teaming as a board-level governance discipline for testing the assumptions under which AI-enabled strategies are approved, funded, and supervised. The report proposes a six-component model for strategic red teaming in AI governance: an explicit assumption register, an adversarial mandate, independence criteria, evidence grading, a board-facing decision record, and a follow-up mechanism for unresolved findings. The model is intended to make strategic uncertainty inspectable before it becomes operational exposure. It treats red teaming not as penetration testing, scenario theatre, or generic risk review, but as structured adversarial testing of the claims on which governance decisions depend. The contribution is conceptual and design-
arXiv:2607.01776v1 Announce Type: new Abstract: In the age of AI, what will be good knowledge? This article, which is accepted and forthcoming in a special issue of Modern Fiction Studies on "Cultural AI" in 2027, applies digital humanities methods to map epistemic virtues (like "true," "accurate," "creative") used in a corpus of 553 journal articles on AI published in 2024. "Creativity" comes in for special attention as an example. Exploring this discourse of value, the article considers how a framework might be developed for evaluating the knowledge-worth of AI -- one less locked into values formed around pre-AI "knowledge work" agents or structures, and more open to the future values of "generativity." The essay is supported by an online digital kit for exploring data models of the corpus of articles on AI it studies.
arXiv:2607.01750v1 Announce Type: new Abstract: Open source software (OSS) is not homogeneous. A project's purpose, governance, and funding shape how its community forms, who contributes, and how the software is maintained, yet empirical research often samples OSS broadly and reports findings as if they held for open source as a whole. We argue that OSS comprises distinguishable sub-genres, and that the sub-genre a study samples bounds how far its findings generalize. Using a light, multi-source review that screens 3,925 unique papers, we synthesize a typology of fourteen OSS sub-genres, from well-studied ones such as community-driven, company-backed, foundation-governed, research and scientific, and open source for social good (OSS4SG), to under-studied ones such as multi-company co-opetition, protestware, and open-source appropriate technology. We place the sub-genres in a framework that records each one's primary driver, governance, and funding, with its maturity in the literature a
arXiv:2607.01460v1 Announce Type: new Abstract: Human-annotated data remains foundational for machine learning and social media analysis. However, traditional data collection often relies on cumbersome pipelines that isolate content from its original source, compromising ecological validity. To address these challenges, we present Social-Annotate, a flexible browser extension that facilitates direct data collection on online platforms. By injecting customizable forms into webpages, the tool captures annotations while users interact with the native environment. Social-Annotate offers no-code design interface for the survey forms for non-technical users. Since injecting custom elements directly into host platforms creates a brittle dependency on evolving interfaces, we integrate a self-healing agent powered by large language models. This automated pipeline autonomously detects structural changes, regenerates valid target selectors, and validates them within a live browser environment. Ou
arXiv:2607.01258v1 Announce Type: new Abstract: The rapid advancement of Artificial Intelligence (AI) has been accompanied by significant increases in computational and environmental costs, driven by large-scale investments in AI infrastructure, hardware, and software. In particular, graphics cards have become central to AI training, with frequent hardware updates required to meet escalating computational demands. However, the environmental damages of graphics cards production remain understudied. This study addresses this gap by estimating the environmental damages associated with graphics cards production over the past decade (2013-2025). We analyze trends in energy consumption, carbon emissions and resource depletion. We compile and provide a dataset documenting the environmental damages of NVIDIA workstation graphics cards production since 2013. Our analysis of this dataset reveals a steady increase in production-related impacts over the period. Our finding highlights the need for
arXiv:2607.01257v1 Announce Type: new Abstract: The rapid digitalisation of financial systems has improved operational efficiency and financial inclusion while simultaneously increasing exposure to sophisticated forms of cyber-enabled fraud and electronic financial misconduct. Conventional auditing systems, which largely depend on retrospective verification and rule-based monitoring, increasingly struggle to address the complexity and speed of modern financial crime. Consequently, financial institutions are progressively adopting Artificial Intelligence (AI)-enabled Accounting Information Systems (AIS) and Natural Language Processing (NLP) technologies to strengthen fraud detection, continuous auditing, and institutional monitoring. This study examined the influence of AI-enabled AIS on auditing and fraud detection effectiveness within Nigeria's financial services sector while additionally evaluating the moderating role of NLP. Anchored on the Fraud Diamond Theory and the Technology Ac
arXiv:2607.01256v1 Announce Type: new Abstract: Overwhelmed courts in the United States review millions of default judgments each year. Unfortunately, such manual reviews are time-consuming and prone to error. In an audit of 188 debt collection cases granted default judgment by the Superior Court of Los Angeles, we find that 4% contained major defects that should have entirely prevented default judgment, 10% contained inconsistencies requiring reduced judgments, and 32% contained errors requiring amendment prior to judgment. To support courthouses in default judgment review, we collaborated with courthouse attorneys and judges in designing a Default Assistant. The Default Assistant employs large language models to evaluate a case with respect to predetermined legal requirements and provide cited recommendations for an expert user's review. We equip users to verify these recommendations by grounding the assistant's explanations in cited quotes and tables from the original case filings.
arXiv:2607.01255v1 Announce Type: new Abstract: Universities have responded to generative artificial intelligence (GenAI) in noticeably different ways, both internationally and within Spain. So far, the dominant reaction has been defensive, this is, most institutions frame the debate around AI detection, plagiarism, academic integrity and a presumed drop in student effort, prioritizing basic training for academic staff over students. Other group of pioneering universities is doing the opposite, pursuing deeper adoption, and assuming that any policy built on prevention or sanction will not hold. This paper sides with that second view. Obsessing about detection is a dead end, since generated text is increasingly hard to distinguish from human writing, and detectors still misfire too often to be trusted. What universities need instead is a coordinated effort to set clear, course-by-course rules for GenAI use, redesign assessment toward authentic and interdisciplinary assessment that foste
arXiv:2607.01254v1 Announce Type: new Abstract: Benchmarks are the primary instruments through which AI capability is measured, compared, and governed. This paper argues that the validity of frontier AI benchmarks is a function of the quality of human judgment embedded in their construction, and that this quality is structurally scarce in ways that standard scaling narratives obscure. As foundation models approach ceiling performance on existing evaluation suites, discriminating signal concentrates in the hardest benchmark items, precisely those requiring elite expert judgment to design. We term this the benchmark ceiling problem: the progressive exhaustion of evaluation signal as models saturate the easy majority of items while the difficult tail, authored by a thin stratum of highly expert evaluators, remains the only source of genuine discrimination. The paper develops this argument in three steps. First, we present a formal model of benchmark signal depreciation. Benchmark scores a
arXiv:2607.01253v1 Announce Type: new Abstract: Rationale. The diagnostic radiologist's role in 2035 will not look like it does today. Imaging AI is already changing how worklists are organized, how reports are generated, and which cases require a radiologist's attention. What remains genuinely contested is not whether the role changes but how. Approach. Three subject-matter experts (two radiologists and one health tech professional with more than 20 years of experience in medical imaging IT) independently authored 2035 job descriptions for the diagnostic radiologist using a shared template. Each author wrote from a distinct vantage point: one optimistic, one framed as a trade-off view incorporating workforce economics, and one structured around professional stratification. The three versions were published openly and subjected to a structured comparison across seven dimensions. Key findings. The three versions agree on direction but disagree on magnitude. All three describe a radiolog
arXiv:2607.01252v1 Announce Type: new Abstract: Background: Dermatology AI has mainly focused on image-based diagnosis, while chronic disease workflows have received less attention. We surveyed Indian dermatologists to map routine clinical challenges, with a focus on atopic dermatitis (AD), and assess current AI use. Methods: A nationwide cross-sectional survey commissioned by the Society for Eczema Studies included 377 practicing Indian dermatologists. The survey assessed clinical challenges, AD workflow barriers, AI use, adoption barriers, and ethical concerns. Analyses used descriptive statistics, chi-square tests, false discovery rate correction, and multivariable logistic regression. Results: Patient adherence (61.3%) and treatment planning in difficult or refractory cases (57.0%) were reported more often than diagnostic uncertainty (48.0%). In AD care, severity scoring was reported as a challenge by 47.7% and had the lowest satisfaction among measured workflow areas. Current AI u
arXiv:2607.01251v1 Announce Type: new Abstract: Debate, where AI agents argue opposing positions, has emerged as a key approach to scalable oversight. However, debate faces a fundamental tension: models are incentivized to be persuasive to the judge, which may not always align with epistemic honesty. In this work, we propose an alternative paradigm: disagreement resolution, which reframes the interaction mechanism from adversarial debate to collaborative truth seeking. Drawing on principles from human mediation and conflict resolution, where mediators facilitate dialogue to help disputing parties reach consensus rather than adjudicating between them, we design an automated pipeline that adapts these strategies to AI oversight. Unlike standard debate where models argue for fixed positions, our pipeline directs models to collaboratively identify points of disagreement, examine the evidence for conflicting claims, and converge toward consensus or isolate the specific ''crux'' of their dis
arXiv:2607.01250v1 Announce Type: new Abstract: Sociotechnical alignment concerns the social desirability of AI behavior and is thus inherently normative, not merely technical. While NLP research increasingly addresses its technical aspects, it often leaves underspecified what such "social desirability" entails. We argue that this reflects a fundamental gap: the absence of a systematic way to specify how sociotechnical alignment defines, justifies, and evaluates socially desirable AI behavior. To address this gap, we introduce a human-centered framework for specifying sociotechnical alignment. We draw on social-scientific accounts of sociobehavioral desirability to ground the basis for behavioral desirability judgments and use this framework to analyze how alignment is specified in practice. Our systematic literature review identifies recurring patterns: normative concepts grounding desirability judgments are often unspecified or conflated with alignment targets for (desired) system be
arXiv:2607.01248v1 Announce Type: new Abstract: Large language models are increasingly used for knowledge acquisition, code generation, academic writing, and agent-based automation. In these settings, users may obtain highly structured answers, plans, and judgments without sufficient domain practice. This paper proposes a practice auditing framework for LLM use and AI-generated content governance. It introduces collective empiricism to describe how LLMs compress and reorganize large-scale human experience into outputs that appear empirical and rational, and pseudo-rational cognition to describe how users may mistake AI-generated structured expression for their own rational understanding. The paper analyzes AI subjectivity illusion, subjectivity structures in input materials, template loops in AI-AI conversations, statistical misjudgment in AIGC detection, and memory pollution when generated content enters future contexts, long-term memory, retrieval spaces, or agent skill systems. To r
arXiv:2607.01247v1 Announce Type: new Abstract: Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of intermediate steps. They are also difficult to grade at scale because instructors must apply partial-credit rubrics consistently while giving feedback that helps students repair misconceptions. This paper evaluates six contemporary large language model (LLM) configurations, Gemini 3.1 Pro Extended, Gemini 3.5 Flash, ChatGPT 5.5 Pro Extended, ChatGPT 5.5 Thinking, Claude Pro Opus 4.7, and Claude Sonnet 4.6, as grading assistants for an undergraduate discrete mathematics examination. The study compares two grading policies. The BASELINE policy uses a stricter rubric-following prompt that emphasizes explicit evidence and complete justification. The LIBERAL policy was added after preliminary grading showed that the baseline condition sometimes applied harsh point deductions and failed to recognize valid parti
arXiv:2607.01246v1 Announce Type: new Abstract: Recent proposals for HTTP-based sustainability disclosure focus on \textbf{what} environmental information should be transmitted at the protocol boundary, for example through response headers, but leave open the practical question of \textbf{how} such per-request values can be generated in realistic deployments. This paper addresses that implementation gap. We present a model-based approach for estimating resource consumption and $CO_2e$ per HTTP request without requiring fine-grained production power telemetry. The approach benchmarks endpoints offline under controlled conditions, derives compact endpoint-specific energy models from observable request features, and evaluates these models online at the HTTP server boundary. We implement this mechanism as an nginx extension that loads a JSON model registry and emits per-request metadata for energy, grid intensity, embodied emissions, and total request-level impact. We show that heterogeneo
New edtech products that have caught our attention this month
Large language models (LLMs) are entering clinical training as digital standardized patients (DSPs), simulated patient encounters the model scripts and portrays. Demographic associations learned from corpus co-occurrences can enter at two points: the written case scripts and the live portrayals improvised in role-plays. We audited both pathways using an HIV pre-exposure prophylaxis (PrEP) screening scenario across six demographic factors (age, gender, marital status, sexual orientation, education, and race or ethnicity) in a 216-cell factorial design. With a generated arm and a template-substituted control arm, we simulated 4,320 conversations with fixed audit questions, measuring three channels: the case scripts, the composite role-play trainees receive, and role-play under identical control cases. Differences concentrated where corpus associations were strongest and reinforced predictable stereotypes. Anal sex appeared in 100% of gay mens cases, 83% of bisexual mens and 3% of heteros
Background Gambling harm is a significant public health concern that is systematically under-recognised in clinical practice. Despite the recent inclusion of gambling disorder in the General Medical Council's Medical Licensing Assessment content map, gambling harm has been largely absent from undergraduate medical education in the United Kingdom, and structured evaluations of gambling harm teaching delivered to medical students have not, to our knowledge, been reported. Methods A single-group, pre-post study of a teaching session on gambling harm was conducted across three sequential cohorts of Year 4 medical students at King's College London during the 2025-2026 academic year. Outcome measures were collected immediately after the teaching with no follow-up. The session was delivered online by Gambling Harm UK, a registered UK charity, and comprised lived experience testimony and teaching with public health and clinical components. Self-reported confidence across six domains was assess
Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outc
Background. A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification, requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners. Objective. To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it. Methods. In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss' kappa, Krippendorff's alpha, Gwet's AC1, and ICC(2,1)/ICC(2,k). Results. Human item agreement was moderate (kappa 0.409 to 0.467), as was
Background: Exclusive breastfeeding may protect infants against common infections and support healthy growth and development. Working mothers may face constraints on exclusive breastfeeding arising from work schedules, separation from their infants, and inadequate breastfeeding support. National evidence on the individual, healthcare-related, and contextual factors associated with exclusive breastfeeding among working Ghanaian mothers appears to remain limited. Design: Cross-sectional secondary analysis. Setting: Nationally representative survey covering urban and rural communities across all 16 administrative regions of Ghana. Participants: The analysis included 620 currently working mothers whose youngest living infants were aged 0-5 completed months and lived with them. The complete-case multivariable analysis included 619 mother-infant pairs. Primary outcome measure: Current exclusive breastfeeding, defined using the standard 24-hour infant-feeding indicator. Infants were classifie
Evidence-based medicine (EBM) concepts are difficult for medical students to grasp. We developed a Python-Streamlit web application providing interactive visualizations to enhance EBM education. Preliminary use with first year medical students demonstrated high engagement and improved conceptual understanding, supporting the feasibility of integrating interactive, web-based tools into EBM curricula.
Objectives: This study aimed to examine physiotherapy students reaction and learning, following Simulation Based Education (SBE) during a Cardiorespiratory module. A further aim was to see if any learning translated into clinical placement. Design: A mixed methods research design consisting of a questionnaire (Phase1) after the activity which was underpinned by the Kirkpatrick model of evaluation. This was followed by a focus groups (Phase 2) after completion of clinical placement. Participants: 92 final year physiotherapy students at a single institution were eligible to participate in the SBE session with n=81 (88%) students completing the survey, and 8 students participating across 2 focus group sessions. Results: Survey: Over 80% of students strongly agreed on a positive initial reaction to SBE. Learning yielded a 76% and above response of strongly agree in all areas except confidence. Students valued SBE as a preferred learning and teaching strategy and wanted more. They welcomed
Within our synchronous, online global health-focused upper-level microbiology course, we find that students struggle to translate learning to real-world applications. For examples, consider the recent measles outbreaks and major events such as the COVID-19 pandemic, which prompt many questions about how basic microbiological information is used by public health professionals. To address these points, we created a simulation-based curriculum that places students in an action role during an infectious disease outbreak. Our stand-alone curriculum walks through a historical measles outbreak that introduces outbreak investigation, community communication, and how these depend on microbiological knowledge. By using breakout groups, students have an opportunity to decide classifications, public messaging, and intervention metrics. We provide students with an outbreak investigation reference worksheet and interweave breakout rooms with didactic vignettes covering background information while r
Medical education in resource-limited settings faces significant challenges in providing diverse clinical exposure and fostering essential skills such as clinical reasoning, communication, and empathy. Due to the inability to afford immersive technologies such as virtual reality (VR) and Augmented Reality (AR), constrained by financial, infrastructural, and structural barriers, interactive simulation videos (ISV) constitute an innovative, cost-effective educational tool that can bridge the gap between theoretical knowledge and practical clinical experience, while enhancing student engagement and learning outcomes. This study aimed to assess the educational value and user experience of ISV as a supplementary tool in teaching medical semiology among medical students in Mozambique. A quantitative, descriptive, cross-sectional study was conducted among 4th-year medical students at Alberto Chipande University. Descriptive and inferential statistical analyses were performed, including a one-
Medical education in resource-limited settings faces significant challenges in providing diverse clinical exposure and fostering essential skills such as clinical reasoning, communication, and empathy. Due to the inability to afford immersive technologies such as virtual reality (VR) and Augmented Reality (AR), constrained by financial, infrastructural, and structural barriers, interactive simulation videos (ISV) constitute an innovative, cost-effective educational tool that can bridge the gap between theoretical knowledge and practical clinical experience, while enhancing student engagement and learning outcomes. This study aimed to assess the educational value and user experience of ISV as a supplementary tool in teaching medical semiology among medical students in Mozambique. A quantitative, descriptive, cross-sectional study was conducted among 4th-year medical students at Alberto Chipande University. Descriptive and inferential statistical analyses were performed, including a one-
Objective: To establish a standardized training program for endoscopic pathogen visualization literacy (EPVL) based on fluorescence rapid on-site evaluation (ROSE) technology for gastroenterologists, and to evaluate its training efficacy. Methods: A prospective quasi-experimental study was conducted. A total of 54 gastroenterology trainees were non-randomly allocated into the EPVL training group (Group A, n=28, 16-hour comprehensive training) and the control group (Group B, n=26, 3.5-hour traditional teaching). Pre- and post-training assessments included theoretical examinations, fluorescence ROSE image interpretation tests (30 parallel images per set), interpretation speed measurement, and clinical decision-making integration evaluation. The primary outcome was the change in image interpretation accuracy, analyzed by ANCOVA with pre-test scores as the covariate. Results: Baseline characteristics were comparable between groups (P greater than 0.05 for all demographic variables and pre-
Background. Journal clubs (JCs) are a popular education format. Interest in studying their impact is high, and authors often report positive results related to subjective parameters. Objective assessments of effectiveness are limited and contradictory. In this paper, we share our experience and describe our journal club effectiveness. Methods. We conducted a prospective cohort study within our online journal club. Meetings followed a discussion-based format and were held via Zoom, with timing and topics determined by voting in the club Telegram chat. Enrolment occurred in waves and included an application, entry test, and interview. During each recruitment wave, both club members (treatment group) and applicants (control) completed an admission test assessing knowledge of evidence-based medicine and statistics. Results. The JC currently comprises 27 members. Over the past year, 76 meetings were held, with 75% of participants grading their experience with 9 or 10 on a ten-point scale. M
Introduction: Postgraduate clinical training is crucial for developing professional competence, communication skills, and effective teamwork. Although resident physician selection is crucial, little is known about how Japanese residency programs select residents and whether selection practices are associated with difficulties during training. Methods: We conducted a nationwide cross-sectional survey of residency programs participating in Japan's 2023 General Medicine In-Training Examination (GM-ITE). Program directors completed a questionnaire assessing selection methods, interview content, quality-assurance measures, and resident difficulties, defined as at least one postgraduate year 1 or 2 resident physician receiving disciplinary action or a severe warning. Free-text responses were coded using the Situation, Task, Action, and Result framework. Associations between selection methods and resident difficulties were examined using adjusted logistic regression models controlling for hos
Background. Designing high-quality Objective Structured Clinical Examination (OSCE) stations is a time-consuming process. Generative artificial intelligence (AI) represents a promising path to accelerate content creation by automating the generation of scenarios. A growing number of AI tools is now available for this purpose. Objective. To assess the variability between generative AI models in their ability to produce OSCE stations in the field of paediatrics. Methods. A structured prompt was developed based on the French national OSCE guidelines for medical education. Five distinct AI models were provided with this prompt, alongside the neonatal jaundice chapter from the French pediatric reference textbook, to generate 6 complete OSCE stations. Results. Prompt compliance was high for ChatGPT 5.1, ChatGPT 5.2, Gemini 3.0 Pro, and Claude Opus 4.5, while it was lower for Grok 4.1. Expert-rated quality was generally high, with few factual errors or missing information across models. Howev
Problem Authentic patient encounters are the raw material of clinical learning, yet the educational resources learners receive are rarely keyed to the diagnoses in front of them, creating temporal and cognitive gaps. Precision medical education (PME) proposes delivering the right resource to the right learner at the right moment, but practical implementation in the clinical learning environment remains limited. Approach We developed DxMentor, an electronic health record (EHR)-integrated platform that captures each learner's daily inpatient diagnostic exposures from documented International Classification of Diseases, Tenth Revision (ICD-10) codes. Artificial intelligence (AI) is used to match each diagnosis to an educator-curated formulary of micro-learning resources and board-style questions, and to PubMed-derived primary and synthesis literature converted into plain-language evidence summaries. A personalized email "nudge" is delivered before morning rounds, copying supervising atten
Background Good quality maternity care is critical in reducing maternal morbidity and mortality in regions with a high maternal and perinatal mortality. Building capacity of maternity care providers through in service training has been proven as effective in bridging the knowledge, competency, and skills gap during provision of maternity care. Meaningful involvement of stakeholder perspectives in the design and implementation of public health training programs heightens the prospects of achieving long term changes in practice and policy. Methods To explore stakeholder perspectives and experiences on the co-development of the antenatal-postnatal care course in Kenya, Nigeria, Tanzania. Data was collected through nine key informant interviews and observation notes between July - October 2022. Qualitative data was analysed in NVIVO software using inductive thematic analysis. Results Study findings showed that stakeholders were receptive of the blended learning approach for training in ant
Objective: To investigate and establish a baseline for sex and gender considerations in policy, research and curricula across the state of Victoria. Design and setting: Victoria was selected as a case study for Australia, using a mixed-methods approach to examine health and medical university curricula, research organisation policies and research funding between 2020-2025 prior to mandated inclusion. Main outcome measures: Primary outcomes include identification of predefined sex- and gender-related terms in university curricula descriptors and funded grant descriptions; and questionnaire responses from university course coordinators and organisational leads. Results: Data mining across nine Victorian universities (318 courses/3383 units) identified ~93% of units and ~60% of healthcare courses lacked sex and gender terms in their descriptors. Among medical research organisations operating in Victoria, including peak bodies, research institutes, hospitals and universities, ~70% (18/26)
Background and objective: Problem-based learning (PBL) strategies such as case studies and role-playing are increasingly adopted in medical education to promote critical thinking, communication, and teamwork. However, comparative evidence regarding their effectiveness in medical biochemistry remains limited. This study aims to evaluate and compare the impact of case studies and role-playing on learning outcomes, engagement and skill development among first year MBBS students. Material and methods: An educational randomized controlled trial (RCT) with a within-subjects crossover design was conducted in the Department of Biochemistry involving forty-five first-year MBBS students. Participants were randomly assigned to receive case studies and role-playing in a counterbalanced sequence. Learning outcomes were assessed using pre- and post- tests, a structured feedback questionnaire and qualitative reflections. Data were analysed using appropriate parametric and non-parametric tests, along