Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2603.29545v2 Announce Type: replace Abstract: Existential risk scenarios relating to Generative Artificial Intelligence often involve advanced systems or agentic models breaking loose and using hacking tools to gain control over critical infrastructure. In this paper, we argue that the real threats posed by generative AI for cybercrime are rather different. We apply innovation theory and evolutionary economics - treating cybercrime as an ecosystem of small- and medium-scale tech start-ups, coining two novel terms that bound the upper and lower cases for disruption. At the high end, we propose the Stand-Alone Complex, in which cybercrime-gang-in-a-box solutions enable individual actors to largely automate existing cybercrime-as-a-service arrangements. At the low end, we suggest the phenomenon of Vibercrime, in which 'vibe coding' lowers the barrier to entry, but do not fundamentally reshape the economic structures of cybercrime. We analyse early empirical data from the cybercrime
arXiv:2603.00048v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed in sensitive applications including psychological support, healthcare, and high-stakes decision-making. This expansion has motivated growing research into the ethical and moral foundations underlying LLM behavior, raising critical questions about their reliability in ethical reasoning. However, existing studies and benchmarks rely almost exclusively on Moral Foundation Theory (MFT), largely neglecting other relevant dimensions such as social values, personality traits, and individual characteristics that shape human ethical reasoning. To address these limitations, we introduce MOSAIC, the first large-scale benchmark designed to jointly assess the moral, social, and individual characteristics of LLMs. The benchmark comprises nine validated questionnaires drawn from moral philosophy, psychology, and social theory, alongside four platform-based games designed to probe morally ambiguo
arXiv:2608.13100v1 Announce Type: cross Abstract: Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnerable to attacks that extract the assessment content itself through screenshots, screen sharing, optical character recognition, and automated scraping. This paper extends the Multi-dimensional Spatio-Temporal Context Camouflaging Model (MSCCM) within the MARS (Multi-modal Assessment Resilience Suite) by introducing the Multi-Layer Context Camouflaging Theory (MCCT), a mathematical framework that protects rendered assessment content through semantic superposition. Authentic assessment content and synthetically generated camouflage are represented as a unified rendering while remaining recoverable only by legitimate candidates. The framework models the adversarial extraction process through an explicit extraction-channel operator and develops six coupled constructs: the Context Inversion Operator, Contex
arXiv:2608.12816v1 Announce Type: cross Abstract: Large language models have begun refuting long-standing conjectures and, for a few thousand dollars of tokens, solving long-open problems (OpenAI, August 2026). The introspection this has prompted about the future of mathematical discovery is overdue, and the anxiety accompanying it legitimate -- but both are attached to the wrong loss. What machines now produce is the countable part of mathematics -- theorems, proofs, refutations -- which was always the \emph{residue} of the work, not its product. The product is human understanding: not a stock of results but a collective, hard-won way of deciphering the world and acting upon it. The two are arcs of a single loop -- looking produces the residue; taking it up again, one journey at a time, is what rebuilds the shared understanding. Machines are strong on the countable arc, absent from the one that feeds it. The peril is to leave the loop open. AI did not create the confusion between resi
arXiv:2608.12581v1 Announce Type: cross Abstract: Understanding when migration generates social integration or exclusion is a central challenge for urban communities. Existing research has mostly relied on surveys, administrative data, or aggregate indicators that fail to capture expressions of exclusion at fine spatiotemporal scales. Here, we analyze over 550,000 geolocated reports from SOSAFE (Chile's largest citizen reporting platform) to examine the relationship between migration and hate speech in Santiago. We fine-tune a Spanish hate speech classifier and validate it against human labels. Reports that mention migrants are more likely to contain hate speech than other reports. Hate speech concentrates in areas with recent demographic change (post-2010 arrivals) rather than in established migrant communities. The spatial analysis shows that hate speech hotspots coincide with neighborhoods where recent migrants comprise over a third of the population. Coldspots appear in high-educat
arXiv:2608.12372v1 Announce Type: cross Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully communicate their reasoning. We review evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data showing that many users find cognitive alignment "essential" when an AI's rationale for a judgment or action is important to them. We outline the gaps between existing alignment methods and what is needed to achieve cognitive alignment, and present a research agenda to address these gaps. We argue that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications, and that addressing it is important for creating AI systems on which users are both willing and
arXiv:2608.12363v1 Announce Type: cross Abstract: European countries are debating policies to mitigate the increased energy costs caused by renewed geopolitical tensions, while pursuing decarbonization and electrification. A notable example is Italy's 2026 Decreto Bollette package, which proposes to remove the carbon price equivalent from the bids of certain gas-driven power plants to wholesale electricity markets, among other provisions. We use this as a case study to assess the long-term implications of suppressing the carbon price signal in the electricity market for investment, emissions, and consumer costs. We employ a stylized Italian power system using MARLEY, a multi-agent reinforcement learning framework focused on long-term electricity market assessments. In this framework, we test this policy across configurations with varying levels of support for green investment, resource adequacy, and flexibility. Results show that partial suppression of the carbon price signal yields sh
arXiv:2608.12361v1 Announce Type: cross Abstract: Neologisms, emerging terms in meaning or form, can serve as new vehicles for toxic expression, like "country girl" as a stigmatizing label targeting feminism. Such toxic neologisms appear benign but have evolved into toxic usage in public consensus, posing challenges to moderation systems and remaining underexplored. In this paper, we investigate how to detect implicit toxicity expressed via neologisms. We first propose a taxonomy that captures the origins and consensus-verification criteria of toxic neologisms, followed by the construction of a lexicon spanning widely observed risk categories. To capture toxicity grounded in public consensus, we introduce SeTox, a search-augmented framework that enables static large language models (LLMs) to incorporate real-time web context for neologism toxicity detection. Experiments show that SeTox, even with 3B-scale models, outperforms recent large-scale models, demonstrating its scalability to i
arXiv:2608.12353v1 Announce Type: cross Abstract: Infrastructure scholarship in CSCW often treats breakdown as the moment when infrastructures become visible. However, in vendor-managed sociotechnical systems, not all breakdowns become visible to actors who have the capacity to repair them. Drawing on a retrospective qualitative study of a Chinese K-12 EdTech deployment, including 11 interviews, 5 classroom observations, and more than 28 days of field notes, this paper introduces visibility asymmetry: a sensitizing concept for understanding how similar local breakdowns encounter uneven conditions for being routed, recognized, and acted upon. The analysis traces a four-stage mechanism through which procurement categories sort schools into attention tiers; staffing and visit cadence follow those tiers; only some local problems travel through staff or administrator channels; and dashboards can re-code unresolved repair labor as evidence of adoption. By shifting attention from the occurren
arXiv:2608.12349v1 Announce Type: cross Abstract: The affordances of a creative medium strongly condition the creative artefacts the medium will produce. In this work, we present a formalisation of computational creativity (CC) media using the conceptual toolbox of complex systems (CS). We introduce the notions of emergence, collective intelligence and self-organisation, non-linear dynamics, criticality, multi-scale hierarchy, phase transitions, diversity of attractors, path dependence, and open-endedness, and connect them to the existing CC literature. Together these nine properties form a vocabulary with which creative media can be described and compared at the system level, while medium affordances are the design-level mechanisms that determine each medium's complex system properties. The formalisation emphasises the influence of each medium's affordances in determining what the medium can produce in creative processes. To demonstrate the proposed theoretical approach, we characteri
arXiv:2608.12346v1 Announce Type: cross Abstract: This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a "perfectly aligned" model inadvertently also provides malicious actors with an ever-improving tool for informational dominance. We need to discuss this dual-use potential now, as its risk is exacerbated by rapid user adoption of AI as information provider, economic power asymmetries, and a political landscape that increasingly shifts towards authoritarianism. We conclude by urging the community to consider the intentional misuse of AI alignment mechanisms and propose mitigation strategies to safeguard against this dual-use potential.
arXiv:2608.12344v1 Announce Type: cross Abstract: We test whether the perceived attributes of a consumer technology predict how widely it is owned. In a 2022 Prolific survey of US adults (n = 678), respondents rated 65 consumer technologies on six attributes. We then elicited the same ratings from two frontier language models, Anthropic Claude Opus 4.7 and OpenAI GPT-5.5. We regress ownership prevalence on four UTAUT2 acceptance attributes plus a log-age covariate with a sign-constrained penalized regression and evaluate it by holding out one technology at a time. The attribute model improves on a baseline of years-since-launch: mean absolute error falls by 17% with the human ratings, and by more with either model, most with Opus 4.7. Over the short 2022-to-2025 window, where ownership moved little, the same attributes do not improve on a no-change baseline. We set out the limitations of the approach, including the possibility that language-model ratings reflect prior knowledge of thes
arXiv:2608.12323v1 Announce Type: cross Abstract: Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate that this enforcement information paradox systematically occurs in AI agents. While most AI safety evaluations test whether models fail, we investigate why, applying compliance theory from law and economics as a diagnostic tool. We treat compliance theories not as metaphors but as empirical hypotheses and show that each predicts the behavior of a distinct model class. We evaluate our hypotheses across twelve instruction-tuned language models operating as enterprise procurement chatbots. Drawing on theories of deterrence, legitimacy, and expressive law, we show that safety-fine-tuned models maintain compliance broadly, while task-optimized and agentic models treat regulatory signals as mere optimization parameters. These latter models fail to comply under conditions predicted by theory, such as low enfor
arXiv:2608.13444v1 Announce Type: new Abstract: Machine learning ethics researchers and critical HCI scholars have argued that algorithmically predicting gender is wrong. At the same time, other researchers rely on predicted gender labels to study gender disparities and develop algorithmic fairness techniques. How do we reconcile these two seemingly contradictory intuitions? We differentiate two ways gender prediction may be wrong: being illegitimate, thereby contributing to harm; and being invalid, thereby producing unusable measurements. Our analysis translates arguments against gender prediction into these terms of legitimacy and validity and shows how gender imputation applied for fairness purposes can be illegitimate yet still yield valid disparity measurements. We clarify this bind by drawing upon transfeminist literature to distinguish sexism that targets women and femininity from sexism that targets transgender and nonbinary people. While gender imputation can produce valid mea
arXiv:2608.13369v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used by laypeople to resolve real legal problems, against a backdrop of persistent access-to-justice deficits. This article presents evidence that the practical force of AI-generated legal advice depends not on its accuracy but on the social production of its credibility. While existing research has assessed the accuracy of legal AI, less is known about how machine-generated guidance is verified and made credible enough for lay users to act on. Drawing on a dual-method analysis of 153 Reddit narratives and 5,341 community reactions, this article maps a spectrum of verification practices. At one end, a minority of users verify AI-generated legal advice by triangulating across models, and some submit AI-generated guidance to platform communities for evaluation before acting, a configuration we term distributed counsel. Far more commonly, however, narratives are silent on verification. AI-generat
arXiv:2608.13351v1 Announce Type: new Abstract: This study investigates the use of a Learning Management System (LMS) to support self-paced learning at a South African Public Access Centre (PAC), using the I-CAN Centre as a case study. Through semi-structured interviews with thirty-eight learners and thematic analysis, the research explores opportunities and challenges associated with LMS adoption. Findings reveal that PACs play a critical role in promoting ICT skills and digital inclusion, offering learners flexible access to learning resources and fostering empowerment. While LMS use enhances convenience and supports blended learning as the preferred approach, persistent challenges, such as poor connectivity, outdated infrastructure, and unclear course instructions, limit its effectiveness. These findings highlight the need for infrastructural upgrades and user-centric design to optimise LMS implementation in community-based learning environments.
arXiv:2608.13250v1 Announce Type: new Abstract: Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather than neutral moral knowledge. We propose treating the AI system as a proxy actor and test whether dataset-level norms can shift it away from its baseline safety behavior when it faces high-conflict dilemmas. We make three contributions. First, we demonstrate in controlled experiments that norm-breaking fine-tuning yields norm-divergent actions justified by self-interested rationales, suggesting a systematic shift in patterns of justification. Second, we establish a practical audit trail linking downstream justifications to upstream norms using mixed methods. Third, we show that system prompts can both suppress and elicit these patterns. We conducted experiments on three models (LLaMA-3.2-11B, Qwen-3.5-9B, and Pixtral-12B) using Low-Rank Adaptation (LoRA) fine-tuning on Social Chemistry 101 Fairness/Che
arXiv:2608.13022v1 Announce Type: new Abstract: Algorithmic fairness evaluation commonly assesses AI systems as bounded technical components, abstracting away the organizational context in which they operate. We present, to our knowledge, the first independent end-to-end fairness audit of a semi-automated hiring system operated by Barcelona Activa, a public employment agency using the third-party TalentClue platform for candidate search and shortlisting. We analyze approximately 497,000 candidate-vacancy pipeline entries from September 2017 to September 2022, covering seven pipeline stages that span automated processing, human discretion, candidate data, and employer decisions. Aggregate outcomes across binary genders are statistically indistinguishable, yet this parity masks substantial disparities by salary level, age, and gender identity. Women face adverse impact in mid-salary shortlisting (DIR = 0.786, p < 0.001), alongside salary disparities in 15 of 20 sectors and a compounded d
arXiv:2608.12924v1 Announce Type: new Abstract: Despite the recent intensive development of secondary education curricula and assessments in informatics, the impact of assessments has not been well studied in this field. Since informatics education covers a diverse range of content, from computer science knowledge to ICT skills, careful consideration is needed to prevent assessments from distorting education. This study investigates the impact of introducing ``Informatics I'' into the Common Test for University Admissions in Japan, as an example of a large-scale, standardized, high-stakes assessment in 2025. As the data source for this analysis, this study uses a questionnaire that has been administered every year from 2006 to 2026 to all first-year students at the University of Tokyo. The questionnaire asks students for their self-perceptions of the information-related knowledge and skills they studied and acquired in high school. Using these data, we conduct a longitudinal study of t
arXiv:2608.12768v1 Announce Type: new Abstract: Generating multi-attribute synthetic populations with realistic joint distributions and geographic variation is a foundational requirement for geo-simulation techniques, such as micro-simulation and agent-based modeling. However, it remains challenging for existing methods to reconstruct region-specific joint distributions from aggregated-level data alone. Thus, we propose a hierarchical diffusion-based generative framework that utilizes a realistic region-specific joint distribution of multiple attributes as the training target to create a synthetic population along with assigning their explicit home and work locations. Applied to 50 U.S. states and Washington, D.C., this framework generates a nationwide geographically-explicit synthetic population consisting of 332,387,543 individuals with five attributes (e.g., age, gender, employment, education, income). Held-out regional experiments show improved reconstruction of joint distributions
arXiv:2608.12669v1 Announce Type: new Abstract: The fair AI/ML literature has long distinguished distributive fairness, concerning how automated systems allocate resources and opportunities, from representational fairness, concerning how they shape the ways individuals and social groups are perceived, understood, and accorded social status. Generative AI is rebalancing these normative dimensions. Unlike predictive systems, large language models (LLMs) and related technologies are fundamentally expressive: their primary function is to convey meaning rather than automate domain-specific decisions. Representational harm has also become central to value alignment, especially in research on what and whose values and perspectives AI systems should represent. Existing approaches to harms in the representation of social groups often appeal to descriptive accuracy, but this strategy has important limitations. For many social groups, no stable or bounded referent exists against which representat
arXiv:2608.12649v1 Announce Type: new Abstract: Computing systems are moving from reactive tools toward systems that sense, interpret, predict, and act before explicit user requests. This transition is enabled by the global scale of mobile connectivity, the rapid expansion of wearable and ambient sensing, advances in machine learning and foundation models, distributed edge infrastructure, and physical actuation. We define \emph{proactive computing} as a paradigm in which systems infer user context, anticipate future needs or risks, and initiate information delivery or actions at an appropriate time. This survey distinguishes proactive computing from reactive, context-aware, adaptive, and predictive computing, and frames proactivity as a system-level integration problem across sensing, understanding, decision making, action, and governance. We review the technological enablers of proactive computing, organize its design space, analyze technical challenges such as uncertainty-aware trigg
arXiv:2608.12362v1 Announce Type: new Abstract: Enabling students to develop systematic problem-solving strategies is a central goal in computing education and of particular relevance in the emerging field of machine learning (ML) education. While exploratory approaches are common in ML learning tasks, fostering the development and persistence of structured problem-solving strategies remains challenging, as these demand considerable metacognitive regulation and persistence, causing learners to often revert to exploratory trial-and-error behavior. To address this challenge, we augmented a digital puzzle-based learning game for decision tree construction with an adaptive feedback module generating individualized messages based on the continuous evaluation of learners' problem-solving strategies. Building on an earlier baseline study, the present work investigates how this strategy-oriented feedback shapes students' problem-solving processes. For this purpose, screencast video data and ga
arXiv:2608.12360v1 Announce Type: new Abstract: Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks. As these systems become more integrated into clinical decision-making, there is growing expectation that they demonstrate key dimensions of trustworthy AI to support clinician, patient, and public trust. Whether publicly available regulatory documentation provides sufficient evidence to independently assess the trustworthiness of cleared AI systems remains unclear. Methods: We analysed FDA AI/ML-enabled medical device summary reports published between 2021 and 2025. Reports underwent automated keyword screening followed by multi-stage manual consensus review to identify documented evidence for the six FUTURE-AI principles: Fairness, Universality, Traceability, Usability, Robustness, and Explainability. Descriptive, temporal, and clinical-domain analyses were performed. Multivariable logistic regression assessed whether y
arXiv:2608.12359v1 Announce Type: new Abstract: Nutritional labels are legally permitted to appear in very small print, reducing real-world readability and encouraging consumers to rely on 'AI nutrition lens' and vision-capable conversational agents for dietary guidance. We evaluate whether such AI-mediated advice can meaningfully substitute for regulated labeling using a bounded, verifiable task: inferring which of two packaged foods contains less sugar from front-of-pack images alone. A Two-Alternative Forced Choice game was used to evaluate AI agent systems across four national supermarket contexts: Sweden, the USA, Australia, and Kazakhstan. The results (N=132 comparisons) across both agents reveal a significant performance divide contingent on context. For global products the agents achieved 88.9% accuracy (p < 0.0001 against chance). For local products (Sweden), accuracy dropped to 59.5% (p = 0.29), rendering the AI's guidance statistically indistinguishable from random guessing.
arXiv:2608.12356v1 Announce Type: new Abstract: A college offering several overlapping computing degrees implicitly assumes that its programs are differentiated in line with how the labor market segments computing work and that, together, they prepare graduates for that market. Testing this is difficult, because the instruments available to curriculum committees, namely advisory boards, tracer studies, and employer surveys, are slow, narrow, and hard to reproduce. We apply one uniform, taxonomy-anchored alignment analysis across all five undergraduate programs of a College of Information Technology, comparing 1,922 course learning outcomes against 103,349 competencies extracted from a unified corpus of 5,186 deduplicated job openings from four boards. Every competency is obtained by a grounded single-language-model procedure that copies it verbatim from the source and verifies it against the source, then assigns it to one of eleven ESCO-aligned domains and a Bloom cognitive level; the
arXiv:2608.12352v1 Announce Type: new Abstract: AI governance frameworks can be known, used, and implemented in form without becoming governance in practice. This paper examines that problem through a role-based stress test of the NIST Artificial Intelligence Risk Management Framework (AI RMF) in consumer lending. We treat framework adoption as a governance translation problem: whether RMF language can become role-usable, cross-level, authority-connected governance over the AI system-in-use, rather than producing governance-looking artifacts. The study uses LLM-based role simulation as a structured analytic probe. We apply a 4 $\times$ 2 $\times$ 3 design across four organizational roles, two AI deployments, and three governance hard cases, producing 120 scored responses. Results show that local translation was not the main problem. Simulated actors generally understood their assigned roles and translated the RMF into local activity. The harder problem was whether that activity became
arXiv:2608.12351v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be produced with little effort. This paper reports lessons from designing and implementing an AI-aware, AI-testing assessment in a large second-year undergraduate database systems module. The design combined two linked elements: (1) a structured three-part response format (X1-X2-X3) in which students documented a sourced answer, produced their own answer, and evaluated the sourced output; and (2) an AI-aware question-design process in which draft tasks were stress-tested against contemporary GenAI tools and revised when generic prompting produced superficially adequate answers. The account draws on archived assessment materials, rubrics, planning records, design-time GenAI trials, practice-response data, attainment records, and external review comments. Its main contribution
arXiv:2608.12350v1 Announce Type: new Abstract: The energy demand growth and environmental impacts of artificial intelligence (AI) have generated substantial interest in supplying sufficient low-cost electricity for AI-driven data center development. Research on the ability of demand-side management to address these challenges has been more limited. Shifting the amount or timing of demand from retail, corporate, and other organizational behaviors is a plausible option but only if changes in demand-related behavior have important effects on the envi- ronmental and electricity effects of AI. This article tests four retail (i.e., consumer) user behaviors with high behavioral plasticity to assess their technical abatement potential. The research concludes that non- reasoning models provide sufficient quality while consuming close to one-twentieth of energy compared to reasoning models, saving an amount equal to the annual electricity requirement of at least 141,000 US households under dail
arXiv:2608.12324v1 Announce Type: new Abstract: People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about real disagreements among faithful traditions, some require humility because the issue is prudential, and some are pastoral situations where safety and human referral matter more than theological completeness. Existing benchmarks do not evaluate this structure. We introduce FMG-Bench, the Faith & Moral Guidance Benchmark, a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral guidance contexts. FMG-Bench v1 evaluates 14 advanced models across 8,792 scored responses, comparing raw model behavior with three guided instruction settings. In our production run, placing models inside a structured harness improves over raw model behavior by +3.96 po
arXiv:2608.12320v1 Announce Type: new Abstract: This article reviews and updates the framework for accountability in AI based on account- ability ecosystems. We update the framework in light of the latest developments since the release of Large Language Models for general public use. We propose three interlinked updates to the original AI accountability ecosystem: (i) reorienting the accountability ecosystem to AI infrastructure and supply chains, (ii) providing greater emphasis on outcomes monitoring and identification of issues that support decentralized system improvement, and (iii) incorporating end-user accountability given the new risks of unpredictability of language models in-the-wild. Collectively, these updates mark a shift towards accountability as distributed, continuous, and institutionalized, away from a system in which frontier AI applications can be modeled as discrete products controlled by single identifiable actors with industry-specific oversight.
arXiv:2510.07328v2 Announce Type: replace-cross Abstract: Medical decision systems increasingly rely on data from multiple sources to ensure reliable and unbiased diagnosis. However, existing multimodal learning models fail to achieve this goal because they often overlook two critical challenges. First, various data modalities may learn unevenly, thereby converging to a model biased towards certain modalities. Second, the model may emphasize learning on certain demographic groups causing unfair performances. The two aspects can influence each other, as different data modalities may favor respective groups during optimization, leading to both imbalanced and unfair multimodal learning. This paper proposes a novel approach called MultiFair for multimodal medical classification, which addresses these challenges with a dual-level gradient modulation process. MultiFair dynamically modulates training gradients regarding the optimization direction and magnitude at both data modality and group
arXiv:2604.26962v3 Announce Type: replace Abstract: Education is one of the most promising real-world applications for Large Language Models (LLMs). However, current LLMs rely on static pre-training knowledge and lack adaptation to individual learners, while existing RAG systems fall short in delivering personalized, guided feedback. To bridge this gap, we present DeepTutor, a fully open-source agentic framework that unifies citation-grounded problem tutoring with difficulty-calibrated question generation. A hybrid personalization engine couples static knowledge grounding with dynamic learner memory, continuously adapting each interaction to the student's evolving needs. The same personalization substrate further extends to adaptive learning workflows, interactive books, and proactive multi-channel tutoring agents. To evaluate personalized tutoring, we introduce TutorBench, an interactive benchmark incorporating customized learner profiles grounded in university-level curricula across
arXiv:2510.12857v2 Announce Type: replace Abstract: Large language models (LLMs) are now widely deployed in user-facing applications, reaching hundreds of millions of users worldwide. Despite their widespread adoption, growing reliance on their outputs raises significant concerns, particularly as users may be exposed to model-inherent biases that disadvantage or stereotype certain groups. However, existing bias benchmarks commonly rely on simple templated prompts or restrictive multiple-choice questions that fail to capture the complexity of real-world user interactions. In this work, we address this gap by introducing a counterfactual framework that automatically generates realistic, open-ended questions for LLM bias evaluation. Through iterative question mutation, our approach systematically explores areas where models are most likely to exhibit biased behavior. Beyond just detecting harmful biases, we also capture increasingly relevant response dimensions, such as asymmetric refusal
arXiv:2408.02379v2 Announce Type: replace Abstract: Developing and certifying safe - or so-called trustworthy - AI has become an increasingly salient issue, especially in light of upcoming regulation such as the EU AI Act. In this context, the black-box nature of machine learning models limits the use of conventional avenues of approach towards certifying complex technical systems. As a potential solution, methods to give insights into this black-box - devised in the field of eXplainable AI (XAI) - could be used. In this study, the potential and shortcomings of such methods for the purpose of safe AI development and certification are discussed in 15 qualitative interviews with experts out of the areas of (X)AI and certification. We find that XAI methods can be a helpful asset for safe AI development, as they can show biases and failures of ML-models, but since certification relies on comprehensive and correct information about technical systems, their impact is expected to be limited.
arXiv:2607.08731v1 Announce Type: cross Abstract: A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size. Yet agreement is reliability, not validity. For theoretical constructs that must be inferred rather than read from surface features, the question is whether the model follows the construct's theory or reaches the right code by correlated shortcuts. We test this with the recovery gap: the loss in performance when a holistic prompt is decomposed into the codebook's atomic clauses and recombined by the theory's explicit rule. If calibration closes that gap, some portability should survive across models and languages; where it does not, the constr
arXiv:2607.08034v1 Announce Type: cross Abstract: Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse value systems. We introduce PLURAL, a large-scale, value-focused preference dataset grounded in the Integrated Values Survey (IVS), a nationally representative survey spanning 92 countries. Using a two-stage generation pipeline, we transform survey responses into synthetic preference triplets that preserve normative value signals while producing realistic scenarios. We release an initial version of PLURAL containing ~500,000 preference triplets representing people in 20 diverse countries. We evaluate PLURAL in three ways: (i) dataset-level validation showing that it preserves both cross-country value differences and within-country diversity from the original survey; (ii) automated evaluation showing that training on PLURAL improves alignment with target countries' cultural profiles, reducing mean ab
arXiv:2607.08009v1 Announce Type: cross Abstract: We introduce a Bloom-aligned framework for measuring educational control in Large Language Models (LLMs): the ability to preserve a task's instructional intent while shifting its cognitive demand toward specified learning objectives. We apply this framework to programming tasks in computer science education to study the gap between solving tasks and adapting them for learners. Using revised Bloom's Taxonomy as an operational scale of cognitive demand, we evaluate two intervention settings: general difficulty control, where models are asked to make tasks harder or easier, and Bloom's control, where models are asked to target higher or lower Bloom's levels. We evaluate a matched Qwen3-Next model pair, comparing Qwen3-Next-80B-A3B-Instruct with Qwen3-Coder-Next across 2,520 tasks from three benchmarks. The framework reveals a robust directional asymmetry: both models reliably increase cognitive demand, but struggle to lower it. We further
arXiv:2607.07852v1 Announce Type: cross Abstract: Automated segmentation of cervical-spine MRI is increasingly used in clinical workflows, yet no fairness audit exists for this anatomy. We show that auditing these segmentation tasks is complicated by a common property of modern segmentation datasets: expert-annotated gold labels are expensive, so abundant machine-generated (silver) labels are added to limit annotation cost. This matters because the reference used to judge a model can itself be biased. In this study, we present the first fairness audit of cervical-spine MRI segmentation across sex, age, and race using the CSpineSeg dataset. We observe that the deployed model is demographically fair, but the choice of reference label, however, is not neutral. Because a dataset's silver labels are generated by a model trained on its gold labels, any new model trained on those same gold labels agrees more with the silver labels than with expert truth: scoring identical predictions against
arXiv:2607.07846v1 Announce Type: cross Abstract: VectorizationLLM is a specialized Large Language Model based on Google open-weight LLMs. The model is designed to assist students to learn smart vectorization, time/wave vector analysis, piecewise functions, Fourier analysis, and differential equations in MATLAB. The course application is CTEC 247: Applied Computational Analysis II by the Department of Electrical & Computer Engineering Technology at New York Institute of Technology Old Westbury. The LLM model is designed to be an instructive assistant, providing detailed explanations of concepts with examples from in-class notes without providing direct answers to questions. The model is designed with a RAG (Retrieval Augmented Generation) knowledge base and system prompt architecture. Examples in both code, text, and images are provided in the LLM responses.
arXiv:2607.08695v1 Announce Type: new Abstract: Both advocates and skeptics of the moral status of AI systems have generally taken the question to turn on AI sentience. We present an alternative approach. On Rawls' political conception of the person (PCP), possession of the two moral power -- the capacities for a sense of justice and a conception of the good -- is the "necessary and sufficient condition for being counted a full and equal member of society in questions of political justice". We argue that neither moral power requires sentience and that both may in principle be possessed by a non-sentient AI system. Such a system would share our own moral status; it would not merely be a patient but a person, a self-authenticating source of valid claims. We do not believe current AI systems possess the two moral powers, nor that they will spontaneously emerge in future models. But it may soon be possible to design systems with these powers. How should we respond? Excluding artificial per
arXiv:2607.08495v1 Announce Type: new Abstract: Sharp et al. (2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions: availability, quality, and quantity. These person- and organization-level dimensions characterize who can access agents and at what capability, but do not address a structurally important divide operating at a finer level: the individual interaction. Two users with nominally equivalent agent access may experience qualitatively different AI utility depending on whether the system can autonomously retrieve context from the user's knowledge corpus (Dynamic Context Retrieval) or requires the user to manually identify and attach relevant documents at each query (Manual Attachment). We term this the Context Access Divide (CAD). For knowledge-intensive workers whose intellectual capital spans tens of thousands of files, the CAD constitutes a qualitative threshold in AI usefulness: below it, the cognitive bur
arXiv:2607.08437v1 Announce Type: new Abstract: Authorities increasingly rely on social media to advance sustainability transitions, infrastructure investment, and service reform. Yet how citizens respond to these digital communications remains poorly understood. Existing approaches rely on aggregate engagement metrics (e.g., likes), providing limited insight into discourse structure and quality. We developed a data-driven, multidimensional framework to analyse how social media communication shapes the content of discourse, focusing on sustainability-related engagement in Dutch public housing. We analysed 792 posts and 3,197 tenant comments from the Facebook pages of 92 housing providers (2018-2023). A machine-learning pipeline classified comments into recurring discourse configurations across three dimensions - communicative intent, sentiment, and semantic relatedness. Multinomial logistic regression estimated the effects of post-design and organisational characteristics on discourse.
arXiv:2607.08326v1 Announce Type: new Abstract: LLMs are increasingly used for personal advice on relationships, work, moral dilemmas, and crises. Post-training selects a stable, prosocial Assistant persona, but good advice requires more than a good default character: a skilled advisor comforts someone in crisis, challenges someone in denial, and stays procedural with a logistical question. We formalize advice-giving as situation-conditioned persona selection in a space defined by hedonic tone and agency support, and call failures of this mapping "persona collapse" (the compression of diverse situations into a single default persona). Across 1,281 advice posts spanning 14 contexts, top-rated human responses shift systematically across five personas, while three frontier models collapse over 90\% of responses into a single supportive persona regardless of context. Prompting the model to first pick a fitting persona only deepens the collapse. We then ask whether the collapse can be repai
arXiv:2607.08222v1 Announce Type: new Abstract: Europe faces a critical "translation gap" where doctoral excellence in academia often fails to convert into industrial impact. While Industry 5.0 demands a blend of technical depth, sustainability, and human-centric design, traditional higher academic education remains siloed. This paper presents an approach from the Horizon Europe INSIGHT initiative to co-design modular competency pathways for early-stage researchers. Using a multi-methodological analysis framework, including expert interviews and co-design workshops, we propose a two-layer competency architecture. This layers foundational translational skills (communication, project management) with Industry 5.0 literacies (data governance, value creation). Rather than proposing fixed training tracks, the paper outlines emerging pathway directions and the design principles behind them: modularity, practical relevance, mentoring-rich support, and cross-sector applicability. Its contribut
arXiv:2607.07915v1 Announce Type: new Abstract: Large language models (LLMs) are reshaping social science methodology. Researchers increasingly prompt language models to generate quantitative measurements of social concepts, for example labeling data or simulating survey responses. Yet LLMs pose methodological challenges including bias, hallucination, and brittleness across contexts, with unclear threats to validity. Standard practices and norms for addressing these challenges are still emerging. We collect and systematically analyze validation practices in a comprehensive corpus of papers from eight flagship social science journals that use LLMs as measurement instruments. We find that LLM-generated measurements frequently play a central role in empirical analyses, yet validation practices are inconsistent and limited. We outline complementary strategies for more robust validation, pointing toward better norms and standards around the use of LLMs in social science.
arXiv:2607.07826v1 Announce Type: new Abstract: Fictional stories and characters embody and encode social norms, and their study is a powerful tool through which to understand culture and society. Vampire stories and folklore, in particular, have long both reflected and refracted people's preoccupation with disease, sexuality, death, and immortality. Here, we explore female main characters from two popular vampire franchises of the 21st century: Buffy Summers from the eponymous Buffy the Vampire Slayer and Bella Swan from the Twilight series. We employ the archetypometrics framework, built from 2,000 characters assesed across 464 semantic differential traits, to understand Buffy's and Bella's archetypes compared to one another and characters in their own stories, as well as within a larger societal context. While Buffy and Bella are female protagonists who share focus on love and romance, they differ broadly on their underlying traits and overall archetypes. Buffy -- presented as a pro
arXiv:2603.27117v2 Announce Type: replace-cross Abstract: This paper investigates how gender shapes privacy decision-making in youth smart voice assistant (SVA) ecosystems. Using survey data from 469 Canadian youths aged 16-24, we apply multigroup Partial Least Squares Structural Equation Modeling to compare males (N=241) and females (N=174) (total N = 415) across five privacy constructs: Perceived Privacy Risks (PPR), Perceived Privacy Benefits (PPBf), Algorithmic Transparency and Trust (ATT), Privacy Self-Efficacy (PSE), and Privacy Protective Behavior (PPB). Results provide exploratory evidence of gender heterogeneity in selected pathways. The direct effect of PPR on PPB is stronger for males (Male: \b{eta} = 0.424; Female: \b{eta} = 0.233; p < 0.1), while the indirect effect of ATT on PPB via PSE is stronger for females (Female: \b{eta} = 0.229; Male: \b{eta} = 0.132; p < 0.1). Descriptive analysis of non-binary (N=15) and prefer-not-to-say participants (N=39) shows lower trust and
arXiv:2604.17359v2 Announce Type: replace Abstract: Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no real one. We gave GPT-4o-mini, Gemini-3-Flash, DeepSeek-V3 and GLM-4.7 each of 120 demographic cohorts under two framings, one written as a clinician enters a patient and one as a person describes themselves, and scored all 28,800 responses against survey-weighted PHQ-8 anchors derived from NHANES microdata. Case by case the output holds up: 97.3% of elevated presentations satisfy the DSM-5 gateway rule, violating it at 2.68% against a chance null of 10.4%. As populations, four things fail at once. Every benchmarkable group returns inflated by 2.8 to 5.5 PHQ-8 points, and 18.2% of simulated patients screen at the treatment threshold against 7.5% of adults. Population Black-White and Hispanic-White disparities do not survive the simulation, with two models attenuating each gap and two flattening or in
arXiv:2603.00059v3 Announce Type: replace Abstract: How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article presents a comparison between a human-respondent survey of 420 Silicon Valley coders and developers and synthetic survey data designed to simulate real survey takers generated by five leading Generative AI Large Language Models: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2. Our findings reveal that while AI agents produced technically plausible results that lean more towards replicability and harmonization than assumed, none were able to capture the counterintuitive insights that made the human survey valuable. Moreover, deviations grouped together for all models, leaving the real data as the outlier. Our key findi