Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2608.12792v1 Announce Type: new Abstract: Contribution statements are an increasingly common way to make research labor visible, reduce academic malfeasance, and provide broader transparency. Despite this potential value, they remain uncommon in visualization and HCI. To explore this gap, we conducted an online study with (N=21) visualization and HCI researchers. We find a range of differing opinions about the utility of contribution statements, which are set against a background of tensions relating to contribution frameworks that inadequately fit contribution types in HCI and especially visualization, power dynamics between authors, bias in authorship perceptions, and the tedium of providing yet another form of documentation. From these factors, we offer a modest recommendation to authors: consider contribution statements. There are contexts when they may usefully explicate work, and others where they can cause author-team conflict or become a burdensome chore. Regardless of wh
arXiv:2608.12759v1 Announce Type: new Abstract: The ability to navigate outdoors safely and independently is crucial yet challenging for people with low vision (PLV). While various augmented reality (AR) systems for low vision have been designed and evaluated in ideal lab environments, no research has investigated their real-world feasibility and challenges. We present NavSight, a mobile AR application that assists PLV in outdoor navigation by recognizing important outdoor objects (e.g., curb, vehicle) and rendering real-time visual augmentations. Through a seven-day diary study with 12 PLV in real-world settings, we characterize the impact of NavSight on scene perception, users' configuration strategies on what objects to augment and how to augment them across scenarios, how users made sense of and responded to recognition errors, and the social acceptability of using NavSight in public. We further identify environmental factors affecting recognition, such as weather conditions, light
arXiv:2608.12605v1 Announce Type: new Abstract: Designers require different design spaces across creative stages: broad during exploration, and targeted during refinement. Yet existing agent-driven tools assume a fixed or continuously expanding space, leaving designers to manage and navigate it themselves. Informed by a formative study with five designers, we propose an axis-centered workflow that adaptively broadens and narrows the design space to support structured exploration and refinement. We implemented this workflow in Surprise2Refine, a prototype that allows users to build and reshape an nxn design space through a set of axis-centered interactions as their creative intent evolves. A within-subjects study with 14 designers shows that Surprise2Refine enhances users' sense of control, supports tracking of scaffolding paths, and improves the perceived creativity of design outcomes. We further distill design insights to guide future agent-assisted tools for creative scaffolding.
arXiv:2608.12582v1 Announce Type: new Abstract: AI journaling tools can tailor prompts to a person's own sensed behavior, but it is unclear which behaviors respond to them. We analyzed 369 journal entries from an eight-week passive sensing study. An LLM labeled each entry as expressing an intention to change a behavior or not, and we measured follow-through against 26 sensor features with a 3-day before/after comparison. Responsiveness depended most on whether a behavior involves other people. Behaviors that depend on others improved in only 15 to 22% of cases, while behaviors a person can act on alone improved more often, up to 50 to 63%, though unevenly. How users wrote mattered less. No single text feature separated improved from unimproved entries; writing carried signal only within specific behaviors, most clearly for text messaging and for longer, more personal intention entries. The sample is small, so we treat these as exploratory patterns that point to where AI journaling nudg
arXiv:2608.12580v1 Announce Type: new Abstract: Textiles are increasingly explored as media for interacting with digital information. However, many of the existing approaches rely on visible tags, printed overlays, or electronic modules that compromise the fabric's aesthetic and tactile qualities. To address this, we present InvisIto, a method for weaving visually unobtrusive yet machine-readable infrared markers directly into fabrics using near-infrared (NIR)-absorbing yarns. Although these yarns look similar to standard fibers in ambient light, they produce strong contrast in NIR imaging. Our method includes: (1) a design tool that helps users easily embed infrared markers into weaving drafts, (2) five disguising strategies that further reduce marker visibility under ambient light, and (3) a camera-based detection pipeline for decoding and tracking the woven markers. InvisIto supports both woven QR codes for data encoding and woven ArUco markers for binary input and deformation track
arXiv:2608.12552v1 Announce Type: new Abstract: Norway is among the most digitalized countries in the world, where access to essential services increasingly depends on digital systems. Although universal design of ICT is legally required across public and private sectors, ensuring cognitive accessibility for older adults involves more than technical compliance. We analyzed responses from 294 participants aged 55 to 90 to examine the barriers they encounter when using digital services. Our findings identify four tensions, namely navigational, semantic, procedural, and temporal, which reveal how accessibility barriers emerge through misalignments between service design, system assumptions, and older adults' capabilities. Together, these tensions show that checklist-based compliance does not necessarily translate into lived accessibility when engaging with mandatory or near-mandatory digital services. We discuss the need to move beyond minimum compliance frameworks toward accessibility ap
arXiv:2608.12548v1 Announce Type: new Abstract: Dance imitation integrates motor planning, sensorimotor integration, and social cognition, offering a sensitive framework to characterize motor behavior in autism. In this work, we explore a computational analysis framework to identify potential biomarkers that allow the design and development of improved medical and human-machine systems. We analyzed 3D motion capture data from autistic and neurotypical adults performing dance imitation under solo and socially-framed duo conditions. Methodologically, using Dynamic Time Warping, we quantified movement consistency and propose the Social Context Sensitivity Index (SCSI) to measure modulation of variability by social framing. These features were then used on a classifier to discriminate subjects into autistic or neurotypical groups. Results show that neurotypical adults exhibited increased movement variability in socially-framed imitation, especially in upper and lower limbs, whereas autisti
arXiv:2608.12546v1 Announce Type: new Abstract: At the studied research institute, one professorship oversees approximately 20 theses per semester, while day-to-day supervision is distributed among doctoral and postdoctoral researchers. To manage this supervision demand, the institute uses an expos\'e-first workflow in which students prepare a research proposal before entering the main thesis-writing phase. This paper asks how students, supervisors, and administrators experience the expos\'e-first workflow as a structured process for early thesis preparation, and how it redistributes responsibility, supervision, and administrative coordination work across roles and two digital platforms. Based on a mixed-methods study analyzed through Frauenberger et al.'s four reflective design lenses, the findings show that the expos\'e-first model made thesis preparation more structured by turning early research planning into a staged process of proposal writing, feedback, and approval. Students rep
arXiv:2608.12370v1 Announce Type: new Abstract: There has been limited clarity and consistency regarding what the term infographics, or information graphics, refers to in visualization research and practice. In particular, little is understood about where people's conceptualizations of infographics converge or diverge. To address this gap, we conducted a systematic literature review and a practitioner survey to identify and contrast different perspectives on infographics. We performed inductive coding on 487 sentences from 111 visualization and HCI papers, and analyzed questionnaire responses from 44 domain practitioners. Our findings reveal recurring dimensions of conceptual disagreement on infographics: role of text, relationship with data visualization, and relationships with statistical charts and data comics. Based on the results, we recommend future work to explicitly report any assumptions made along these diverging conceptualizations, to investigate the cognitive origins of the
arXiv:2608.12358v1 Announce Type: new Abstract: Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role. This paper introduces Interaction Readiness as a framework for specifying and evaluating that missing layer of performance. The framework separates content specifications, which govern what an agent knows and says, from interaction specifications, which define how an agent should conduct itself in a role-governed exchange. Interaction specifications require teams to define role purpose, authority boundaries, recurring situations, boundary cases, repair behaviors, and audit criteria before deployment. We operationalize interaction readiness through four agent operations: understanding purpose, calibrating authority, managing tone, and repairing breakdowns. Using StudyChat, a public dataset of student interactions with an AI tutorin
arXiv:2608.12357v1 Announce Type: new Abstract: I think the creative process in conducting open-ended creativity-adjacent research is itself part of the creative process (including ideation, pivots, and tangents excluded from the final work) that deserves greater emphasis. how might we bridge multiple open-ended projects and serve broader non-technical audiences? We may choose to emphasize interaction techniques rather than singular tools in isolation. We may grow a community based on our shareable experiences in creativity interactions, to draw more complete pictures of how we creatively solve creativity problems (or satisfy curiosities). This may lead to new directions. I briefly elaborate on this using the example of "DrawTalking" a drawing+talking interactions work. DrawTalking itself resulted from a winding creative process exploring spontaneous interactive world-building and storytelling when drawing and talking.
arXiv:2608.12355v1 Announce Type: new Abstract: Recent progress in AI coding agent research has led to rapid improvements in agents' ability to autonomously perform complex software engineering tasks, from editing large codebases to executing long-horizon development workflows. As these systems make strides, however, the primary bottleneck to practical usefulness increasingly shifts away from pure task-solving capability, and toward challenges in how users communicate with, supervise, and trust agents. In this position paper, we argue for a reorientation from autonomous to human-centered coding agents: systems designed not only to complete tasks, but to collaborate effectively with people. We identify four core interaction-level dimensions that characterize the human-agent task-solving loop: task alignment, verifiability, steerability, and adaptability. Finally, we outline concrete research directions to advance these dimensions, including user-involved coding environments, comprehensi
arXiv:2606.18479v2 Announce Type: replace-cross Abstract: Reject inference methods are widely used to mitigate survival bias in credit scoring, yet their effectiveness remains poorly understood. We systematically evaluate several such methods and uncover a structural failure mode: in a natural retraining cycle, models whose accuracy improves while recall collapses create an illusion of improvement that leads practitioners to believe the system is getting better when, in fact, its rejection quality -- the ability to correctly screen out defaulters -- is deteriorating. We then propose a controlled exploration strategy that breaks the feedback loop without statistical assumptions: the lender deliberately approves a fraction of rejected applicants and observes their true outcomes. We show that accuracy and rejection quality give opposite recommendations on whether to explore: accuracy favors no exploration, while rejection quality improves with it, confirming that standard evaluation metri
arXiv:2603.29545v2 Announce Type: replace Abstract: Existential risk scenarios relating to Generative Artificial Intelligence often involve advanced systems or agentic models breaking loose and using hacking tools to gain control over critical infrastructure. In this paper, we argue that the real threats posed by generative AI for cybercrime are rather different. We apply innovation theory and evolutionary economics - treating cybercrime as an ecosystem of small- and medium-scale tech start-ups, coining two novel terms that bound the upper and lower cases for disruption. At the high end, we propose the Stand-Alone Complex, in which cybercrime-gang-in-a-box solutions enable individual actors to largely automate existing cybercrime-as-a-service arrangements. At the low end, we suggest the phenomenon of Vibercrime, in which 'vibe coding' lowers the barrier to entry, but do not fundamentally reshape the economic structures of cybercrime. We analyse early empirical data from the cybercrime
arXiv:2603.00048v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed in sensitive applications including psychological support, healthcare, and high-stakes decision-making. This expansion has motivated growing research into the ethical and moral foundations underlying LLM behavior, raising critical questions about their reliability in ethical reasoning. However, existing studies and benchmarks rely almost exclusively on Moral Foundation Theory (MFT), largely neglecting other relevant dimensions such as social values, personality traits, and individual characteristics that shape human ethical reasoning. To address these limitations, we introduce MOSAIC, the first large-scale benchmark designed to jointly assess the moral, social, and individual characteristics of LLMs. The benchmark comprises nine validated questionnaires drawn from moral philosophy, psychology, and social theory, alongside four platform-based games designed to probe morally ambiguo
arXiv:2608.13100v1 Announce Type: cross Abstract: Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnerable to attacks that extract the assessment content itself through screenshots, screen sharing, optical character recognition, and automated scraping. This paper extends the Multi-dimensional Spatio-Temporal Context Camouflaging Model (MSCCM) within the MARS (Multi-modal Assessment Resilience Suite) by introducing the Multi-Layer Context Camouflaging Theory (MCCT), a mathematical framework that protects rendered assessment content through semantic superposition. Authentic assessment content and synthetically generated camouflage are represented as a unified rendering while remaining recoverable only by legitimate candidates. The framework models the adversarial extraction process through an explicit extraction-channel operator and develops six coupled constructs: the Context Inversion Operator, Contex
arXiv:2608.12816v1 Announce Type: cross Abstract: Large language models have begun refuting long-standing conjectures and, for a few thousand dollars of tokens, solving long-open problems (OpenAI, August 2026). The introspection this has prompted about the future of mathematical discovery is overdue, and the anxiety accompanying it legitimate -- but both are attached to the wrong loss. What machines now produce is the countable part of mathematics -- theorems, proofs, refutations -- which was always the \emph{residue} of the work, not its product. The product is human understanding: not a stock of results but a collective, hard-won way of deciphering the world and acting upon it. The two are arcs of a single loop -- looking produces the residue; taking it up again, one journey at a time, is what rebuilds the shared understanding. Machines are strong on the countable arc, absent from the one that feeds it. The peril is to leave the loop open. AI did not create the confusion between resi
arXiv:2608.12581v1 Announce Type: cross Abstract: Understanding when migration generates social integration or exclusion is a central challenge for urban communities. Existing research has mostly relied on surveys, administrative data, or aggregate indicators that fail to capture expressions of exclusion at fine spatiotemporal scales. Here, we analyze over 550,000 geolocated reports from SOSAFE (Chile's largest citizen reporting platform) to examine the relationship between migration and hate speech in Santiago. We fine-tune a Spanish hate speech classifier and validate it against human labels. Reports that mention migrants are more likely to contain hate speech than other reports. Hate speech concentrates in areas with recent demographic change (post-2010 arrivals) rather than in established migrant communities. The spatial analysis shows that hate speech hotspots coincide with neighborhoods where recent migrants comprise over a third of the population. Coldspots appear in high-educat
arXiv:2608.12372v1 Announce Type: cross Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully communicate their reasoning. We review evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data showing that many users find cognitive alignment "essential" when an AI's rationale for a judgment or action is important to them. We outline the gaps between existing alignment methods and what is needed to achieve cognitive alignment, and present a research agenda to address these gaps. We argue that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications, and that addressing it is important for creating AI systems on which users are both willing and
arXiv:2608.12363v1 Announce Type: cross Abstract: European countries are debating policies to mitigate the increased energy costs caused by renewed geopolitical tensions, while pursuing decarbonization and electrification. A notable example is Italy's 2026 Decreto Bollette package, which proposes to remove the carbon price equivalent from the bids of certain gas-driven power plants to wholesale electricity markets, among other provisions. We use this as a case study to assess the long-term implications of suppressing the carbon price signal in the electricity market for investment, emissions, and consumer costs. We employ a stylized Italian power system using MARLEY, a multi-agent reinforcement learning framework focused on long-term electricity market assessments. In this framework, we test this policy across configurations with varying levels of support for green investment, resource adequacy, and flexibility. Results show that partial suppression of the carbon price signal yields sh
arXiv:2608.12361v1 Announce Type: cross Abstract: Neologisms, emerging terms in meaning or form, can serve as new vehicles for toxic expression, like "country girl" as a stigmatizing label targeting feminism. Such toxic neologisms appear benign but have evolved into toxic usage in public consensus, posing challenges to moderation systems and remaining underexplored. In this paper, we investigate how to detect implicit toxicity expressed via neologisms. We first propose a taxonomy that captures the origins and consensus-verification criteria of toxic neologisms, followed by the construction of a lexicon spanning widely observed risk categories. To capture toxicity grounded in public consensus, we introduce SeTox, a search-augmented framework that enables static large language models (LLMs) to incorporate real-time web context for neologism toxicity detection. Experiments show that SeTox, even with 3B-scale models, outperforms recent large-scale models, demonstrating its scalability to i
arXiv:2608.12353v1 Announce Type: cross Abstract: Infrastructure scholarship in CSCW often treats breakdown as the moment when infrastructures become visible. However, in vendor-managed sociotechnical systems, not all breakdowns become visible to actors who have the capacity to repair them. Drawing on a retrospective qualitative study of a Chinese K-12 EdTech deployment, including 11 interviews, 5 classroom observations, and more than 28 days of field notes, this paper introduces visibility asymmetry: a sensitizing concept for understanding how similar local breakdowns encounter uneven conditions for being routed, recognized, and acted upon. The analysis traces a four-stage mechanism through which procurement categories sort schools into attention tiers; staffing and visit cadence follow those tiers; only some local problems travel through staff or administrator channels; and dashboards can re-code unresolved repair labor as evidence of adoption. By shifting attention from the occurren
arXiv:2608.12349v1 Announce Type: cross Abstract: The affordances of a creative medium strongly condition the creative artefacts the medium will produce. In this work, we present a formalisation of computational creativity (CC) media using the conceptual toolbox of complex systems (CS). We introduce the notions of emergence, collective intelligence and self-organisation, non-linear dynamics, criticality, multi-scale hierarchy, phase transitions, diversity of attractors, path dependence, and open-endedness, and connect them to the existing CC literature. Together these nine properties form a vocabulary with which creative media can be described and compared at the system level, while medium affordances are the design-level mechanisms that determine each medium's complex system properties. The formalisation emphasises the influence of each medium's affordances in determining what the medium can produce in creative processes. To demonstrate the proposed theoretical approach, we characteri
arXiv:2608.12346v1 Announce Type: cross Abstract: This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a "perfectly aligned" model inadvertently also provides malicious actors with an ever-improving tool for informational dominance. We need to discuss this dual-use potential now, as its risk is exacerbated by rapid user adoption of AI as information provider, economic power asymmetries, and a political landscape that increasingly shifts towards authoritarianism. We conclude by urging the community to consider the intentional misuse of AI alignment mechanisms and propose mitigation strategies to safeguard against this dual-use potential.
arXiv:2608.12344v1 Announce Type: cross Abstract: We test whether the perceived attributes of a consumer technology predict how widely it is owned. In a 2022 Prolific survey of US adults (n = 678), respondents rated 65 consumer technologies on six attributes. We then elicited the same ratings from two frontier language models, Anthropic Claude Opus 4.7 and OpenAI GPT-5.5. We regress ownership prevalence on four UTAUT2 acceptance attributes plus a log-age covariate with a sign-constrained penalized regression and evaluate it by holding out one technology at a time. The attribute model improves on a baseline of years-since-launch: mean absolute error falls by 17% with the human ratings, and by more with either model, most with Opus 4.7. Over the short 2022-to-2025 window, where ownership moved little, the same attributes do not improve on a no-change baseline. We set out the limitations of the approach, including the possibility that language-model ratings reflect prior knowledge of thes
arXiv:2608.12323v1 Announce Type: cross Abstract: Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate that this enforcement information paradox systematically occurs in AI agents. While most AI safety evaluations test whether models fail, we investigate why, applying compliance theory from law and economics as a diagnostic tool. We treat compliance theories not as metaphors but as empirical hypotheses and show that each predicts the behavior of a distinct model class. We evaluate our hypotheses across twelve instruction-tuned language models operating as enterprise procurement chatbots. Drawing on theories of deterrence, legitimacy, and expressive law, we show that safety-fine-tuned models maintain compliance broadly, while task-optimized and agentic models treat regulatory signals as mere optimization parameters. These latter models fail to comply under conditions predicted by theory, such as low enfor
arXiv:2608.13444v1 Announce Type: new Abstract: Machine learning ethics researchers and critical HCI scholars have argued that algorithmically predicting gender is wrong. At the same time, other researchers rely on predicted gender labels to study gender disparities and develop algorithmic fairness techniques. How do we reconcile these two seemingly contradictory intuitions? We differentiate two ways gender prediction may be wrong: being illegitimate, thereby contributing to harm; and being invalid, thereby producing unusable measurements. Our analysis translates arguments against gender prediction into these terms of legitimacy and validity and shows how gender imputation applied for fairness purposes can be illegitimate yet still yield valid disparity measurements. We clarify this bind by drawing upon transfeminist literature to distinguish sexism that targets women and femininity from sexism that targets transgender and nonbinary people. While gender imputation can produce valid mea
arXiv:2608.13369v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used by laypeople to resolve real legal problems, against a backdrop of persistent access-to-justice deficits. This article presents evidence that the practical force of AI-generated legal advice depends not on its accuracy but on the social production of its credibility. While existing research has assessed the accuracy of legal AI, less is known about how machine-generated guidance is verified and made credible enough for lay users to act on. Drawing on a dual-method analysis of 153 Reddit narratives and 5,341 community reactions, this article maps a spectrum of verification practices. At one end, a minority of users verify AI-generated legal advice by triangulating across models, and some submit AI-generated guidance to platform communities for evaluation before acting, a configuration we term distributed counsel. Far more commonly, however, narratives are silent on verification. AI-generat
arXiv:2608.13351v1 Announce Type: new Abstract: This study investigates the use of a Learning Management System (LMS) to support self-paced learning at a South African Public Access Centre (PAC), using the I-CAN Centre as a case study. Through semi-structured interviews with thirty-eight learners and thematic analysis, the research explores opportunities and challenges associated with LMS adoption. Findings reveal that PACs play a critical role in promoting ICT skills and digital inclusion, offering learners flexible access to learning resources and fostering empowerment. While LMS use enhances convenience and supports blended learning as the preferred approach, persistent challenges, such as poor connectivity, outdated infrastructure, and unclear course instructions, limit its effectiveness. These findings highlight the need for infrastructural upgrades and user-centric design to optimise LMS implementation in community-based learning environments.
arXiv:2608.13250v1 Announce Type: new Abstract: Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather than neutral moral knowledge. We propose treating the AI system as a proxy actor and test whether dataset-level norms can shift it away from its baseline safety behavior when it faces high-conflict dilemmas. We make three contributions. First, we demonstrate in controlled experiments that norm-breaking fine-tuning yields norm-divergent actions justified by self-interested rationales, suggesting a systematic shift in patterns of justification. Second, we establish a practical audit trail linking downstream justifications to upstream norms using mixed methods. Third, we show that system prompts can both suppress and elicit these patterns. We conducted experiments on three models (LLaMA-3.2-11B, Qwen-3.5-9B, and Pixtral-12B) using Low-Rank Adaptation (LoRA) fine-tuning on Social Chemistry 101 Fairness/Che
arXiv:2608.13022v1 Announce Type: new Abstract: Algorithmic fairness evaluation commonly assesses AI systems as bounded technical components, abstracting away the organizational context in which they operate. We present, to our knowledge, the first independent end-to-end fairness audit of a semi-automated hiring system operated by Barcelona Activa, a public employment agency using the third-party TalentClue platform for candidate search and shortlisting. We analyze approximately 497,000 candidate-vacancy pipeline entries from September 2017 to September 2022, covering seven pipeline stages that span automated processing, human discretion, candidate data, and employer decisions. Aggregate outcomes across binary genders are statistically indistinguishable, yet this parity masks substantial disparities by salary level, age, and gender identity. Women face adverse impact in mid-salary shortlisting (DIR = 0.786, p < 0.001), alongside salary disparities in 15 of 20 sectors and a compounded d
arXiv:2608.12924v1 Announce Type: new Abstract: Despite the recent intensive development of secondary education curricula and assessments in informatics, the impact of assessments has not been well studied in this field. Since informatics education covers a diverse range of content, from computer science knowledge to ICT skills, careful consideration is needed to prevent assessments from distorting education. This study investigates the impact of introducing ``Informatics I'' into the Common Test for University Admissions in Japan, as an example of a large-scale, standardized, high-stakes assessment in 2025. As the data source for this analysis, this study uses a questionnaire that has been administered every year from 2006 to 2026 to all first-year students at the University of Tokyo. The questionnaire asks students for their self-perceptions of the information-related knowledge and skills they studied and acquired in high school. Using these data, we conduct a longitudinal study of t
arXiv:2608.12768v1 Announce Type: new Abstract: Generating multi-attribute synthetic populations with realistic joint distributions and geographic variation is a foundational requirement for geo-simulation techniques, such as micro-simulation and agent-based modeling. However, it remains challenging for existing methods to reconstruct region-specific joint distributions from aggregated-level data alone. Thus, we propose a hierarchical diffusion-based generative framework that utilizes a realistic region-specific joint distribution of multiple attributes as the training target to create a synthetic population along with assigning their explicit home and work locations. Applied to 50 U.S. states and Washington, D.C., this framework generates a nationwide geographically-explicit synthetic population consisting of 332,387,543 individuals with five attributes (e.g., age, gender, employment, education, income). Held-out regional experiments show improved reconstruction of joint distributions
arXiv:2608.12669v1 Announce Type: new Abstract: The fair AI/ML literature has long distinguished distributive fairness, concerning how automated systems allocate resources and opportunities, from representational fairness, concerning how they shape the ways individuals and social groups are perceived, understood, and accorded social status. Generative AI is rebalancing these normative dimensions. Unlike predictive systems, large language models (LLMs) and related technologies are fundamentally expressive: their primary function is to convey meaning rather than automate domain-specific decisions. Representational harm has also become central to value alignment, especially in research on what and whose values and perspectives AI systems should represent. Existing approaches to harms in the representation of social groups often appeal to descriptive accuracy, but this strategy has important limitations. For many social groups, no stable or bounded referent exists against which representat
arXiv:2608.12649v1 Announce Type: new Abstract: Computing systems are moving from reactive tools toward systems that sense, interpret, predict, and act before explicit user requests. This transition is enabled by the global scale of mobile connectivity, the rapid expansion of wearable and ambient sensing, advances in machine learning and foundation models, distributed edge infrastructure, and physical actuation. We define \emph{proactive computing} as a paradigm in which systems infer user context, anticipate future needs or risks, and initiate information delivery or actions at an appropriate time. This survey distinguishes proactive computing from reactive, context-aware, adaptive, and predictive computing, and frames proactivity as a system-level integration problem across sensing, understanding, decision making, action, and governance. We review the technological enablers of proactive computing, organize its design space, analyze technical challenges such as uncertainty-aware trigg
arXiv:2608.12362v1 Announce Type: new Abstract: Enabling students to develop systematic problem-solving strategies is a central goal in computing education and of particular relevance in the emerging field of machine learning (ML) education. While exploratory approaches are common in ML learning tasks, fostering the development and persistence of structured problem-solving strategies remains challenging, as these demand considerable metacognitive regulation and persistence, causing learners to often revert to exploratory trial-and-error behavior. To address this challenge, we augmented a digital puzzle-based learning game for decision tree construction with an adaptive feedback module generating individualized messages based on the continuous evaluation of learners' problem-solving strategies. Building on an earlier baseline study, the present work investigates how this strategy-oriented feedback shapes students' problem-solving processes. For this purpose, screencast video data and ga
arXiv:2608.12360v1 Announce Type: new Abstract: Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks. As these systems become more integrated into clinical decision-making, there is growing expectation that they demonstrate key dimensions of trustworthy AI to support clinician, patient, and public trust. Whether publicly available regulatory documentation provides sufficient evidence to independently assess the trustworthiness of cleared AI systems remains unclear. Methods: We analysed FDA AI/ML-enabled medical device summary reports published between 2021 and 2025. Reports underwent automated keyword screening followed by multi-stage manual consensus review to identify documented evidence for the six FUTURE-AI principles: Fairness, Universality, Traceability, Usability, Robustness, and Explainability. Descriptive, temporal, and clinical-domain analyses were performed. Multivariable logistic regression assessed whether y
arXiv:2608.12359v1 Announce Type: new Abstract: Nutritional labels are legally permitted to appear in very small print, reducing real-world readability and encouraging consumers to rely on 'AI nutrition lens' and vision-capable conversational agents for dietary guidance. We evaluate whether such AI-mediated advice can meaningfully substitute for regulated labeling using a bounded, verifiable task: inferring which of two packaged foods contains less sugar from front-of-pack images alone. A Two-Alternative Forced Choice game was used to evaluate AI agent systems across four national supermarket contexts: Sweden, the USA, Australia, and Kazakhstan. The results (N=132 comparisons) across both agents reveal a significant performance divide contingent on context. For global products the agents achieved 88.9% accuracy (p < 0.0001 against chance). For local products (Sweden), accuracy dropped to 59.5% (p = 0.29), rendering the AI's guidance statistically indistinguishable from random guessing.
arXiv:2608.12356v1 Announce Type: new Abstract: A college offering several overlapping computing degrees implicitly assumes that its programs are differentiated in line with how the labor market segments computing work and that, together, they prepare graduates for that market. Testing this is difficult, because the instruments available to curriculum committees, namely advisory boards, tracer studies, and employer surveys, are slow, narrow, and hard to reproduce. We apply one uniform, taxonomy-anchored alignment analysis across all five undergraduate programs of a College of Information Technology, comparing 1,922 course learning outcomes against 103,349 competencies extracted from a unified corpus of 5,186 deduplicated job openings from four boards. Every competency is obtained by a grounded single-language-model procedure that copies it verbatim from the source and verifies it against the source, then assigns it to one of eleven ESCO-aligned domains and a Bloom cognitive level; the
arXiv:2608.12352v1 Announce Type: new Abstract: AI governance frameworks can be known, used, and implemented in form without becoming governance in practice. This paper examines that problem through a role-based stress test of the NIST Artificial Intelligence Risk Management Framework (AI RMF) in consumer lending. We treat framework adoption as a governance translation problem: whether RMF language can become role-usable, cross-level, authority-connected governance over the AI system-in-use, rather than producing governance-looking artifacts. The study uses LLM-based role simulation as a structured analytic probe. We apply a 4 $\times$ 2 $\times$ 3 design across four organizational roles, two AI deployments, and three governance hard cases, producing 120 scored responses. Results show that local translation was not the main problem. Simulated actors generally understood their assigned roles and translated the RMF into local activity. The harder problem was whether that activity became
arXiv:2608.12351v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be produced with little effort. This paper reports lessons from designing and implementing an AI-aware, AI-testing assessment in a large second-year undergraduate database systems module. The design combined two linked elements: (1) a structured three-part response format (X1-X2-X3) in which students documented a sourced answer, produced their own answer, and evaluated the sourced output; and (2) an AI-aware question-design process in which draft tasks were stress-tested against contemporary GenAI tools and revised when generic prompting produced superficially adequate answers. The account draws on archived assessment materials, rubrics, planning records, design-time GenAI trials, practice-response data, attainment records, and external review comments. Its main contribution
arXiv:2608.12350v1 Announce Type: new Abstract: The energy demand growth and environmental impacts of artificial intelligence (AI) have generated substantial interest in supplying sufficient low-cost electricity for AI-driven data center development. Research on the ability of demand-side management to address these challenges has been more limited. Shifting the amount or timing of demand from retail, corporate, and other organizational behaviors is a plausible option but only if changes in demand-related behavior have important effects on the envi- ronmental and electricity effects of AI. This article tests four retail (i.e., consumer) user behaviors with high behavioral plasticity to assess their technical abatement potential. The research concludes that non- reasoning models provide sufficient quality while consuming close to one-twentieth of energy compared to reasoning models, saving an amount equal to the annual electricity requirement of at least 141,000 US households under dail
arXiv:2608.12324v1 Announce Type: new Abstract: People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about real disagreements among faithful traditions, some require humility because the issue is prudential, and some are pastoral situations where safety and human referral matter more than theological completeness. Existing benchmarks do not evaluate this structure. We introduce FMG-Bench, the Faith & Moral Guidance Benchmark, a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral guidance contexts. FMG-Bench v1 evaluates 14 advanced models across 8,792 scored responses, comparing raw model behavior with three guided instruction settings. In our production run, placing models inside a structured harness improves over raw model behavior by +3.96 po
arXiv:2608.12320v1 Announce Type: new Abstract: This article reviews and updates the framework for accountability in AI based on account- ability ecosystems. We update the framework in light of the latest developments since the release of Large Language Models for general public use. We propose three interlinked updates to the original AI accountability ecosystem: (i) reorienting the accountability ecosystem to AI infrastructure and supply chains, (ii) providing greater emphasis on outcomes monitoring and identification of issues that support decentralized system improvement, and (iii) incorporating end-user accountability given the new risks of unpredictability of language models in-the-wild. Collectively, these updates mark a shift towards accountability as distributed, continuous, and institutionalized, away from a system in which frontier AI applications can be modeled as discrete products controlled by single identifiable actors with industry-specific oversight.
Article URL: https://www.youtube.com/watch?v=SgxpxSjD8zA Comments URL: https://news.ycombinator.com/item?id=45222312 Points: 1 # Comments: 0
Artificial intelligence promises big gains for faculty in higher education, including greater efficiencies and elevated learning outcomes. To realize the wins, professors need to get up to speed on the tools. While many are experimenting on their own, some institutions are taking steps to accelerate that learning. At Ventura College, a California community college, leaders recently stood up communities of practice around AI use. A CoP brings together individuals with a shared interest in a topic or technology; in this case, AI. The group then works together to learn more about the topic or…
In a K–12 setting, deepfakes hold a lot of power. These falsified images or videos, virtually impossible to identify with an untrained eye, can be wielded to harm educators’ reputations, cyberbully vulnerable students, and blackmail individuals and schools. With artificial intelligence image generation, the problem is growing rapidly. Super-realistic images can be created quickly and deployed easily, creating a concerning scalability. Faced with the malicious use of AI-generated images — both of students and school officials — leaders must redouble their efforts around deepfake detection,…
Lately, school-related data breaches seem to keep coming. PowerSchool and Canvas made major headlines this year. Countless smaller incidents may not hit the news, but they disrupt instruction and expose sensitive student data just the same. For K–12 IT leaders, threats to their district are inevitable. The question is whether their teams will be ready when those threats materialize. After years of conducting maturity assessments, working alongside district security teams and witnessing the aftermath of incidents, we can say with confidence that most districts aren’t there yet — not because…
Innovative Leader Award - Kimberly Zajac discusses why digital accessibility is important beyond compliance
Key points: The demographic cliff higher education has been warned about for years isn’t coming; it’s already here. The post-2008 ... Read more The post Why the old enrollment playbook no longer works appeared first on eCampus News .