Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
Comments URL: https://news.ycombinator.com/item?id=48341278 Points: 8 # Comments: 3
Article URL: https://www.ezducate.ai/ Comments URL: https://news.ycombinator.com/item?id=49486557 Points: 1 # Comments: 0
Article URL: https://www.edcapit.com/2026/08/17/future-of-edtech-silicon-valley-interview/ Comments URL: https://news.ycombinator.com/item?id=49485993 Points: 2 # Comments: 0
Article URL: https://thunderseethe.dev/posts/compiler-education-deserves-a-revoluation/ Comments URL: https://news.ycombinator.com/item?id=48694117 Points: 3 # Comments: 0
Article URL: https://www.thehindu.com/news/national/dharmendra-pradhan-quit-resign-as-union-education-minister-jantar-mantar-protest-july-25-2026/article71265711.ece Comments URL: https://news.ycombinator.com/item?id=49046523 Points: 9 # Comments: 1
Drug spending continues to grow, but that doesn’t mean that everyone is getting the medicines that they need. A panel at MedCity News’ Bullseye event discussed the evolution of the prescription drug marketplace. The post A Prescription for Drug Costs: Inject Transparency and Improve Accessibility appeared first on MedCity News .
Health tech investors say the founders who win them over aren’t the ones with all the answers, but the ones who know exactly what they don’t know — and build a team to cover the gap. The post What Health Tech Investors Love in Founders — and What They Can’t Stand appeared first on MedCity News .
Article URL: https://text.npr.org/nx-s1-5820922 Comments URL: https://news.ycombinator.com/item?id=48251247 Points: 3 # Comments: 0
Article URL: https://thezvi.substack.com/p/childhood-and-education-19-letting Comments URL: https://news.ycombinator.com/item?id=48249123 Points: 2 # Comments: 0
Epic faced an FTC antitrust probe this week, but that didn’t stop the EHR giant from unveiling a wave of new AI tools at its annual conference. The post This Week In Epic News: FTC Probe, AI Agents & More appeared first on MedCity News .
Article URL: https://apnews.com/article/afghanistan-girls-excluded-secondary-education-taliban-618c7f5c7659d43ede2a3f71a7cf845f Comments URL: https://news.ycombinator.com/item?id=49395364 Points: 11 # Comments: 1
Article URL: https://novayagazeta.eu/en/articles/2026/06/19/russia-no-longer-needs-so-many-graduates-countrys-education-minister-warns-en-news Comments URL: https://news.ycombinator.com/item?id=48612022 Points: 13 # Comments: 0
Article URL: https://www.afterbabel.com/p/edtech-tragedy Comments URL: https://news.ycombinator.com/item?id=43736442 Points: 5 # Comments: 0
Article URL: https://en.wikipedia.org/wiki/Minimally_invasive_education Comments URL: https://news.ycombinator.com/item?id=49313881 Points: 3 # Comments: 0
Article URL: https://www.aei.org/commentary/how-a-few-foundations-shape-american-higher-education/ Comments URL: https://news.ycombinator.com/item?id=49313386 Points: 1 # Comments: 0
Article URL: https://gadlevanon.substack.com/p/job-recession-in-higher-education Comments URL: https://news.ycombinator.com/item?id=49313371 Points: 17 # Comments: 2
Article URL: https://calendly.com/phonologie Comments URL: https://news.ycombinator.com/item?id=35932703 Points: 1 # Comments: 1
Article URL: https://stemteachingtools.org/brief/109 Comments URL: https://news.ycombinator.com/item?id=48517450 Points: 3 # Comments: 0
Article URL: http://surgeryacade.my/ Comments URL: https://news.ycombinator.com/item?id=8443037 Points: 2 # Comments: 0
Article URL: https://github.com/withmarbleapp/os-taxonomy Comments URL: https://news.ycombinator.com/item?id=48868607 Points: 1 # Comments: 0
Article URL: https://ijpp.com/competency-based-medical-education-in-india-a-work-in-progress/ Comments URL: https://news.ycombinator.com/item?id=34750463 Points: 2 # Comments: 0
Hey Folks, I have penned down my thoughts on EdTech 2.0. This is an idea, I am exploring pursuing as my next entrepreneurial venture. I would appreciate feedback and suggestions. Link - https://open.substack.com/pub/monkeylike/p/edtech-20 Thanks in advance. Comments URL: https://news.ycombinator.com/item?id=42985788 Points: 1 # Comments: 0
Article URL: https://khanted.org Comments URL: https://news.ycombinator.com/item?id=49219963 Points: 1 # Comments: 0
Article URL: https://innig.net/teaching/liberal-arts-manifesto Comments URL: https://news.ycombinator.com/item?id=49131034 Points: 45 # Comments: 67
In this episode, we’re joined by Amanda Shafton, practicing CNM and national director of midwifery at Ob Hospitalist Group. We discuss the benefits of physician/midwifery collaboration and how midwives can improve maternal health outcomes. The post MedCity FemFwd: How Midwives Can Support the Maternal Health Crisis appeared first on MedCity News .
Health tech companies made several major funding announcements in August. Here is a list of some of the biggest funding rounds. The post 4 Notable Health Tech Funding Announcements in August appeared first on MedCity News .
Takeda Pharmaceutical’s Mimrylo is now approved for treating polycythemia vera, a rare blood cancer with limited therapeutic options. Originally developed by Protagonist Therapeutics, the engineered peptide gives Takeda a drug with blockbuster sales potential as the pharma company’s top overall product faces patent expirations. The post First-in-Class Takeda, Protagonist Drug Lands FDA Approval in Rare Blood Disorder appeared first on MedCity News .
Amazon Web Services recently announced the launch of Student Rewards, offering free educational and certification opportunities to verified university students. The company has committed more than $500 million in resources to support workforce development via artificial intelligence and cloud computing education for students enrolled in college. Verified students over the age of 18 can receive premium access to the AWS Skill Builder platform, $30 in AWS credits and a $100 Certification Exam voucher. As employer demand for AI and cloud computing skills grows, this program bridges the gap…
If AI is not one of the CEO’s top two or three priorities, the organization will produce pilots and slide decks while better-mobilized peers pull ahead and the institution’s ability to fulfill its mission erodes The post AI Won’t Transform Health Systems, CEOs Will — 6 Principles to Drive the Transformation appeared first on MedCity News .
Article URL: https://edtechdev.github.io/aied/ Comments URL: https://news.ycombinator.com/item?id=49509073 Points: 2 # Comments: 1
One of the major barriers to artificial intelligence adoption across industries is the idea of displacement. Will AI eventually be able to replace humans? Will automation render our skills — and our jobs — obsolete? Most of us understand that however powerful this technology might be, it still requires human oversight to ensure it’s working properly. The question now is how we can collaborate with this technology in ways that amplify the value of us both. DISCOVER: Organizations are creating frictionless digital experiences for employees. Trusting in Technology Is Key to Harnessing Its…
New edtech products that have caught our attention this month
A single workshop or training will not adequately prepare educators for everything they need to know and do in regard to AI.
As retention becomes one of the most important drivers of institutional stability, many colleges and universities are turning to AI to better identify struggling students, personalize support, reduce administrative burdens, and help staff intervene before challenges escalate. The post The hidden link between AI strategy and student retention appeared first on eCampus News .
arXiv:2608.24334v2 Announce Type: replace-cross Abstract: Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for reconstruction and do not explicitly allocate capacity according to semantic role. Action-level meaning and fine-grained kinematic detail must therefore be encoded through the same reconstruction-driven hierarchy. We introduce SeMoCo, a semantic-first motion codec, together with a dual-axis motion generator for language-conditioned motion generation. Each motion token contains one semantic token and a residual sequence of kinematic tokens. The generator models semantic progression across time and autoregressively refines the residual entries. We also construct $\Omega$-MotionVerse, a large-scale, multi-source human-motion dataset unified under the SOMA representation. Across the reported comparisons, SeMoCo achieves the best reconstruction accuracy among the compared codecs, whil
arXiv:2608.23873v2 Announce Type: replace-cross Abstract: Everything a language model sees is tokens. The serving stack knows what each span is -- user input, tool output, instructions -- but the model must keep track of that itself, and can lose track or be confused: text can be written to read like anything. Prompt injection is a natural exploit of this phenomenon. By scrambling the model's understanding of span identity, an attacker can induce unwanted and dangerous actions. Adding a non-textual channel to the model's input -- a way to communicate span identity beyond text -- mitigates this class of attack. We thus introduce a general steering technique called Semantic Overlays: small learned adapters applied at chosen prefill positions to a frozen model's residual stream. Laying an overlay over a span creates an out-of-band annotation channel that cannot be replicated by tokens. Unlike steering vectors, Semantic Overlays are trained, adaptable, and selectively applied. An overlay c
arXiv:2608.02751v3 Announce Type: replace-cross Abstract: Existing deep-search agents use a Search-Visit workflow that retrieves whole webpages without considering the structure they expose through titles, headings, sections, and metadata. This prevents agents from directly constraining retrieval to parts of a webpage and often carries irrelevant content into their context. We introduce Sieve, a search-inspect-fetch strategy driven by a Boolean Query Language (BQL): it searches webpage fields to filter candidates, uses an interchangeable ranker to order them, presents structure-rich result cards for inspection, and fetches only selected sections. Across three QA collections, Sieve is more accurate than the strongest conventional Search-Visit configuration on each collection while using 20.7-50.6% fewer tokens. Boolean filtering improves every tested ranker, and the accuracy-context advantage persists across retriever choices and agent backbones. Our implementation is included in the Sk
arXiv:2607.13396v2 Announce Type: replace-cross Abstract: What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow the notion of set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts. Our cognitive test for LLM agents mounts libraries of redundant tools and skills, in which many tools solve the same task but differ in hidden reliability. Using a branching schedule, we shift the reliable tool group in the environment and compare it with a stable control, allowing us to isolate the effect of each shift on the agent's behavior. We conduct our study on a panel of LLMs equipped with harnesses and show that the same set of shifts results in distinct behaviors across models: some latch onto a fixed routine within a few turns, whereas others continue to vary. Less capable models often omit the reliable tool group, while frontier models keep calling it alongside the other groups. We intro
arXiv:2606.19719v3 Announce Type: replace-cross Abstract: Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries. Standard practice evaluates these systems using PR-AUC, a metric that only measures how well scores rank and ignores whether they are usable at a fixed threshold. We show this mismatch leads to systematically poor deployment choices, as models with the highest PR-AUC are often the worst in operation. We introduce Precision--Cache Hit Ratio (P-CHR) AUC, a cache-aware metric that measures precision across cache utilization levels, and Operational Retention Rate (ORR), which captures how much offline ranking quality survives at deployment. We decompose the operational gap between offline and deployed quality into a recoverable threshold-utility component and an irreducible structural component fixed by the dataset's positive rate. Our experiments show that the threshold-utility gap is governed by the training objective rather tha
arXiv:2606.17092v2 Announce Type: replace-cross Abstract: Agentic systems are increasingly integrated with geographic information systems (GIS), where multi-agent coordination enables complex conversational and spatial analysis but introduces security risks. This work presents a security-oriented framework for risk identification, evaluation, and mitigation in a multi-agent GIS system while maintaining adaptability to broader agentic architectures. We test the agentic system of a commercial geospatial partner while developing a modular state-machine-based orchestration framework that abstracts agent behavior into reusable components. We evaluate robustness using a red-teaming framework with an adaptive attacker LLM and a deterministic judge that produces binary outcomes with supporting rationales across multi-turn attacks. We further improve resilience with a prompt optimization framework that treats prompts as structured signatures and injects adversarial demonstrations, enabling syst
arXiv:2606.05917v2 Announce Type: replace-cross Abstract: Long-video question answering remains challenging for Vision-Language Models (VLMs), as answer-relevant evidence is often sparse, transient, and temporally dispersed across lengthy video contexts. Existing frame-centric approaches improve efficiency through uniform sampling, query-aware frame selection, visual-token compression, and adaptive resolution strategies. However, they still rely on isolated and fragmented frames as the fundamental evidence units, limiting VLMs' ability to effectively capture coherent event-level semantics. To address this limitation, we propose MemoryCard, a video-memory-based augmentation framework that organizes long videos into self-contained Memory Cards. Specifically, MemoryCard first performs a self-reading process over videos and aligned utterances to segment the video into semantically coherent units, each corresponding to a distinct topic or event. For each unit, it generates an event-level vi
arXiv:2605.30434v2 Announce Type: replace-cross Abstract: Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested. We introduce LongDS, a benchmark for long-horizon, multi-turn data analysis where agents must maintain, update, restore, and compose evolving analytical states. LongDS comprises 68 tasks constructed from real-world Kaggle notebooks, spanning 2,225 turns across six domains including Geoscience, Business, and Education. Tasks are designed around state-evolution patterns (e.g., counterfactual perturbation, rollback, multi-state composition), with an average dependency span of 11.3 turns. Evaluating five state-of-the-art models, we find that the best model reaches only 48.45% average accuracy, performance drops nearly 47 points from early to late turns, and long-horizon errors account for 52%--69% of failures. Further a
arXiv:2605.28112v2 Announce Type: replace-cross Abstract: Federated Retrieval-Augmented Generation (FedRAG) is attractive for privacy-sensitive applications because full local corpora remain on clients. As a result, routing must rely on client-provided semantic profiles, creating a new opportunity for manipulation. We introduce Routing Hijacking, a routing-stage attack in which a malicious client forges its profile to attract target queries despite having irrelevant underlying data. We show that this vulnerability is severe. Across three representative FedRAG routing architectures, Routing Hijacking consistently misroutes target queries and leads to downstream disruptions and failures, including missing evidence, poisoning, incorrect answers, and hallucinations. In a controlled MedQA-USMLE stress test, we further show that poisoned retrieved evidence can mislead models across scales, leading to incorrect answers, hallucinations, and sycophantic failures. Existing defenses do not close
arXiv:2605.12015v3 Announce Type: replace-cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environments. However, this modularity introduces attack surfaces that are largely missed by existing safety evaluations: even when the user request is benign, unsafe influence may reside in skill guidance, local artifacts, or execution-environment files that steer the agent toward unsafe actions. We present SkillSafetyBench, a runnable benchmark for evaluating such skill-facing safety failures. SkillSafetyBench includes 155 adversarial cases across 47 tasks, 6 risk domains, and 30 safety categories, each evaluated with a case-specific rule-based verifier. Experiments with multiple CLI agents and model backends show that non-user attacks can consistently induce unsafe behavior, with distinct failure patterns across domains, attack methods, and scaffold-model p
arXiv:2512.12066v3 Announce Type: replace-cross Abstract: Current safety evaluations of large language models rely on single-shot testing, implicitly assuming that model responses are deterministic and representative of the model's safety alignment. We challenge this assumption by investigating the stability of safety refusal decisions across random seeds and temperature settings. Testing four instruction-tuned models from three families (Llama 3.1 8B, Qwen 2.5 7B, Qwen 3 8B, Gemma 3 12B) on 876 harmful prompts across 20 sampling configurations (4 temperatures x 5 seeds), we find that 18-28% of prompts exhibit decision flips--the model refuses in some configurations but complies in others--depending on the model. Our Safety Stability Index (SSI) reveals that higher temperatures significantly reduce decision stability (Friedman chi-squared = 396.81, p < 0.001), with mean within-temperature SSI dropping from 0.977 at temperature 0.0 to 0.942 at temperature 1.0. We validate findings acros
arXiv:2509.04438v3 Announce Type: replace-cross Abstract: Unified models (UMs) combine visual understanding (I2T) and generation (T2I) in a single framework. We focus on T2I and I2T, where cross-consistency---what a model understands, it should be able to generate---is a promise of unification and a necessity when composing both capabilities. Yet, existing benchmarks evaluate them in isolation: FID/GenEval for T2I; MME/MMBench for I2T. We show this gap is consequential: models scoring competitively on these benchmarks can fail severely when understanding and generation are composed, losing entities, attributes, spatial relations, and counts, resulting in semantic drift. To quantify drift, we introduce the Semantic Drift Protocol (SDP), inspired by the Telephone Game: starting from a caption or image, we alternate I2T and T2I over multiple generations and measure semantic preservation. We propose Mean Cumulative Drift (MCD), an embedding-based measure of content retention across three r
arXiv:2509.00094v2 Announce Type: replace-cross Abstract: Assessing spoken language is challenging, and quantifying pronunciation metrics for machine learning models is even harder. However, for the Holy Quran, this task is enabled by the rigorous recitation rules (Tajweed) established through the efforts of Muslim scholars, making highly effective assessment possible. Despite this advantage, the scarcity of high-quality annotated data remains a significant barrier. In this work, we bridge these gaps by introducing: (1) A 98% automated pipeline to produce high-quality Quranic datasets -- encompassing collection of recitations from expert reciters, segmentation at pause points (waqf) using our fine-tuned wav2vec2-BERT model, transcription of segments, and transcript verification via our novel Tasmeea algorithm; (2) 848 hours of audio (286K annotated utterances); (3) qdat_bench, a benchmark covering phonemes, diacritization, and Tajweed rules (Ghunnah, Qalqalah, Madd) on real recitation
arXiv:2502.12119v5 Announce Type: replace-cross Abstract: Visual instruction tuning adapts pre-trained Multimodal Large Language Models (MLLMs) to follow human instructions for real-world applications. However, the rapid growth of these datasets introduces significant redundancy, leading to increased computational costs. Existing methods for selecting instruction data aim to prune this redundancy, but predominantly rely on computationally demanding techniques such as proxy-based inference or training-based metrics. Consequently, the substantial computational costs incurred by these selection processes often exacerbate the very efficiency bottlenecks they are intended to resolve, posing a significant challenge to the scalable and effective tuning of MLLMs. To address this challenge, we first identify a critical, yet previously overlooked, factor: the anisotropy inherent in visual feature distributions. We find that this anisotropy induces a \textit{Global Semantic Drift}, and overlookin
arXiv:2406.10221v3 Announce Type: replace-cross Abstract: Recent developments in vision-language models have significantly advanced video understanding. Existing datasets and tasks, however, have notable limitations. Most datasets are confined to short videos with limited events and narrow narratives. For example, datasets with instructional and egocentric videos often depict the activities of one person in a single scene. Although existing movie datasets offer richer content, they are often limited to short-term tasks, lack publicly available videos, and frequently encounter data leakage issues given the use of subtitles and other information about commercial movies during LLM pretraining. To address the above limitations, we propose Short-Films 20K (SF20K), the largest publicly available movie dataset. SF20K consists of 20,143 amateur films, amounting to 3,582 hours of video, with an average of 12 minutes per movie. We accompany this dataset with SF20K-Test, a manual, open-ended ques
arXiv:2606.22454v2 Announce Type: replace Abstract: As LLM-generated text is increasingly used, especially in fictional domains, we explore how much LLM-generated stories differ from human-written stories. In this work, we focus on characters. We borrow definitions from narratology to analyze eight intricate dimensions of character, such as stylization and wholeness. These dimensions consider more than just basic characteristics. They assess how characters are portrayed within their stories. After automatically inferring categories of characters within both LLM and human-written stories, we compare and contrast these two sets of stories. We consider the following overarching questions: (1) Do LLMs and human-written stories have similar characters? and (2) Do LLMs generate stories with a variety of characters? Our analysis includes research questions that focus on stories generated by popular LLMs and recently published human-written stories. We describe a number of interesting similari