EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

behavior Wed, 05 Nov 2025 10:00:00 +0000
eSchool News

Room to grow: Creating a classroom built for success

For decades, curriculum, pedagogy, and technology have evolved to meet the changing needs of students. But in many schools, the classroom environment itself hasn’t kept pace.

Source ↗
technology Wed, 05 Aug 2026 23:48:31 +0000
MedCity News

6 Ways UnityPoint Health Is Trying to Make Clinicians’ Jobs Easier

UnityPoint Health’s chief medical officer, Gregory Johnson, shared several ways the system is using AI to lift administrative burden off clinicians and get them back to patient care — including ambient scribing, real-time capacity tracking and AI-powered chart review. The post 6 Ways UnityPoint Health Is Trying to Make Clinicians’ Jobs Easier appeared first on MedCity News .

Source ↗
technology Wed, 05 Aug 2026 22:13:20 +0000
MedCity News

Curative, Wondr Health Team Up for Weight and Metabolic Health Support

Curative Insurance Company partnered with Wondr Health to expand weight management support for its members. The post Curative, Wondr Health Team Up for Weight and Metabolic Health Support appeared first on MedCity News .

Source ↗
audience Wed, 05 Aug 2026 20:43:00 -0400
Higher Ed Dive

Fitch: New visa limits could hurt international enrollment and revenue

Financial pressure will be greatest on institutions with weaker student demand and greater reliance on tuition and fee revenue, analysts wrote.

Source ↗
technology Wed, 05 Aug 2026 18:37:16 +0000
MedCity News

Attovia’s Upsized IPO Brings In $289M to Fund New Kind of Biologic for Immune Disorders

Attovia Therapeutics’ biologic medicines could offer advantages over currently available immune disorder drugs, such as Sanofi’s blockbuster medicine Dupixent. The IPO comes as Attovia looks to expand clinical development of its pipeline in 2027. The post Attovia’s Upsized IPO Brings In $289M to Fund New Kind of Biologic for Immune Disorders appeared first on MedCity News .

Source ↗
regulation Wed, 05 Aug 2026 18:30:00 +0000
The 74

Maryland Gov. Offers Late Admission, $800 Credit to Disenrolled Howard Students

Maryland Gov. Wes Moore (D) said Thursday the state would provide an $800 credit and expedited admission processing for any of the hundreds of incoming Howard University freshmen who were abruptly told last week that they no longer had a place at the school. The governor said in a statement that he directed the Maryland […]

Source ↗
technology Wed, 05 Aug 2026 18:20:16 -0400
EdTech Mag (Higher)

Why Higher Ed IT Can't Afford To Wait on Quantum Security

Sometimes, dealing with cybersecurity can feel like the boardwalk arcade game Whac-A-Mole. Just when you’ve addressed a potential threat with hardware, software and operating system patches, another threat comes from a direction you never anticipated. A new threat to many organizations’ cybersecurity is looming, and it could be the biggest one of all: quantum computing. The good news is that universities and other organizations can take steps now to protect their systems that might be vulnerable in a few years. Quantum computing, where processors use the principles of quantum physics, is…

Source ↗
regulation Wed, 05 Aug 2026 17:43:00 -0400
K-12 Dive

Kids Online Safety Act clears key Senate panel with bipartisan support

Three other bills aimed at protecting youth online also won approval Wednesday from the Commerce, Science and Transportation Committee.

Source ↗
regulation Wed, 05 Aug 2026 16:30:00 +0000
The 74

The Schools Turning to AI to Rate Students Beyond A-F Grades

Artificial Intelligence will soon help teachers rate students’ academic and “soft” skills at schools using nontraditional report cards that skip A-F grading after an AI company purchased the non-profit Mastery Transcript Consortium last month. The New York-based company Legend.org is developing AI tools to give teachers and students immediate feedback on papers, exams, recorded presentations […]

Source ↗
regulation Wed, 05 Aug 2026 14:30:00 +0000
The 74

2-K Offers Are Out: 2,000 NYC Toddlers Get Spots, 5,700 Apply

Roughly 2,000 families received offers for a spot in the inaugural year of 2-K, New York City’s new free childcare program for 2-year-olds, officials announced Tuesday. Around 5,700 families applied for the initial 2,000 seats, which were restricted to priority neighborhoods in Manhattan, Brooklyn, Queens, and the Bronx. “Today represents a monumental step on the […]

Source ↗
regulation Wed, 05 Aug 2026 14:00:00 -0400
K-12 Dive

Teaching students how to Google effectively in the age of AI

As Google AI overviews become the norm, educators should push students to dig deeper and do their own sourcing, a prominent school librarian says.

Source ↗
regulation Wed, 05 Aug 2026 14:00:00 -0400
K-12 Dive

Immigrant history delivers the past through the eyes of everyday people

Through a summer institute, New York City’s Tenement Museum is providing “turnkey” ideas that make lessons more relevant to students, participants say.

Source ↗
technology Wed, 05 Aug 2026 13:29:00 +0000
MedCity News

The Sleep Crisis We May Be Missing: Why Millions of Americans Remain Undiagnosed

Despite the silk masks, oils, and sleepy girl mocktails swirling around social media, one of the greatest causes of non-restorative sleep is rarely mentioned by wellness influencers online: actual sleep disorders. The post The Sleep Crisis We May Be Missing: Why Millions of Americans Remain Undiagnosed appeared first on MedCity News .

Source ↗
technology Wed, 05 Aug 2026 13:21:00 +0000
MedCity News

The Healthcare Leaders Driving Change Need to Get Better at Explaining It

Healthcare leaders who shape that future need more than innovation. They need a visible, credible voice that helps the people around them — patients, clinicians, employees, partners, and investors — understand what the change is, why it matters, and why this is the right person to lead it. The post The Healthcare Leaders Driving Change Need to Get Better at Explaining It appeared first on MedCity News .

Source ↗
technology Wed, 05 Aug 2026 12:56:16 -0400
EdTech Mag (K-12)

Reliable Networks Make Modern Classrooms Possible

The technology that turns the most heads in K–12 schools tends to be the flashy equipment — interactive displays, 3D printers, powerful gaming computers, lifelike robots. But IT leaders know that the technology that matters most is often invisible. It’s a fast network designed to handle a classroom full of devices, clear policies that guide smart decisions, and cybersecurity strategies that guard student and institutional data. Click the banner below to find the tools you need to protect your district from cyberthreats.

Source ↗
behavior Wed, 05 Aug 2026 12:42:46 +0000
District Admin

How to elevate college and career readiness in 3 steps

Readiness requires deliberate decisions about what students are expected to learn, how they are supported, and how progress is measured. The post How to elevate college and career readiness in 3 steps appeared first on District Administration .

Source ↗
regulation Wed, 05 Aug 2026 12:30:00 +0000
The 74

Opinion: As the Education Department Is Dismantled, Who Protects the Right to Learn?

A first grader reads an entire page on her own after months of specialized instruction. A middle school student with autism delivers his first classroom presentation. These moments are not medical breakthroughs, they are educational ones. Unfortunately, our history — and too often our present — shows that many students still do not experience classrooms […]

Source ↗
regulation Wed, 05 Aug 2026 12:22:14 -0400
K-12 Dive

How schools can tackle mental well-being for teens in military families

These adolescents face distinct hurdles, including frequent moves and parental deployments, not experienced by "civilian" teens.

Source ↗
behavior Wed, 05 Aug 2026 11:44:46 +0000
District Admin

9 million U.S. students lost school time for weather in the 2024-2025 school year

The nonprofit UndauntedK12 maps instances across the U.S. when a school day or school activity was canceled, delayed, or cut short by extreme weather. The post 9 million U.S. students lost school time for weather in the 2024-2025 school year appeared first on District Administration .

Source ↗
behavior Wed, 05 Aug 2026 11:38:24 +0000
District Admin

How the costly digital age is leaving behind rural, underfunded schools

Matthew Splain, superintendent of the Otto-Eldred School District, said his schools are struggling to keep aging, refurbished laptops from dying on students. The district’s struggles are not isolated. The post How the costly digital age is leaving behind rural, underfunded schools appeared first on District Administration .

Source ↗
regulation Wed, 05 Aug 2026 10:30:00 +0000
The 74

Measles Cases Hit 35-Year Record High Just As Kids Head Back to School

As millions of kids across the country prepare to head back to school — and many have already returned — the national count of measles cases continues to climb, reaching its highest number in 35 years. Utah and South Carolina, which both allow for various exemptions to school-based vaccine mandates, carry most of the disease burden, accounting […]

Source ↗
behavior Wed, 05 Aug 2026 10:00:00 +0000
eSchool News

In a school emergency, 911 is still flying blind

School safety emergencies are inherently chaotic, frightening, and made even more perilous by a lack of real-time visibility once an emergency is underway. In an active shooter incident, first responders often arrive at the scene with little or no knowledge about what is happening inside the school.

Source ↗
technology Wed, 05 Aug 2026 09:00:00 +0000
Tech & Learning

What Is Dabble And Why Is It Helpful to Teachers?

Dabble is a writing tool that helps users organize, save, and produce long writing drafts, and is perfect for thesis projects, novels, and more.

Source ↗
technology Wed, 05 Aug 2026 09:00:00 +0000
eCampus News

How we recovered student trust after big Wi-Fi challenges

Every day I come to work at Gettysburg College wearing a college-branded cap that I bought with a $25 gift card, presented to me by the Student Senate for leading an effort to improve student access to the campus Wi-Fi. The post How we recovered student trust after big Wi-Fi challenges appeared first on eCampus News .

Source ↗
audience Wed, 05 Aug 2026 07:00:00 +0000
Inside Higher Ed

From Knowledge Building to Wisdom Building in Higher Education

From Knowledge Building to Wisdom Building in Higher Education jdimaggio@upcea.edu Wed, 08/05/2026 - 03:00 AM In years gone by, we said that students went to college to gather knowledge . Now, college is appropriately expected to be much more about building wisdom . Byline(s) Ray Schroeder

Source ↗
audience Wed, 05 Aug 2026 07:00:00 +0000
Inside Higher Ed

Detentions of International Scholars at Airports Raise Alarm

Detentions of International Scholars at Airports Raise Alarm Sara Weissman Wed, 08/05/2026 - 03:00 AM A Johns Hopkins University researcher was detained at an airport by immigration officials and later released. Cases like hers make travel fraught for foreign scholars. Byline(s) Sara Weissman

Source ↗
audience Wed, 05 Aug 2026 07:00:00 +0000
Inside Higher Ed

Vanderbilt’s Board Declares War Against Academic Freedom

Vanderbilt’s Board Declares War Against Academic Freedom sara.custer@in… Wed, 08/05/2026 - 03:00 AM A new policy punishes faculty for speech the trustees deem strays from the core purpose of the university. Byline(s) John K. Wilson

Source ↗
audience Wed, 05 Aug 2026 07:00:00 +0000
Inside Higher Ed

AI Detectors Are Out, New Assessments Are In

AI Detectors Are Out, New Assessments Are In kathryn.palmer… Wed, 08/05/2026 - 03:00 AM A growing number of universities have prohibited AI detectors, arguing they are ‘unreliable’ indicators of cheating. But where does that leave faculty, who know all too well that cheating is rampant in the age of AI? Byline(s) Kathryn Palmer

Source ↗
audience Wed, 05 Aug 2026 07:00:00 +0000
Inside Higher Ed

Cornell Bans Processing Wild Animals in Communal Kitchens

Cornell Bans Processing Wild Animals in Communal Kitchens jessica.blake@… Wed, 08/05/2026 - 03:00 AM Byline(s) Jessica Blake

Source ↗
audience Wed, 05 Aug 2026 07:00:00 +0000
Inside Higher Ed

Where Enrollment Stands as Fall 2026 Approaches

Where Enrollment Stands as Fall 2026 Approaches Johanna Alonso Wed, 08/05/2026 - 03:00 AM Although six in 10 enrollment leaders said their colleges had met their fall 2026 enrollment goals, that number was lower for small colleges, which also discounted more and relied on campus visits to convert prospective students. Byline(s) Johanna Alonso

Source ↗
audience Wed, 05 Aug 2026 07:00:00 +0000
Inside Higher Ed

UMich College Shifts to P/NC Grading for First-Year Students

UMich College Shifts to P/NC Grading for First-Year Students Susan H. Greenberg Wed, 08/05/2026 - 03:00 AM Byline(s) Susan H. Greenberg

Source ↗
audience Wed, 05 Aug 2026 07:00:00 +0000
Inside Higher Ed

How HBCUs Help Students Find Their Way Back to College

How HBCUs Help Students Find Their Way Back to College Joshua.Bay Wed, 08/05/2026 - 03:00 AM InsideTrack’s five-year coaching initiative across 39 HBCUs helped stopped-out students return and stay on track toward degree completion. Byline(s) Joshua Bay

Source ↗
audience Wed, 05 Aug 2026 07:00:00 +0000
Inside Higher Ed

Texas A&M Faculty Challenge System’s Race and Gender Rules

Texas A&M Faculty Challenge System’s Race and Gender Rules Katherine Knott Wed, 08/05/2026 - 03:00 AM Byline(s) Katherine Knott

Source ↗
audience Wed, 05 Aug 2026 07:00:00 +0000
Inside Higher Ed

Nursing Coalition Challenges ED’s Professional Degrees List

Nursing Coalition Challenges ED’s Professional Degrees List jessica.blake@… Wed, 08/05/2026 - 03:00 AM Byline(s) Jessica Blake

Source ↗
audience Wed, 05 Aug 2026 05:00:00 -0400
Higher Ed Dive

Culture eats strategy for breakfast — but not always for college mergers

While cultural discordance needs careful consideration, it doesn't trump strong strategy, argues an expert on college consolidation.

Source ↗
regulation Wed, 05 Aug 2026 05:00:00 -0400
K-12 Dive

Trump administration targets teen pregnancy prevention grants

School and community partnerships to improve pregnancy and STI prevention are under threat after HHS suddenly cut 53 of 66 existing grants.

Source ↗
technology Wed, 05 Aug 2026 00:45:55 +0000
MedCity News

QuantHealth Snags $45M to Simulate Clinical Trials Before Patients Ever Enroll

QuantHealth raised $45 million in Series B funding to expand its AI platform, which simulates clinical trials before they run and predicts patient outcomes for pharma companies. The startup is aiming to cut the industry’s 90% trial failure rate and speed up the timelines for effective drugs to reach the market. The post QuantHealth Snags $45M to Simulate Clinical Trials Before Patients Ever Enroll appeared first on MedCity News .

Source ↗
behavior Wed, 05 Aug 2026 00:00:00 GMT
EdSurge

The Library That Sparked a STEM Revolution for Our Students

To truly close the digital equity gap, educators must become architects of student possibility.

Source ↗
behavior Wed, 05 Aug 2026 00:00:00 GMT
EdSurge

What Happens When AI Policy Meets a Real Classroom?

One national study, one local leader, and a reality check on how schools are really regulating AI.

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

arXiv:2606.08044v2 Announce Type: replace-cross Abstract: Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and answers benign ones. But refusing on the prompts an auditor happens to try does not show that the model is far from harmful behavior. Behavioral tests observe outputs; they do not measure how easily an intervention on the model turns a refusal into compliance. We call the gap between what static audits certify and what an intervention can reach the audit gap, and we show it is realizable: one can build a model that matches its safety-aligned base on every static audit yet gives way to a small, known perturbation of its internal state. We construct such dissociated models from three safety-aligned bases (Gemma 2 2B, Llama 3.2 3B, Qwen 2.5 3B) and audit the base, dissociated, and openly harmful models with the same soft interventions in parameter and latent space; the latent attacks are summarized

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

CaliDist: Calibrating Large Language Models via Behavioral Robustness to Distraction

arXiv:2606.05799v2 Announce Type: replace-cross Abstract: Existing calibration methods for Large Language Models (LLMs) often overlook a critical dimension of trustworthiness: a model's behavioral robustness to irrelevant or misleading information. In this paper, we argue that a model's true confidence should reflect its stability under cognitive pressure. We introduce CaliDist, a novel post-hoc calibration approach that directly measures and penalizes a model's susceptibility to distraction. CaliDist quantifies how an LLM's predictions and uncertainty change when its input prompt is perturbed with semantic distractors. This stability (or lack thereof) signal is then used to adaptively scale the model's initial confidence score. Our extensive experiments on seven Natural Language Understanding classification benchmarks using six distinct LLMs show that CaliDist consistently achieves lower Expected Calibration Error (ECE) and Brier Score compared with strong baselines. Remarkably, our m

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search

arXiv:2605.20244v2 Announce Type: replace-cross Abstract: We present Lean Refactor, a plug-and-play retrieval-augmented agentic framework for multi-objective, controllable, and version-robust refactoring of Lean proofs. LLM-generated proofs are notoriously correct-but-verbose and brittle across library versions, yet existing refactoring works overlook three practical challenges: 1) Lean refactoring is natively multi-objective (proof length, compilation cost, and version compatibility are often in tension); 2) Lean repositories have fragile compatibility, whereas LLM releases are unaware of Lean/Mathlib versions; 3) Training-based pipelines require repeated fine-tuning with each new LLM release, scaling neither with model churn nor with Lean's release cycle. Lean Refactor steers a frozen agentic LLM with retrievals from a curated database of multi-objective refactoring strategies, each densely annotated with metadata such as supported Lean/Mathlib versions and expected compilation-cost

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

arXiv:2604.23586v2 Announce Type: replace-cross Abstract: Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating high-level semantics and low-level details in a fully entangled manner. This is suboptimal for talking head synthesis: while audio and facial motion are semantically correlated, their low-level realizations (acoustic signals and visual textures) follow distinct rendering processes. Enforcing joint modeling across all levels causes unnecessary entanglement and reduces efficiency. We propose Talker-T2AV, an autoregressive diffusion framework where high-level cross-modal modeling occurs in a shared backbone, while low-level refinement uses modality-specific decoders. A shared autoregressive language model jointly reasons over audio and video in a unified patch-level token space. Two lightweight diffusio

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations

arXiv:2604.22207v2 Announce Type: replace-cross Abstract: Due to the textual and repetitive nature of many Requirements Engineering (RE) artefacts, Large Language Models (LLMs) have proven useful to automate their generation and processing. In this paper, we discuss a possible approach for automating the Goal-Oriented Requirements Engineering (GORE) process by extracting functional goals from software documentation through three phases: actor identification, high and low-level goal extraction. To implement these functionalities, we propose a chain of LLMs fed with engineered prompts. We experimented with different variants of in-context learning and measured the similarities between input data and in-context examples to better investigate their impact. Another key element is the generation-critic mechanism, implemented as a feedback loop involving two LLMs. Although the pipeline achieved 61% accuracy in low-level goal identification, the final stage, these results indicate the approach

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

Geometry-Aware Localized Watermarking for Copyright Protection in Embedding-as-a-Service

arXiv:2604.11344v2 Announce Type: replace-cross Abstract: Embedding-as-a-Service (EaaS) has become an important semantic infrastructure for natural language and multimedia applications, but it is highly vulnerable to model stealing and copyright infringement. Existing EaaS watermarking methods face a fundamental robustness--utility--verifiability tension: trigger-based methods are fragile to paraphrasing, transformation-based methods are sensitive to dimensional perturbation, and region-based methods may incur false positives due to coincidental geometric affinity. To address this problem, we propose GeoMark, a geometry-aware localized watermarking framework for EaaS copyright protection. GeoMark uses a natural in-manifold embedding as a shared watermark target, constructs geometry-separated anchors with explicit target--anchor margins, and activates watermark injection only within adaptive local neighborhoods. This design decouples where watermarking is triggered from what ownership i

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

arXiv:2604.02486v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they often fail on tasks that require fine-grained visual perception, even when the required information is still present in their internal representations. Prior work has attributed this ``hidden-in-plain-sight'' gap to the language model, but the cause remains unexplained. In this work, we demonstrate that this gap arises from the language model's lack of semantic labels for fine-grained visual details: when visual entities can be mapped to known concepts, VLMs bypass visual comparison and reason through language; when they cannot, VLMs resort to brittle and hallucinated descriptions. We verify this across semantic correspondence, synthetic shape matching, and face matching, and find that VLMs perform much better when the relevant entities are nameable than when they are unnamable. Mechanistically, Logit Lens an

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics

arXiv:2603.24929v2 Announce Type: replace-cross Abstract: Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment. However, traditional evaluation approaches provide limited insight into model confidence at individual token positions during generation. To address this issue, we introduce LogitScope, a lightweight framework for analyzing LLM uncertainty through token-level information metrics computed from probability distributions. By measuring metrics such as entropy and varentropy at each generation step, LogitScope reveals patterns in model confidence, identifies potential hallucinations, and exposes decision points where models exhibit high uncertainty, all without requiring labeled data or semantic interpretation. We demonstrate LogitScope's utility across diverse applications including uncertainty quantification, model behavior analysis, and production monitoring. The framework is model-agnostic, computationally efficien

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

Assessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language Models

arXiv:2603.19087v2 Announce Type: replace-cross Abstract: Creativity is the ability to come up with novel ideas, a capacity crucial for human development and flourishing. Are large language models (LLMs) creative in the same way humans are, and can the same interventions increase creativity in both? We study a promising but largely untested intervention for creativity: forcing creators to draw an analogy from a random, remote source domain (''cross-domain mapping''). Human participants and LLMs generated novel designs for ten daily products (e.g., backpack, TV) under two prompts: (i) cross-domain mapping, which required drawing inspiration from a randomly assigned source (e.g., octopus, cactus, GPS), and (ii) user need, which required proposing innovations targeting unmet user needs. We show that humans reliably benefit from randomly assigned cross-domain mappings, while LLMs, on average, generate more original ideas than humans and do not show a statistically significant effect of cro

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

MIDI-LLM: Improving Text-to-MIDI Music Generation via Adapting Large Language Models

arXiv:2511.03942v2 Announce Type: replace-cross Abstract: We present MIDI-LLM, a recipe that improves multitrack text-to-MIDI generation via adapting Large Language Models (LLMs). MIDI-LLM expands an LLM's text vocabulary to include MIDI tokens and employs a two-stage training pipeline: (i) unimodal continued pretraining on music-adjacent text and standalone MIDIs, and (ii) multimodal supervised finetuning on text-MIDI pairs. Our instantiation of MIDI-LLM based on Llama 3.2 (1B) outperforms the recent Text2midi model in both text control and musical quality, and readily integrates with optimized inference ecosystems like vLLM. To align with real-world songwriting workflows, we further finetune our MIDI-LLM on the TheoryTab dataset for text-conditioned lead sheet (i.e., melody + chords) generation and infilling. A comprehensive ablation study validates the synergy between LLM text pretraining, standalone MIDI pretraining, and supervised text-to-MIDI finetuning. Finally, an in-the-wild b

Source ↗
technology Wed, 05 Aug 2026 00:00:00 -0400
arXiv cs.CL

Don't Walk the Line: Boundary Guidance for Filtered Generation

arXiv:2510.11834v3 Announce Type: replace-cross Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tune the generator to reduce the probability of being filtered, but this can be suboptimal: it often pushes the model toward producing samples near the classifier's decision boundary, increasing both false positives and false negatives. We propose Boundary Guidance, a reinforcement learning fine-tuning method that explicitly steers generation away from the classifier's margin. On a benchmark of jailbreak, ambiguous, and longcontext prompts, Boundary Guidance improves both the safety and the utility of outputs, as judged by LLM-as-a-Judge evaluations. Comprehensive ablations across model scales and reward designs demonstrate the robustness of our approach.

Source ↗
Showing 1601–1650 of 18349 signals
← Prev Page 33 of 367 Next →