EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18624 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Mon, 10 Aug 2026 22:02:33 +0000
MedCity News

Jazz Pharma Expands Again in Epilepsy With $820M Actio Bio Acquisition

Jazz Pharmaceuticals, maker of the blockbuster epilepsy drug Epidiolex, is acquiring Actio Biosciences and its lead drug in pivotal testing for a rare type of epilepsy with no FDA-approved treatments. It’s Jazz’s second epilepsy deal in the past year. The post Jazz Pharma Expands Again in Epilepsy With $820M Actio Bio Acquisition appeared first on MedCity News .

Source ↗
technology Mon, 10 Aug 2026 17:36:14 -0400
EdTech Mag (Higher)

How Universities Are Turning Their Connected Campuses Into Smart Cities

A university is a small city in itself: There’s lighting, waste management, traffic control. That means, when researchers look to optimize a campus, they can simultaneously build a “smart cities” proof of concept for a municipality. It’s a logical fit. “Universities are often at the cutting edge of technology. What they lack is a city use case,” says Chris Lucero, technology and design director for The Connective, a smart city consortium for the greater Phoenix region. Cities meanwhile “have resource constraints and budget issues but would like to pilot new technologies to solve municipal…

Source ↗
technology Mon, 10 Aug 2026 17:32:24 -0400
EdTech Mag (K-12)

AI as an Extra Set of Eyes: Using AI Observation and Analytics To Support Student Success

Artificial intelligence has transformed many areas of education, from lesson planning and personalized learning to administrative tasks. It can analyze data, identify patterns in student learning that might otherwise go unnoticed and provide insights that support instructional planning and earlier intervention. AI systems can analyze patterns in student participation, attendance, assessment data and even student engagement within digital learning environments. These systems may identify patterns suggesting that a student is at risk academically long before those concerns become apparent…

Source ↗
technology Mon, 10 Aug 2026 17:20:44 +0000
MedCity News

MedCity Pivot: How Aetna is Adopting Agentic AI

In this episode of the MedCity Pivot Podcast, CVS Health Aetna’s chief digital and technology officer talks about how the insurance business is using AI, especially agentic AI, in healthcare. Point to note: AI is not used in automatic denials of coverage. The post MedCity Pivot: How Aetna is Adopting Agentic AI appeared first on MedCity News .

Source ↗
technology Mon, 10 Aug 2026 14:38:28 +0000
HN: education

The Medium Is the Mind: What AI Is Doing to Education

Article URL: https://medium.com/@henry.fisher.n/the-medium-is-the-mind-what-ai-is-doing-to-education-614bc2fbb493 Comments URL: https://news.ycombinator.com/item?id=49244310 Points: 3 # Comments: 0

Source ↗
technology Mon, 10 Aug 2026 13:09:00 +0000
MedCity News

AI in Drug Development is Not a Data Problem — It’s An Evidence Problem

Drug development isn’t limited by the amount of data we collect, but by our ability to preserve the context, meaning and relationships that transform data into evidence. The post AI in Drug Development is Not a Data Problem — It’s An Evidence Problem appeared first on MedCity News .

Source ↗
technology Mon, 10 Aug 2026 11:52:17 +0000
MedCity News

How to Keep Clinicians Satisfied, the Ascension Way

Ascension is pushing AI beyond note-taking to tackle the billing, order entry and nursing handoff work still weighing on clinicians after a patient visit ends. The post How to Keep Clinicians Satisfied, the Ascension Way appeared first on MedCity News .

Source ↗
technology Mon, 10 Aug 2026 11:50:00 +0000
MedCity News

Compliance Alone Won’t Fix the Home Health Fraud Problem

[Sponsored] With home health fraud growing more sophisticated through AI, post-pay audits are no longer enough. Payers must shift to proactive, tech-driven prevention to catch improper claims before paying them. The post Compliance Alone Won’t Fix the Home Health Fraud Problem appeared first on MedCity News .

Source ↗
technology Mon, 10 Aug 2026 09:00:00 +0000
Tech & Learning

Teaching With CoChat

One of the best deep research AI tools we’ve come across, CoChat acts as a lightning-fast research assistant.

Source ↗
technology Mon, 10 Aug 2026 09:00:00 +0000
Tech & Learning

What is NoodleTools And How Can I Use It to Teach?

NoodleTools is the research and citation assistance platform built to help make student life easier.

Source ↗
technology Mon, 10 Aug 2026 09:00:00 +0000
eCampus News

Beyond a keynote speaker: A strategic roadmap for AI in higher education

The atmosphere inside the Global Connect/American Accounting Association conference at Caesar’s Palace in August 2026 was a study in contrasts. Outside, the Las Vegas sun was punishing, an environment so relentless that I joked that the swag bags should have swapped out sunblock for flame retardant. The post Beyond a keynote speaker: A strategic roadmap for AI in higher education appeared first on eCampus News .

Source ↗
technology Mon, 10 Aug 2026 03:10:54 +0000
MedCity News

Braveheart Bio’s $382M IPO for Heart Drug Leads a Big Week for Biotech IPOs

In addition to pricing the biggest biotech IPO of the week, Braveheart Bio’s underwriters exercised their option to purchase more shares, boosting the clinical-stage company’s total haul to nearly $440 million. Braveheart’s lead drug could be competitive with a blockbuster Bristol Myers Squibb medicine for a type of cardiomyopathy. The post Braveheart Bio’s $382M IPO for Heart Drug Leads a Big Week for Biotech IPOs appeared first on MedCity News .

Source ↗
technology Mon, 10 Aug 2026 02:46:13 +0000
MedCity News

MedCity FemFwd: Employer Strategies for Women’s Health

In this episode, we’re joined by Cheryl Larson, president and CEO of the Midwest Business Group on Health, for our first in-person recording at MedCity News’ recent Bullseye event. The post MedCity FemFwd: Employer Strategies for Women’s Health appeared first on MedCity News .

Source ↗
technology Mon, 10 Aug 2026 02:05:28 +0000
MedCity News

Two Differing POVs Emerge About the Hinge-Cylinder Deal

Hinge Health’s acquisition of Cylinder expands it into GI care, but experts differ on whether employers will favor convenience or specialized solutions. The post Two Differing POVs Emerge About the Hinge-Cylinder Deal appeared first on MedCity News .

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Ratchet: How Reliable Must an LLM Judge Be to Retire a Skill?

arXiv:2605.22148v3 Announce Type: replace-cross Abstract: A large language model (LLM) agent that writes and edits its own skill library must also decide which skills to keep, from one noisy scalar per skill. The answer is exact: a judge scoring failures as passes at rate $(1-\tau)/2$ or above retires nothing, at any sample size, for eviction margin $\tau$. Audits find that machinery is rarely built: LLM-written skills are worth $+0.0$ percentage points (pp) against a no-skill control, human-written ones $+16.2$pp. Unmaintained, a library enters \emph{library drift}, growing until injecting a skill scores worse than injecting nothing. \textbf{Ratchet} repairs this: it evicts each skill on its measured contribution, caps the library at width $C$, and constrains synthesis, lifting held-out $pass@1$ by $+0.328$ on a hard MBPP+ slice. The matching non-divergence bound is finite for exactly two reasons, $C$ and $\tau$. Our contribution is the condition this repair carries and no deployed sy

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought

arXiv:2605.18079v2 Announce Type: replace-cross Abstract: Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconnect them from the models used in practice. We bridge this gap by analyzing standard transformer decoders with softmax attention and rounding of activations and attention weights, while allowing depth and width to grow logarithmically with the context length. As an intermediate step, we construct hardmax transformers with ternary activations and well-separated attention scores that simulate Turing machines using Chain-of-Thought (CoT). This lets us convert the constructions to equivalent softmax transformers without the unrealistic parameter magnitudes or activation precision that prior approaches would require. Using the same technique, we analyze a recently proposed summarized CoT paradigm and show that it simulates Turing machines more efficiently, with model size scaling logarit

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages

arXiv:2605.01720v3 Announce Type: replace-cross Abstract: Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings. While such resources are important for semantic understanding, they do not directly provide a unified interface for open-world recognition and translation, or for modern pose-driven sign language video generation frameworks: 1. RGB-based pretrained recognition models depend heavily on fixed backgrounds or clothing conditions during recording, and are less robust in open-world settings than style-agnostic pose-processing models. 2. Recent pose-guided image/video generation models mostly use a unified keypoint representation such as DWPose as their control interface. At present, the sign language field still lacks a data resource that can directly interface with this modern pose-native paradigm while also targeting real-world open scenarios. We present SignVerse-2M,

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

From Plan to Action: How Well Do Agents Follow the Plan?

arXiv:2604.12147v3 Announce Type: replace-cross Abstract: Agents are commonly instructed to follow a task-specific plan for guidance. However, it is unknown to what extent agents actually follow instructed plans. Without such an analysis, determining the extent agents comply with a given plan, it is impossible to assess whether a solution was reached through correct strategic reasoning or through other means, e.g., data contamination or overfitting to a benchmark. This paper presents the first extensive, systematic analysis of plan compliance in programming agents, examining 21,120 trajectories from SWE-agent across four LLMs on SWE-bench Verified and SWE-bench Pro under eight plan variations. Without an explicit plan, agents fall back on internalized workflows during training, which are often incomplete, overfit, or inconsistently applied. Providing the standard plan improves issue resolution, and we observe that periodic plan reminders can mitigate plan violations and improve task su

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Counterfactual Simulation Training for Chain-of-Thought Faithfulness

arXiv:2602.20710v2 Announce Type: replace-cross Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known problems with CoT faithfulness severely limit what insights can be gained from this practice. In this paper, we introduce a training method called Counterfactual Simulation Training (CST), which aims to improve CoT faithfulness by rewarding CoTs that enable a simulator to accurately predict a model's outputs over counterfactual inputs. We apply CST in two settings: (1) CoT monitoring with cue-based counterfactuals, to detect when models rely on spurious features, reward hack, or are sycophantic, and (2) counterfactual simulation over generic model-based counterfactuals, to encourage models to produce more faithful, generalizable reasoning in the CoT. Experiments with models up to 235B parameters show that CST can substantially improve monitor accuracy on cue-based counterfactuals (by 35 accuracy po

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

MetaSICL: Globalizing Auditory LLMs for Underserved Speakers and Languages via Meta Speech In-Context Learning

arXiv:2601.18904v3 Announce Type: replace-cross Abstract: Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditory Large Language Models (LLMs) are still largely trained and evaluated on high-resource data. Globalizing such systems requires handling low-resource settings, where the target speakers, languages, or tasks are poorly represented in training data. In these regimes, collecting enough labeled in-domain data is often impractical, and the small corpora available may still under-represent the test distribution, making direct fine-tuning brittle under domain shift. In-Context Learning (ICL) offers an alternative: instead of updating model parameters for every underserved community, an auditory LLM can adapt at inference time by conditioning on a few local demonstrations. However, vanilla speech ICL remains limited because most auditory LLMs are not explicitly trained to use such demonstrations effe

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

HFS: Holistic Query-Aware Frame Selection for Efficient Video Understanding

arXiv:2512.11534v2 Announce Type: replace-cross Abstract: Key frame selection is essentially a set-level optimization problem: the quality of the selected subset depends on the interactions among frames, rather than the score of any single frame. Existing methods generally exhibit three major limitations. Point-wise methods score each frame independently and ignore inter-frame dependencies. Although the training-free set-level methods explicitly model the inter-frame relationships, their selection criteria are fixed and cannot be adapted through downstream task feedback. Learnable methods can leverage data-driven training; however, they lack an explicit, differentiable set-quality objective and rely on offline-generated supervision signals. To address these limitations, we propose an end-to-end trainable and task-adaptive framework for frame selection. A Chain-of-Thought prompt conditions a Small Language Model (SLM) to extract task-specific latent query vectors, which are combined wit

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

arXiv:2511.07885v5 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains this paradigm faster than providers can scale. Two advances create an opportunity to rethink it: small, local LMs (<=20B active parameters) now achieve competitive performance to frontier models on many tasks, and local accelerators (e.g., Apple M4 Max) can host these models at interactive latencies. This raises the question: can local inference viably redistribute demand from centralized infrastructure? This requires measuring both whether local LMs can accurately answer real-world queries and whether they can do so efficiently on power-constrained devices (e.g., laptops). We propose intelligence per watt (IPW), task accuracy per unit of power, as a unified metric for the capability and efficiency of local inference across model-accelerator configurations. We evaluate 20+ state-of-the-art

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Diffusion-MF: Approximate Structured Diffusion for Sequence Labelling

arXiv:2606.18856v3 Announce Type: replace Abstract: We introduce Diffusion-MF, a discrete diffu- sion sequence labeller that places a linear-chain conditional random field (LCRF) inside the denoising loop. Unlike prior diffusion labellers, it performs structured inference at every step; parallel Mean-Field makes this efficient. Across multilingual POS, CoNLL-2003 NER, and joint Chinese Segmentation and POS, Diffusion-MF achieves the best primary result across all experimental settings.

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

A Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation

arXiv:2606.15059v2 Announce Type: replace Abstract: Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on short or pre-segmented speech rather than long-form, continuous input. Prior approaches are difficult to reproduce and make assumptions that do not hold for end-to-end systems. We present a practical evaluation method for long-form SimulS2ST. Given source speech, pre-segmented source transcripts, and reference translations, we run automatic speech recognition (ASR) and forced alignment on the generated target speech to recover token-level timestamps, then apply a sentence-embedding-based aligner to match the target text to its corresponding source sentences. This enables sentence-level computation of latency and quality metrics, including YAAL and xCOMET, which are then aggregated into final system-level scores. Experiments on representative SimulS2ST systems show that the method is effect

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

PolyFact: Comparing Consistency-Driven Post-training Methods for Cross-Lingual Factual Recall

arXiv:2606.06586v2 Announce Type: replace Abstract: Large language models (LLMs) trained predominantly on English data encode substantial world knowledge, yet often fail to express it reliably in other languages, a phenomenon known as cross-lingual factual inconsistency. To study this, we introduce PolyFact, a fully parallel multilingual factual QA dataset of 60K Wikidata-grounded facts across 12 typologically diverse languages, and propose consistency-driven GRPO with cross-lingual reward pooling. We compare our method against supervised fine-tuning (SFT) and the consistency-enhancement baselines DCO and CM-Align on OLMo-2-1124-7B and Qwen-2.5-7B, and analyze whether light continual pretraining (CPT) on parallel data provides a useful foundation for post-training. No single method dominates: SFT maximises in-distribution accuracy but not consistency, DCO yields the strongest consistency gains but fails to transfer to free-form generation, and our GRPO variant achieves the strongest tr

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

arXiv:2605.29738v2 Announce Type: replace Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible. We introduce Multi-Legal-Bench, the first cross-jurisdictional legal benchmark that evaluates identical tasks across six countries (Ukraine, France, Netherlands, Poland, Czech Republic, Lithuania), four language families, and 165 million full-text court decisions. The benchmark defines five tasks (court-type classification, judgment form classification, case-outcome prediction, legal norm extraction, and cause category prediction) mapped to structured metadata from national court registries, forming a deliberately sparse 5x6 task-jurisdiction matrix (20 of 30 cells filled). We evaluate 7 frontier LLMs under zero-shot and 3-shot prompting via AWS Bedrock, with 4 additional small/medium models (3-12B) for scaling analysis. Our results reveal that: (1) few-shot gains

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

VISHC at PsyDefDetect: Mitigating Data Scarcity in Psychological Defense Classification with Context-Aware Synthetic Augmentation

arXiv:2605.14380v2 Announce Type: replace Abstract: Psychological defense mechanisms (PDMs) are unconscious cognitive processes that modulate how individuals perceive and respond to emotional distress. Automatically classifying PDMs from text is clinically valuable but severely hindered by data scarcity and class imbalance, challenges which generative augmentation alone cannot resolve without psychological grounding. In this work, we address these challenges in the PsyDefDetect shared task (BioNLP@ACL 2026) by proposing a context-aware synthetic augmentation framework combined with a hybrid classification model. Our hybrid model integrates contextual language representations with basic clinical features, along with 150 annotated defense items. Experiments demonstrate that definition quality in prompting directly governs generation fidelity and downstream performance. Our method surpasses DMRS Co-Pilot, reaching an accuracy of 58.26% (+40.25%) and a macro-F1 of 24.62% (+15.99%), thereby

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages

arXiv:2605.02608v2 Announce Type: replace Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers---the Biaffine LSTM, Stack-Pointer Network, AfroXLMR-large, and RemBERT---across twelve typologically diverse languages, with a focus on low-resource African languages. We find that the Biaffine LSTM consistently outperforms transformer models in low-resource regimes, with transformers recovering their advantage as training data increases. The crossover falls within a resource range typical of treebanks for under-resourced languages. Morphological complexity (measured via MATTR) emerges as a significant secondary predictor of transformers' relative disadvantage after controlling for corpus size. These results indicate that the Biaffine LSTM may be better suited for syntactic tool development in low-resource regimes u

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens

arXiv:2604.26355v4 Announce Type: replace Abstract: Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains underexplored. We observe that reasoning tokens split into two functional types: low-entropy structural tokens (recurring phrases that scaffold the reasoning process) and higher-entropy organic tokens (problem-specific content that drives toward a solution). This asymmetry motivates a simple, model-agnostic compression pipeline: apply cross-word BPE merges on a model's own reasoning traces to derive \textit{supertokens} that capture frequent structural patterns, then teach the model to adopt them via supervised fine-tuning. Across three model families and five mathematical reasoning benchmarks, our approach shortens reasoning traces by 8.1% on average; under a TOST equivalence analysis at a +/- 2pp margin, accuracy is equivalent or inconclusive on 13/15 model -- benchmark cells (2 pass equ

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Skill-RAG: Failure-State-Aware Retrieval Augmentation via Hidden-State Probing and Skill Routing

arXiv:2604.15771v4 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has emerged as a foundational paradigm for grounding large language models in external knowledge. While adaptive retrieval mechanisms have improved retrieval efficiency, existing approaches treat post-retrieval failure as a signal to retry rather than to diagnose -- leaving the structural causes of query-evidence misalignment unaddressed. We observe that a significant portion of persistent retrieval failures stem not from the absence of relevant evidence but from an alignment gap between the query and the evidence space. We propose Skill-RAG, a failure-aware RAG framework that couples a lightweight hidden-state prober with a prompt-based skill router. The prober gates retrieval at two pipeline stages; upon detecting a failure state, the skill router diagnoses the underlying cause and selects among four retrieval skills -- query rewriting, question decomposition, evidence focusing, and an exit skill

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Joint Optimization of Reasoning and Dual-Memory for Self-Learning Diagnostic Agent

arXiv:2604.07269v2 Announce Type: replace Abstract: Clinical expertise improves not only by acquiring medical knowledge, but by accumulating experience that yields reusable diagnostic patterns. Recent LLMs-based diagnostic agents have shown promising progress in clinical reasoning for decision support. However, most approaches treat cases independently, limiting experience reuse and continual adaptation. We propose SEA, a self-learning diagnostic agent with cognitively inspired dual-memory module. We design a reinforcement training framework tailored to our designed agent for joint optimization of reasoning and memory management. We evaluate SEA in two complementary settings. On standard evaluation with MedCaseReasoning dataset, SEA achieves 92.46% accuracy, outperforming the strongest baseline by +19.6%, demonstrating the benefit of jointly optimizing reasoning and memory. On the long-horizon with ER-Reason dataset, SEA attains the best final accuracy (0.7214) and the largest improvem

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality

arXiv:2604.06756v2 Announce Type: replace Abstract: Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible to surface-level biases. One possible reason is that these judges lack sufficient information in assessing answer correctness. With the rise of reasoning-capable models, exposing a generator's reasoning content to the judge provides richer information and is a natural candidate for improving judgment accuracy. However, its actual impact on judge behavior remains understudied. In this paper, we systematically investigate how access to reasoning chains affects LLM-based judgment across factual question answering (QA) and mathematical reasoning benchmarks. We find that weak judges are easily swayed by reasoning presence, frequently accepting incorrect answers accompanied by fluent reasoning, while strong judges can partially leverage reasoning as informative evidence. Nevertheless, even stron

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Improving Attributed Long-form Question Answering with Intent Awareness

arXiv:2603.27435v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports. However, while these models are trained on diverse academic papers and reports, they are not exposed to the reasoning processes and intents that guide authors in crafting these documents. We hypothesize that enhancing a model's intent awareness can significantly improve the quality of generated long-form reports. We develop and employ structured, tag-based schemes to better elicit underlying implicit intents to write or cite. We demonstrate that these extracted intents enhance both zero-shot generation capabilities in LLMs and enable the creation of high-quality synthetic data for fine-tuning smaller models. Our experiments reveal improved performance across various challenging scientific report generation tasks, with an average improvement of +2.9 and +12.3 absolute points for large and small models over baselines, respect

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

AfriNLLB: Efficient Translation Models for African Languages

arXiv:2602.09373v2 Announce Type: replace Abstract: In this work, we present AfriNLLB, a series of lightweight models for efficient translation from and into African languages. AfriNLLB supports 15 language pairs (30 translation directions), including Swahili, Hausa, Yoruba, Amharic, Somali, Zulu, Lingala, Afrikaans, Wolof, and Egyptian Arabic, as well as other African Union official languages such as Arabic (MSA), French, Portuguese, and Spanish. Our training data covers bidirectional translation between English and 13 languages, and between French and two languages (Lingala and Wolof). AfriNLLB models are based on NLLB-200 600M, which we compress using iterative layer pruning and quantization. We fine-tune the pruned models on parallel corpora we curated for African languages, employing knowledge distillation from a larger teacher model. Our work aims at enabling efficient deployment of translation models for African languages in resource-constrained settings. Our evaluation results

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Kimi K2.5: Visual Agentic Intelligence

arXiv:2602.02276v2 Announce Type: replace Abstract: We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5 introduces Agent Swarm, a self-directed parallel agent orchestration framework that dynamically decomposes complex tasks into heterogeneous sub-problems and executes them concurrently. Extensive evaluations show that Kimi K2.5 achieves state-of-the-art results across various domains including coding, vision, reasoning, and agentic tasks. Agent Swarm also reduces latency by up to $4.5\times$ over single-agent baselines. We release the post-trained Kimi K2.5 model checkpoint to facilitate future research and real-world applications of agentic i

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction

arXiv:2505.16170v4 Announce Type: replace Abstract: We study the internal mechanisms that govern when LLMs choose to retract wrong answers, i.e., spontaneously and immediately acknowledge errors in their previously generated false assertions. Using model-specific testbeds, we find that while LLMs are capable of retraction, they do so only rarely, even when they can recognize their mistakes when asked in a separate interaction. We identify a reliable predictor of retraction: the model's momentary belief, as measured by a linear probe on its internal representation. The probe is trained to predict the correctness of answers on external datasets unrelated to retraction, then applied to settings where models should retract. A model retracts only when it "believes" its answers to be incorrect during generation; these beliefs frequently diverge from models' parametric knowledge as measured by factoid questions. Steering experiments further demonstrate that model belief causally drives retrac

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Prediction

arXiv:2411.00028v3 Announce Type: replace Abstract: Socioeconomic prediction aims to leverage various urban data to predict the socioeconomic indicators of regions such as population and commercial activity level, which plays an important role in understanding urban regions and supporting decision-making. Existing studies leverage knowledge graphs (KG) to model heterogeneous urban data, and further apply graph representation learning methods for socioeconomic prediction. However, these approaches heavily rely on heuristic ideas and expertise to extract task-relevant knowledge from diverse data, which may not be optimal for specific tasks. Additionally, they tend to overlook the inherent relationships between different indicators, limiting the prediction accuracy. Motivated by the remarkable abilities of large language models (LLMs), in this work, we propose a synergistic framework of LLM agents and KG, which integrates the reasoning and representation learning on KG with LLM agents. We

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent

arXiv:2608.07449v1 Announce Type: cross Abstract: LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that are loaded into the agent's context without weight updates. Recent methods refine skills through iterative task execution, failure diagnosis, and trajectory-guided text-space updates. However, existing frameworks lack explicit diagnosis--outcome feedback and treat deletion as a generic edit operation rather than a dedicated mechanism for consolidating accumulated knowledge. We introduce SkillProx, a proximal-gradient-inspired forward--backward framework that couples closed-loop diagnostic evolution with utility-aware proximal refinement. Motivated by a composite objective balancing task loss and skill complexity, the forward stage re-executes diagnosis-driven edits on the same task batch, rolls back regressions, and feeds measured outcomes into subsequent diagnoses. The backward s

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

arXiv:2608.07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and challenge current models. We present SABRE, a scalable, automated pipeline that converts a Test Primer (a Markdown Task Design with Data Schema) into structured specifications, generated or edited images, and question-answer pairs. Automated filtering removes candidates solved by a Filtering VLM, while human review verifies candidate validity and supports annotation correction and localized image repair. We instantiate SABRE-Prior to test whether VLMs follow visual evidence instead of relying on world priors -- learned expectations about familiar objects and scenes. Its 600 images and 1,000 questions span Context (unexpected entities in familiar scenes), Texture (counterfactual materials), Attribute (noncanonical

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

ResidencyRL: Reinforcement Learning in Simulated Clinical Environments

arXiv:2608.07418v1 Announce Type: cross Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy. Much of clinical reasoning relies on the patient encounter, a dialogue in which a clinician elicits history, refines diagnostic hypotheses, and decides management under uncertainty. While large language models (LLMs) excel on static medical benchmarks, methods to optimize the full sequence of clinical decisions remain underdeveloped. We present ResidencyRL, a reinforcement learning (RL) method for training clinical artificial intelligence (AI) agents through simulated multi-turn clinical encounters (up to 60 dialogue turns and 8 tool calls per trajectory). ResidencyRL pairs the policy agent with LLM simulators capable of complex, adversarial behaviors, training against a structured reward aligned to diagnostic accuracy

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

arXiv:2608.07411v1 Announce Type: cross Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning

arXiv:2608.07371v1 Announce Type: cross Abstract: Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. However, a completed rollout can yield many such signals, leaving their appropriate allocation across turns unclear. We introduce TRIAL, a trajectory-relative hindsight distillation framework with a unified turn-aligned scoring protocol. For each decision turn, TRIAL extracts an outcome view of that decision's realized consequence and evaluates the same response under ordinary and hindsight-conditioned contexts. The signed log-probability gap determines the direction and local strength of token-level supervision, while turn-level magnitudes are normalized jointly over the realized trajectory. The resulting allocation multipliers have an eligible-token-weighted mean of one, redistributing dense supervision across turns while fixing its average multiplier. Experiments on WebShop and ALFWorld with different backbones show that TRIAL outperform

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications

arXiv:2608.07250v1 Announce Type: cross Abstract: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will reduce disease burdens. However, relevant evidence is dispersed and infeasible for humans to comprehensively synthesize. LLMs may enable scalable, expert-level systematic evidence synthesis to identify microbe-cancer pairs; however, such capabilities have not yet been demonstrated. Domain experts were recruited to create a dataset to benchmark LLM performance (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, GPT-5 Nano) on 24 research papers using MMTV-LV and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal, consisting of MCQ, Likert-scale, multi-select, and free-text question types (77 items across 24 papers). Agreement between (1) experts and (2) experts and each LLM was determined per question instance using novel metrics. LLMs were assessed by c

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models

arXiv:2608.07243v1 Announce Type: cross Abstract: Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM creativity by adapting FunSearch to recipe generation for the 2024 Pillsbury Bake-Off and evaluating outputs against human benchmarks using TTCT-based LLM evaluation. Across two experiments, we test iteration count, generator temperature, and in-loop selection-scorer model size. Results show that iterative generation-selection can produce recipes with creativity scores comparable to human benchmarks, but additional iterations alone do not improve creativity. The in-loop evaluator matters most: a smaller selection scorer yields significantly higher scores across most TTCT dimensions, while temperature has limited effects except for originality. These findings suggest that evaluator design is a first-order design

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Modular TTT: Rethinking Test-Time Training as Composable Modules

arXiv:2608.07110v1 Announce Type: cross Abstract: Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. To address this, we propose Modular TTT, a framework that represents the inner learner as a directed acyclic graph and exposes the fast-weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions. Modular TTT automatically composes primitive-level train-view forward, train-view backward, and causal query-view rules into the full graph-level TTT computation, including the fast-weight state transition. Using Modular TTT, we systematically ablate the components of TTT and find that small learning-rate initialization, weight decay, and a single-layer non

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding

arXiv:2608.07067v1 Announce Type: cross Abstract: Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static retrieval and fragile cross-round memory. Mainstream single-round methods commit to a fixed top-$k$ page set at the outset and struggle to recover from early retrieval errors; recent iterative approaches allow multi-round evidence acquisition, but they do not investigate the propagation mechanism of cross-round states, making it difficult to track the dynamic changes in page relevance. To address these limitations, we propose DocMemo, a memory-guided framework that formulates long-document reasoning as dynamic evidence exploration. DocMemo maintains a tri-level retrieval state consisting of Document Schema Memory, Page Belief Memory, and Question Episodic Memory, which respectively capture structural priors, dynamic relevance estimation, and query-specific reasoning trajectories. During

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding

arXiv:2608.06869v1 Announce Type: cross Abstract: We describe DAEP, team BIGC's submission to NLPCC 2026 Shared Task 1 Track 3: Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The task requires retrieving the target video from 50 candidates and localizing the answer-supporting span. DAEP ranks videos with subtitle, visual, and procedural-context evidence, expands high-scoring anchors into temporal spans, and reranks spans for final output. Its main design is to convert the task-provided simple/complex input label into an inference-time evidence plan controlling modality weights, Top-K aggregation, boundary threshold, expansion length, and reranking strength. In the official evaluation, BIGC ranks first among ten systems with an Average score of 0.2728. Validation ablations show that visual evidence, procedural context, and difficulty-aware planning improve ranking quality, with the largest gain on complex questions.

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes

arXiv:2608.06795v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) enables efficient specialization and distribution of large language models through compact adapters. However, untrusted adapters introduce a supply-chain threat: a backdoored adapter can cause a model to generate harmful content, malicious code, political propaganda, or covert advertisements when an input contains a hidden trigger. Adapter-agnostic defenses merge the adapter with the base model, which dilutes backdoor signals and reduces detection performance. Existing adapter-aware methods do not address how to safely use a potentially backdoored adapter. Instead, they either train a defensive adapter to repair a backdoored base model, addressing the inverse problem rather than securing the adapter itself, or rely on a classifier that flags the entire adapter as suspicious and requires separate mitigation. These methods overlook the distinct latent-space signatures produced by trigger-bearing inputs in backdo

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Genotypic Triggers: Exposing Pharmacogenomic Blind Spots via Host-Specific Backdoors in Generative Antimicrobial Peptide Models

arXiv:2608.06779v1 Announce Type: cross Abstract: Large Language Models (LLMs) have accelerated drug discovery, particularly in the automated design of antimicrobial peptides (AMPs). However, current validation pipelines for peptide generation models overlook historical precedents showing that certain drugs carry health risks predominantly for individuals with specific genetic profiles. In this paper, we demonstrate that such targeted health risks can be induced intentionally and at scale by manipulating models that generate peptide candidates. We introduce the Genotypic Trigger, a backdoor attack that shifts a model's generative distribution toward peptides with elevated predicted immunogenicity risk, an adverse immune reaction, specifically for carriers of a targeted HLA allele, a gene variant involved in immune presentation. Across popular peptide generation models, the attack increased the predicted immunogenicity risk score for target-allele carriers by 743% on average relative to

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.CL

Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence

arXiv:2608.06778v1 Announce Type: cross Abstract: Mapping cyber threat intelligence (CTI) text to MITRE ATT&CK techniques is essential for structured threat analysis, yet manual annotation is costly and does not scale. The ATT&CK taxonomy comprises several hundred attack techniques, and a single CTI passage may describe multiple techniques, making accurate and complete extraction challenging. Existing automated approaches fall short in different ways: multi-label classifiers struggle with severe class imbalance and the large label space, while LLM-based methods--retrieval pipelines and fine-tuned generators--optimize token-level objectives that treat technique annotation as sequence generation rather than set prediction, lacking direct supervision on whether the predicted technique set is correct and complete. We propose TTP-R1, a two-stage framework that combines retrieval-augmented supervised fine-tuning (SFT) with reinforcement learning using verifiable rewards (RLVR). A hybrid retr

Source ↗
Showing 8451–8500 of 11029 signals
← Prev Page 170 of 221 Next →