EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals
Admin mode. Curation controls visible. Keep this URL (with token) private.

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

audience Mon, 03 Aug 2026 07:00:00 +0000
Inside Higher Ed

What Parents Prioritize in the College Search

What Parents Prioritize in the College Search Johanna Alonso Mon, 08/03/2026 - 03:00 AM Byline(s) Johanna Alonso

Source ↗
audience Mon, 03 Aug 2026 07:00:00 +0000
Inside Higher Ed

DOD Considers Course Reviews, Axing Tenure at Military Academies

DOD Considers Course Reviews, Axing Tenure at Military Academies kathryn.palmer… Mon, 08/03/2026 - 03:00 AM Byline(s) Kathryn Palmer

Source ↗
audience Mon, 03 Aug 2026 07:00:00 +0000
Inside Higher Ed

Creating a Clearer Transfer Path

Creating a Clearer Transfer Path Joshua.Bay Mon, 08/03/2026 - 03:00 AM California State University’s Transfer Success Pathway provides eligible community college students guaranteed admission and support as they pursue a bachelor’s degree. Byline(s) Joshua Bay

Source ↗
audience Mon, 03 Aug 2026 07:00:00 +0000
Inside Higher Ed

About 40 AAUP Michigan Members Attended Meeting on Endorsing Candidate

About 40 AAUP Michigan Members Attended Meeting on Endorsing Candidate Ryan Quinn Mon, 08/03/2026 - 03:00 AM The AAUP's first-known national-level endorsement of a political candidate has drawn criticism. Some Michigan members have questioned the representativeness of a poll on their feelings. Byline(s) Ryan Quinn

Source ↗
audience Mon, 03 Aug 2026 07:00:00 +0000
Inside Higher Ed

Acupuncture and Integrative Medicine College Closes

Acupuncture and Integrative Medicine College Closes Katherine Knott Mon, 08/03/2026 - 03:00 AM Byline(s) Katherine Knott

Source ↗
audience Mon, 03 Aug 2026 07:00:00 +0000
Inside Higher Ed

Inheriting the Disenchanted

Inheriting the Disenchanted sara.custer@in… Mon, 08/03/2026 - 03:00 AM Persistence matters. Byline(s) Matt Reed

Source ↗
audience Mon, 03 Aug 2026 07:00:00 +0000
Inside Higher Ed

Are Loan Limits Driving Down College Costs?

Are Loan Limits Driving Down College Costs? jessica.blake@… Mon, 08/03/2026 - 03:00 AM The Education Department and some policy experts say yes. But others aren’t so sure. Byline(s) Jessica Blake

Source ↗
audience Mon, 03 Aug 2026 07:00:00 +0000
Inside Higher Ed

Big 10, SEC Back Senate’s College Sports Overhaul

Big 10, SEC Back Senate’s College Sports Overhaul Katherine Knott Mon, 08/03/2026 - 03:00 AM Byline(s) Katherine Knott

Source ↗
audience Mon, 03 Aug 2026 05:00:00 -0400
Higher Ed Dive

Week in review: A $100K fee for OPT?

We’re rounding up recent stories, from a look at the keywords used to target grants to a group of higher education bills advancing in the Senate.

Source ↗
regulation Mon, 03 Aug 2026 05:00:00 -0400
K-12 Dive

Falling back in love with books: What middle schoolers need first

Middle schoolers aren't reading for fun anymore. Here's how to bring the joy back.

Source ↗
regulation Mon, 03 Aug 2026 05:00:00 -0400
K-12 Dive

The school year starts now, so should planning for next summer

Why the most successful school districts are already planning for next summer.

Source ↗
regulation Mon, 03 Aug 2026 05:00:00 -0400
K-12 Dive

Week In Review: Where enrollment is and isn’t in decline

We’re rounding up last week’s news, from the impact of school closures on teacher turnover to districts’ social media lawsuits.

Source ↗
need Mon, 03 Aug 2026 05:00:00 +0000
Hechinger Report

Pace of college closings picks up, with more projected

Pssst — wanna buy a college campus? There’s a sudden market in abandoned campuses, an outgrowth of the accelerating rate of colleges closing in the face of falling enrollment, rising debt and other problems. More than 440 of the nation’s private, nonprofit four-year colleges and universities are now at risk, based on enrollment trends, debt […] The post Pace of college closings picks up, with more projected appeared first on The Hechinger Report .

Source ↗
need Mon, 03 Aug 2026 05:00:00 +0000
Hechinger Report

OPINION: America’s teacher shortage is worse than you think. But there’s a way to fix it

When I was an assistant principal in Chicago, I would post a job opening for a physics teacher and get 20 applications. A call for an English teacher might yield 150. That was the late 1990s. By the time I left the district, the same postings were lucky to bring in one or two qualified […] The post OPINION: America’s teacher shortage is worse than you think. But there’s a way to fix it appeared first on The Hechinger Report .

Source ↗
technology Mon, 03 Aug 2026 01:56:44 +0000
MedCity News

‘We’re Not Going to Be Bullied’: Sentara Digs in on Anthem Contract Fight

Sentara Health notified Anthem that it will let its commercial, Medicare and Medicaid contracts expire at year’s end unless the two sides reach a new rate agreement. The dispute puts coverage for an estimated 380,000 Virginians at risk. The post ‘We’re Not Going to Be Bullied’: Sentara Digs in on Anthem Contract Fight appeared first on MedCity News .

Source ↗
behavior Mon, 03 Aug 2026 00:00:00 GMT
EdSurge

I Used AI to Build AI-Resistant Assignments

A teacher created an app to fight cheating — and discovered a bigger takeaway.

Source ↗
behavior Mon, 03 Aug 2026 00:00:00 GMT
EdSurge

High Schools Need a New Model for a New Economy

Career-connected learning, not just college readiness, will prepare students for jobs that keep shifting.

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models

arXiv:2606.05976v2 Announce Type: replace-cross Abstract: Recent works show that LLM agents struggle to correct errors in their own reasoning traces, despite their ability to correct errors from external sources. We ask whether this reflects a capability deficit or an artifact of the role labeling. To test this, we design a training-free intervention, source-conditioned role relabeling, that keeps the erroneous claim byte-identical and varies only its message role. The claim is presented inside the agent's "", a user message, a tool response, or a system "" block. We test 12 model-domain combinations spanning closed-weight APIs and open-weight models from 70B-class down to smaller families. Relabeling "" to an external role increases the explicit-correction rate by 23 to 93 percentage points, significant in 10 of 12 experimental settings. This suggests that these models' failure to detect a self-generated error is largely an artifact of how the claim is role-labeled in the chat templat

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Few-Shot Contrastive Adaptation for Audio Abuse Detection in Low-Resource Indic Languages

arXiv:2604.09094v2 Announce Type: replace-cross Abstract: Abusive and hateful speech is increasingly spoken rather than written, surfacing in voice notes, calls, and short-form videos. Most detection systems still transcribe speech to text before classifying it, but transcription is unreliable for languages lacking strong speech recognisers, and it discards the tone and emotion that often carry the abuse itself. This paper examines whether abusive speech can instead be detected directly from audio, using CLAP, a model that learns a shared representation of sound and language, evaluated across ten Indic languages in the ADIMA dataset. A lightweight classifier trained on CLAP's existing audio representations, without adapting the model itself, comes within one to three points of a fully supervised system, and far outperforms prompting with no labelled examples at all. Further adaptation with a handful of labelled examples per language yields little extra benefit, varying unpredictably ac

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents

arXiv:2604.04468v2 Announce Type: replace-cross Abstract: Evaluating retail strategies before deployment is difficult, as outcomes are determined across multiple stages, from seller-side persuasion through buyer-seller interaction to purchase decisions. However, existing retail simulators capture only partial aspects of this process and do not model cross-stage dependencies, making it difficult to assess how early decisions affect downstream outcomes. We present RetailSim, an end-to-end retail simulation framework that models this pipeline in a unified environment, explicitly designed for simulation fidelity through diverse product spaces, persona-driven agents, and multi-turn interactions. We evaluate RetailSim with a dual protocol comprising human evaluation of behavioral fidelity and meta-evaluation against real-world economic regularities, showing that it successfully reproduces key patterns such as demographic purchasing behavior, the price-demand relationship, and heterogeneous p

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation

arXiv:2603.17205v3 Announce Type: replace-cross Abstract: Domain-specific finetuning is essential for dense retrievers, yet not all data pairs contribute equally to the learning process. We introduce OPERA, a data pruning framework that exploits this heterogeneity to improve both the effectiveness and efficiency of retrieval model adaptation. We first investigate static pruning (SP), which retains only high-similarity query-document pairs, revealing an intrinsic quality-coverage tradeoff: ranking (NDCG) improves while retrieval (Recall) can degrade due to reduced query diversity. To resolve this tradeoff, we propose a two-stage dynamic pruning (DP) strategy that adaptively modulates sampling probabilities at both query and document levels throughout training, prioritizing high-quality examples while maintaining access to the full training set. Evaluations across eight datasets spanning six domains demonstrate the effectiveness of both approaches: SP improves ranking over standard finet

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation

arXiv:2607.28439v2 Announce Type: replace Abstract: Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation is costly and rater-variant, while LLM-as-a-judge is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users actually perceive the same interface. We propose the Evidence-Grounded, Social-Weighted Persona Panel (ESPP), a three-stage GenUI evaluation method in which a panel of psychologically diverse, evidence-grounded personas independently rates a screenshot, exchanges opinions under a trait-derived, semantically-gated bounded-confidence mechanism, and is aggregated via Delphi-inspired social weighting into a single judgment. ESPP tracks human judgment substantially more closely than a naive single-pass judge, raising Pearson $r$ from $0.716$ to $0.92

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution

arXiv:2607.13683v2 Announce Type: replace Abstract: Large Language Models (LLMs) have enabled capable agents across diverse applications. Beyond the foundation model, the performance of an agent is governed by the surrounding agent harness, including prompts, tools, control loops, etc. Automatically evolving this harness offers a promising pathway to agent improvement, yet existing approaches typically rely on greedy candidate selection and noisy self-generated feedback, rendering their gains susceptible to search collapse, task-specific overfitting, and poor verifiability. To tackle these challenges, we introduce HarnessBank, a trustworthy agent-harness self-evolution framework that pairs a task agent with a separate evolver agent for iterative failure diagnosis, harness generation, and evolution verification. HarnessBank maintains a Harness Gene Bank composed of high-performing harnesses of different semantic coordinates. Those harnesses are reinvented, recombined, screened, and sele

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence

arXiv:2606.21008v2 Announce Type: replace Abstract: The metanym game is a competitive word game for LLMs that measures structural intelligence against established cognitive-science constructs. No content is given in advance; the contestants create all of it -- a new kind of analogy test, analogical production falsifiable sentence by sentence, with no fixed test set to leak into training (contamination-resistant by construction). In the council-of-peers benchmark, the contestants also rate each other's creations. We introduce the first spectral solution, to our knowledge, to the wicked problem of benchmarking LLMs' factual accuracy without golden keys or oracle models: one singular value decomposition of the evaluators' ratings matrix yields their competence as both generators and judges of true statements at once. Competence on the subjective criteria comes from each judge's rating consistency as the yardstick shifts. The factual rating correlates with GPQA Diamond at Pearson r = 0.92.

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Implicit Reasoning for Large Language Model-based Generative Recommendation

arXiv:2606.14142v3 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly adopted as backbones for Generative Recommendation (GR), promising access to pretrained world knowledge. Yet reliably invoking this knowledge for GR remains poorly understood. A key obstacle is that LLM-based GR typically represents items with Semantic IDs (SIDs), disrupting LLMs' natural-language reasoning interface because these tokens are unseen by the LLM during pretraining. Existing approaches address this with expensive multi-stage pipelines that ground SIDs and elicit explicit rationales, but offer limited insight into when and why each stage is necessary. In this work, we systematically decompose explicit reasoning training pipelines for LLM-based GR, revealing three key limitations: weakened world-knowledge verbalization, misalignment between SID and natural-language token embedding spaces, and sensitivity to rationale quality, all of which hurt explicit reasoning performance. To

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Creative Integration: A Decidable Criterion of Creativity

arXiv:2606.13977v2 Announce Type: replace Abstract: "Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes the world cheaper to describe -- from a tidy re-description. Building on the lineage that treats creativity and intelligence as compression, we give such a criterion for creative integration (CI): the resolution of a real conflict between A and B is CI if and only if, under a fixed description language, the description length strictly shrinks (C = L_pre/L_post > 1), with the reduction located in the conflict itself. We make the judgment decidable through four binary, conjunctive gates, and we fix its extension through a taxonomy of pseudo-integration that names and rejects the look-alikes. We back the criterion with a curated, multi-domain corpus and -- crucially -- validate it not by human inter-rater agreement but by four falsifiable tests it could fail: an independent computational check, discrim

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

arXiv:2606.05176v2 Announce Type: replace Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-specific constraints in telecommunications customer support remain limited. In addition, data sovereignty, regulatory constraints, and the handling of sensitive customer and network information complicate the use of externally hosted foundation models in this domain. We present a systematic study of parameter-efficient fine-tuning (PEFT) using Low-Rank Adaptation (LoRA) applied to Qwen2.5-3B to build a domain-specific conversational assistant. We introduce a combinatorial synthetic data generation approach based on a glossary of 52 industry-specific terms, producing approximately 30,000 training examples across 1,560 distinct problem scenarios via a generative pipeline powered by Gemini 2.0 Flash. We evaluate 16 LoRA configurations by varying hyperparameters and target modules. Our eval

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining

arXiv:2606.01049v2 Announce Type: replace Abstract: Biomedical figures are explained not by captions alone but by body-text passages that discuss them. Yet current multimodal corpora typically reduce figures to isolated image-caption pairs, discarding this crucial context. Existing pipelines either omit this context or append it without enforcing the figure references that support each attachment, which can create unsupported image-text attachments and incoherent discourse. We introduce context-grounded reconstruction, a source-grounded framework that converts PubMed Central Open Access (PMC-OA) records into referentially coherent interleaved sequences. It recovers captions and source text, attaches context only through article-native figure references, repairs non-contiguous context, and prunes unsupported images. Starting from these reconstructed sequences, PMC-InterCPT first filters records for text quality and medical relevance, then applies evidence-aware allocation to form a 9.63

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

arXiv:2605.07699v2 Announce Type: replace Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that admit multiple valid interpretations. Despite the prevalence of such ambiguities in practice, existing agent benchmarks largely assume unambiguous, well-specified policies, leaving a critical evaluation gap. We introduce DRIP-R, a benchmark that systematically exploits real-world retail policy ambiguities to construct scenarios in which no single correct resolution exists. DRIP-R comprises a curated set of policy-ambiguous return scenarios paired with a realistic customer personas, a full-duplex conversational simulation with tool-calling capabilities and a multi-judge evaluation framework covering policy adherence, dialogue quality, behavioral alignment, and resolution quality. Our experiments show that frontier models fundamentally disagree on identical po

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Escaping Mode Collapse in LLM Generation via Geometric Regulation

arXiv:2605.00435v3 Announce Type: replace Abstract: Mode collapse is a persistent challenge in generative modeling and appears in autoregressive text generation as behaviors ranging from explicit looping to gradual loss of diversity and premature trajectory convergence. We take a dynamical-systems view and reinterpret mode collapse as reduced state-space accessibility caused by *geometric collapse*: during generation, the model's internal trajectory becomes confined to a low-dimensional region of its representation space. This implies mode collapse is not purely a token-level phenomenon and cannot be reliably solved by symbolic constraints or probability-only decoding heuristics. Guided by this perspective, we propose *Reinforced Mode Regulation* (RMR), a lightweight, online state-space intervention that regulates dominant self-reinforcing directions in the Transformer value cache (implemented as low-rank damping). Across multiple large language models, RMR substantially reduces mode c

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Disentangling Similarity and Relatedness in Topic Models

arXiv:2603.10619v3 Announce Type: replace Abstract: The recent success of large pre-trained language models (PLMs) has motivated their integration into topic modeling. However, PLM-augmented topic models differ from classical co-occurrence models such as Latent Dirichlet Allocation (LDA) not only in performance, but also in the type of semantic structure they capture. We formalize this distinction along two psycholinguistic axes: thematic relatedness (dog/bone) and taxonomic similarity (dog/wolf). To measure both axes over topic words, we construct a large synthetic benchmark of word pairs using LLM-based annotation and train a neural scorer on it. Across multiple corpora and model families, the scorer places different topic-model families at distinct positions within the joint similarity-relatedness space. The two scores further predict downstream task performance: tasks requiring similarity benefit from similarity-rich topics, whereas tasks requiring relatedness benefit from the conv

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery

arXiv:2603.03322v2 Announce Type: replace Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery. However, rigorously evaluating an AI's capacity for knowledge discovery remains a critical challenge. Existing benchmarks predominantly rely on static datasets, leading to inevitable data contamination where models have likely seen the evaluation knowledge during training. Furthermore, the rapid release cycles of modern LLMs render static benchmarks quickly outdated, failing to assess the ability to discover truly new knowledge. To address these limitations, we propose DBench-Bio, a dynamic and fully automated benchmark designed to evaluate AI's biological knowledge discovery ability. DBench-Bio employs a three-stage pipeline: (1) data acquisition of rigorous, authoritative paper abstracts; (2) QA extraction utilizing LLMs to synthesize scientific hypothesis questions and corresponding discovery answers; an

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation

arXiv:2601.22546v2 Announce Type: replace Abstract: The recent advancements in Large Language Models (LLMs) have attracted interest in exploring their in-context learning abilities and chain-of-thought capabilities. However, there are few studies investigating the specific traits related to the powerful generation capacity of LLMs. This paper aims to delve into the generation characteristics exhibited by LLMs. Through our investigation, we have discovered that language models tend to capture target-side keywords at the beginning of the generation process. We name this phenomenon the Holographic Characteristic of language models. For the purpose of exploring this characteristic and further improving the inference efficiency of language models, we propose a plugin called HOLO, which leverages the Holographic Characteristic to extract target-side keywords from language models within a limited number of generation steps and complements the sentence with a parallel lexically constrained tex

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

arXiv:2601.19827v5 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative retrieval-reasoning loops meaningfully outperform static RAG, particularly in scientific domains requiring multi-hop reasoning over sparse, heterogeneous evidence. We provide the first controlled, mechanism-level diagnostic evaluation of whether synchronized iterative retrieval and reasoning can surpass even an idealized static upper bound (Gold Context) RAG. We benchmark eleven state-of-the-art LLMs under three regimes: (i) No Context, measuring reliance on parametric memory; (ii) Gold Context, where all oracle evidence is supplied at once; and (iii) Iterative RAG, a training-free controller that alternates retrieval, hypothesis refinement, and evidence-aware stopping. Using the chemistry-focused ChemKGMultiHopQA dataset, we isolate questions requiring genuine retrieval and analyze retrieval coverage

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Knowledge Restoration-driven Prompt Optimization: Unlocking LLM Potential for Open-Domain Relational Triplet Extraction

arXiv:2601.15037v2 Announce Type: replace Abstract: Open-domain Relational Triplet Extraction (ORTE) aims to mine structured knowledge without predefined relation schemas. Large Language Models (LLMs) have advanced ORTE toward a prompt-driven paradigm through powerful in-context learning. However, adapting their extraction behavior to varying open-domain contexts remains challenging. Existing methods typically rely on manually crafted prompts that remain fixed across inputs, despite substantial variation in linguistic expressions and contextual structures. This mismatch may lead to unsupported triplets, while the absence of ground-truth annotations makes such deficiencies difficult to identify and correct. Moreover, free-form relation generation produces non-canonical relation surface forms, undermining knowledge graph consistency. To address these challenges, we propose Knowledge Restoration-driven Prompt Optimization (KRPO), a framework for label-free target-corpus adaptation. KRPO r

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

arXiv:2607.29613v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM backbone latents, which is a fundamental mismatch with the partially observable nature of robot control. A naive approach to incorporate observation history into the critic incurs exponential complexity with high-dimensional visual space, and still fails because pure scalar-return regression provides insufficient supervision for learning cross-temporal dynamics. We identify the root cause as a state approximation problem: without an explicit world modeling objective, the critic's representation cannot capture the temporal structure needed for accurate value estimation. To address this, we propose the World Critic Model (WCM), built on a lightweight LeJEPA architect

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings

arXiv:2607.29402v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems synergize retrieval mechanisms with generative language models to enhance the accuracy and relevance of responses. However, bridging the style gap between user queries and relevant information in document text remains a persistent challenge in retrieval-augmented systems, often addressed by runtime solutions (e.g., Hypothetical Document Embeddings (HyDE)) that attempt to improve alignment but introduce extra computational overhead at query time. To address these challenges, we propose Hypothetical Prompt Embeddings (HyPE), a framework that shifts the generation of hypothetical content from query time to the indexing phase. By precomputing multiple hypothetical prompts for each data chunk and embedding the chunk in place of the prompt, HyPE transforms retrieval into a question-question matching task, bypassing the need for runtime synthetic answer generation. This approach does not introduce l

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

arXiv:2607.29241v1 Announce Type: cross Abstract: Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes. While LLM-based agents can automate this trial-and-error process, allowing the LLM to both select modification directions and generate concrete hypotheses often leads to unstable search under limited experiment budgets. Inspired by the above challenge, we propose RecHarness, a Bandit-Routed Agentic Harness for automated recommender model optimization. RecHarness separates the optimization process into two steps: a bandit router selects the next modification direction according to historical validation feedback, while the LLM generates a concrete optimization hypothesis and executable code edit within the selected direction. To sustain long-horizon exploration, RecHarness uses a jump-basin mechanism to activate a structural-jump arm when local edits stagnate. Across multiple recommendati

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

TransMem: Transforming Hidden States into Memory for Large Language Models

arXiv:2607.29032v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly operate over long interaction histories, where effective reasoning requires identifying and exploiting task-relevant evidence distributed across past observations and actions. However, useful information encoded in previously computed representations is often underutilized during subsequent generation. We propose \textbf{TransMem}, a lightweight inference-time parametric memory module that transforms sparse historical hidden states from a frozen LLM backbone into reusable memory representations. TransMem uses a lightweight gating network to dynamically apply the latent intervention to the current hidden states, without repeatedly encoding the preceding context. To learn transferable memory utilization rather than task-specific knowledge, we introduce evidence-conditioned self-distillation. A memory-augmented student processes the full context and matches the predictive distribution of an ev

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

GoldenRetriever: Non-Interactive Homomorphic Encrypted Retrieval for Privacy-Preserving RAG

arXiv:2607.29019v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances large language models by incorporating external knowledge, but existing pipelines typically operate on plaintext data, raising significant privacy concerns. Prior work on privacy-preserving retrieval leverages cryptographic techniques such as homomorphic encryption (HE) and private information retrieval (PIR), but often relies on interactive protocols or ranking-based selection mechanisms that incur high latency and potential information leakage. In this paper, we propose a practical non-interactive encrypted retrieval framework for RAG based on threshold selection. Instead of performing expensive top-$k$ ranking under encryption, our approach selects documents whose similarity scores exceed a predefined threshold, reducing computational complexity from quadratic to linear in the corpus size. We implement this design using CKKS-based homomorphic computation, enabling fully encrypted similari

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning

arXiv:2607.28986v1 Announce Type: cross Abstract: Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers. Existing retrieval-augmented methods score image-text alignment once, at retrieval, then commit the captioner's autoregressive beam under language-model probability alone, leaving the decoder without further visual grounding feedback. Progress has stalled, with no method improving on the strict-regime best since 2024. We propose Adjudicated Captioning, an inference-time multi-agent framework that restores grounding feedback at multiple checkpoints over an unchanged IFCap captioner. First, we install a stronger frozen Retrieval Encoder at the input. Second, between retrieval and decoding we insert a frozen Cross-Attention Verifier that re-ranks the top-9 retrievals to top-5. Third, at the output beam we attach a learned Reranker pairing TriFuse, a mult

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

arXiv:2607.28896v1 Announce Type: cross Abstract: Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about the same audio? Current practice evaluates each capability in isolation on specialized benchmarks, and never asks whether a model can make sense of its own generations. We present TORUS, the first self-coherence test for audio-native unified models. TORUS comprises 48 three-stage self-coherence tests carrying 432 six-option questions spanning speech, sound and music across five task families. We holistically evaluate five open unified models alongside a Cascaded Baseline that combines state-of-the-art specialized generation, editing and understanding models. The best unified model answers 50.5% of questions against the Cascaded Baseline's 63.2% and a 16.7% chance floor. Models struggle on audio editing. Among t

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

arXiv:2607.28818v1 Announce Type: cross Abstract: As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajector

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

arXiv:2607.28692v1 Announce Type: cross Abstract: Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on predefined tool spaces with static semantics limits their applicability to open-world scientific workflows, where tool requirements, capabilities, and boundaries evolve dynamically. To this end, we propose SciToolAgent-Evo, an ontology-aware self-evolving agent for open-world scientific tool acquisition. Driven by an evolving memory of skills, experiences, and an ontologized tool graph, it distills generalizable knowledge from contrastive trajectories during accumulation, whereas during inference, it formulates active requests and utilizes a LinUCB-based bandit gate to dynamically balance exploration and exploitation. Once a novel tool is acquired, its scientific ontology is completed online for seamless integration into the known graph. Moreover, we introduce Ope

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories

arXiv:2607.28674v1 Announce Type: cross Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods rely on output-level signals or collapse processing depth into a single trajectory-level scalar, leaving step-wise effort opaque. We propose Step-Aware Reasoning Energy (SARE), a geometric framework that quantifies effort at the granularity of individual CoT steps via Centered Kernel Alignment (CKA) between Gram matrices of token hidden states across adjacent transformer layers, capturing inter-token relational structure without requiring eigenvector alignment or cluster correspondence. SARE further contextualizes this energy within reasoning's semantic progression by modeling CoT trajectories as transitions among latent semantic states. Across six reasoning benchmarks and three open-weight LLMs, we find that reasoning energy is highly non-uniform across step types, exh

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning

arXiv:2607.28642v1 Announce Type: cross Abstract: Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue that under bounded context windows, the core bottleneck is not trajectory compression or test-time control, but the absence of a reusable intermediate interface that can replace discarded history and support continued solving. We further identify a key failure mode of outcome-reward-driven long-chain reinforcement learning: when the model has not solved the task before the window is nearly exhausted, the final-answer reward encourages premature guessing rather than continued careful reasoning. We propose ThinkReset, a text-space instantiation of this view. ThinkReset explicitly constructs reusable intermediate interfaces through interface writeback and reset, and directly optimizes post-reset continuation success. Across multiple long-horizon reasoning benchmarks, this

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

arXiv:2607.28631v1 Announce Type: cross Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI-generated papers remains an open challenge. We propose and implement a rigorous benchmarking protocol using an automated peer-review system that harnesses frontier large language models to assess scientific papers across four core dimensions: originality, scientific rigor, clarity, and significance. We evaluate four leading AI Scientist frameworks: \textit{Sakana AI (v1 & v2)}, \textit{CycleResearcher}, and \textit{Data-to-Paper}. Each framework was run on a consistent set of 15 research proposals published by a commercial autonomous AI scientist company (FARS), generating 60 papers that we evaluate alongside 15 FARS benchmark papers. Using three independent LLM reviewers (GPT-5.4, Gemini, and Claude), we find that FARS benchmark papers significantly outperform all c

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Evidence-Ledger Adjudication for Claim-Evidence Traceability

arXiv:2607.26512v1 Announce Type: cross Abstract: AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a claim-evidence traceability workflow that pairs each claim with an evidence packet, assigns a support relation, and routes unsupported, contradicted, or mixed-evidence claims back to the author. The empirical core is a 2,335-row blind benchmark built from independent external labels in AVeriTeC, CLIMATE-FEVER, and SciFact. Gold relations and source evidence labels are hidden during prediction and joined only for scoring. On this benchmark, the agent evidence-ledger condition achieves 0.676 relation accuracy and 0.601 macro-F1, compared with 0.383 accuracy and 0.303 macro-F1 for the best non-agent baseline. It also routes 1270/1435 claims whose gold labels indicate contradiction, missing evidence, or mixed evidence, while routing 295/900 supported claims. These results show that evidence

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small tool result, and reuse is hard because even a short append can change token boundaries near the end of the previous sequence. Across 153,951 calls from two agent ecosystems, the median call appends about 1.4K characters, and only 1.0-3.6% of calls start or rebuild a session with contexts of millions of characters. At a 94.1% fleet prompt-cache hit rate, tokenization reaches up to 64% of time to first token. TokTier is a stateful tokenization service with one contract: emitted token IDs are always identical to full reference tokenization of the request text. For a session continuation, it re-tokenizes a small window around the append and splices only after a per-request stable-boundary check, widening the window or falling back to full tokeni

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CL

Evolving language compositionality in a frequency-structured meaning space

arXiv:2607.29642v1 Announce Type: new Abstract: The iterated learning model was introduced to investigate language evolution: the way in which the characteristic properties of human languages have been shaped, at least partly, by repeated transmission from one language user to another. The key finding is that language compositionality can arise spontaneously as a consequence of language being passed repeatedly through a language learning bottleneck. Here we explore how changing the frequency of different meanings, so that some meanings occur much more frequently than others, affects the character of its compositionality. We find that, as observed in natural languages, high-frequency meanings can escape the pressure to conform to the grammar that characterizes lower-frequency meanings. However, when the frequency structure is instead imposed on parts rather than on whole meaning vectors, the language fails to transmit across generations. This occurs despite the fact that the most freque

Source ↗
Showing 10401–10450 of 18694 signals
← Prev Page 209 of 374 Next →