Pith. sign in

REVIEW 2 cited by

SEM: Reinforcement Learning for Search-Efficient Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.07903 v1 pith:D7KPWI3O submitted 2025-05-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords searchexternallearningmodelmodelsreasoningreinforcementwhen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in Large Language Models(LLMs) have demonstrated their capabilities not only in reasoning but also in invoking external tools, particularly search engines. However, teaching models to discern when to invoke search and when to rely on their internal knowledge remains a significant challenge. Existing reinforcement learning approaches often lead to redundant search behaviors, resulting in inefficiencies and over-cost. In this paper, we propose SEM, a novel post-training reinforcement learning framework that explicitly trains LLMs to optimize search usage. By constructing a balanced dataset combining MuSiQue and MMLU, we create scenarios where the model must learn to distinguish between questions it can answer directly and those requiring external retrieval. We design a structured reasoning template and employ Group Relative Policy Optimization(GRPO) to post-train the model's search behaviors. Our reward function encourages accurate answering without unnecessary search while promoting effective retrieval when needed. Experimental results demonstrate that our method significantly reduces redundant search operations while maintaining or improving answer accuracy across multiple challenging benchmarks. This framework advances the model's reasoning efficiency and extends its capability to judiciously leverage external knowledge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A DAPO-trained 4B model modestly improves selective evidence adoption on SelectBench-v2, but the gains are not statistically robust and prompt-injection resistance does not improve.

  2. Agent Safety Alignment via Reinforcement Learning

    cs.AI 2025-07 reject novelty 5.0 of 10

    RL-based safety alignment with an execute-refuse-verify policy improves reported threat resistance for tool-using agents, but utility preservation is not consistently demonstrated.

Pith tools