Pith. sign in

Octothinker: Mid-training incentivizes reinforcement learning scaling.arXiv preprint arXiv:2506.20512, 2025b

4 Pith papers cite this work, alongside 10 external citations. Polarity classification is still indexing.

4 Pith papers citing it
10 external citations · OpenAlex

fields

cs.CL 4

years

2026 3 2025 1

verdicts

UNVERDICTED 4

representative citing papers

MASH: Modeling Abstention via Selective Help-Seeking

cs.CL · 2025-10-01 · unverdicted · novelty 6.0

MASH uses RL with a pay-per-search reward to make LLMs seek external help only when needed, improving multi-hop QA accuracy by 7.6% and enabling competitive abstention without pre-defined knowledge boundaries.

citing papers explorer

Showing 4 of 4 citing papers.