Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:20:20.407043Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2510.10541.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:20:20.407043Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-25T21:15:07.735336Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T19:30:07.857950Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4b1a2395-b57b-4d82-90e2-ecbfe2981f0c · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Training Verifiers to Solve Math Word Problems
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed62fb8-7d48-44dc-98be-f3aef646042c · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Shortcut learning in deep neural networks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a346905-9ddc-414e-bf1d-1a8080810ceb · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Measuring massive multitask language understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56a3c550-052d-4348-b875-16894bba678d · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Adversarial Examples for Evaluating Reading Comprehension Systems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6fa7254-93ae-453e-b732-8cbf73c1e452 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Solving Quantitative Reasoning Problems with Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 433ff859-2a4e-4de6-b343-f9853686de99 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac220ade-97f4-47ec-ae94-9b163ee48c29 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Let's verify step by step
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ec04a5c-12af-419c-a04b-721bdc6b9759 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdbde81f-bb83-4466-8f7c-4181ea1a7913 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Training language models to follow instructions with human feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebcd2427-c0b6-4803-bf06-5750abe77b37 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Curriculum reinforcement learning from easy to hard tasks improves llm reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db9d3758-fb1d-40ca-a4a0-5cb6970d7162 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Direct preference optimization: Your language model is secretly a reward model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 983eeed1-c37b-46be-919d-6036b7315bb1 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Proximal Policy Optimization Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747ce24d-dcb6-4120-835c-625de9a36168 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e73c85-0fb2-4f3b-a262-063dcc4ff5a1 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? HEAD-QA: A Healthcare Dataset for Complex Reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e01159f6-6375-4e28-975d-3c8078019ee7 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f1baf7b-1d70-40de-a153-53e171474b09 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Chain-of-thought prompting elicits reasoning in large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c7ad0c6-f9cd-4d50-9f3f-88f64a9e3845 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7fc6145-cb8b-4ef3-9e68-4cf5b78681d8 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Frame: Feedback-refined agent methodology for enhancing medical research insights
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3795470-4370-46d5-9301-1677b24c1bc2 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 664256ce-62d2-4b7a-9711-f17a87d68503 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85245822-8158-4d77-a0bb-7676a8d9b49a · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Easyr1: An efficient, scalable, multi-modality rl training framework
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af54b007-148c-4ebc-bc7c-6ce0efd556a9 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? write newline
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32270c72-0952-4a5f-aea5-d7edcf434404 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? @esa (Ref
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 700a9b75-6981-47c6-ba3c-aea8b9ed55e9 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce937af4-a397-4290-9587-029f73a6da46 · outbound
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8fbd985-0893-40d4-a625-a2a3988894ea · inbound
Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.