Pith. sign in

Paper Citation Record · LEDGER

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

As of 13 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2510.10541.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.10541 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:20:20.407043Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-25T21:15:07.735336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T19:30:07.857950Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b1a2395-b57b-4d82-90e2-ecbfe2981f0c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Training Verifiers to Solve Math Word Problems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.047448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.047448Z digest=sha256:3d353b501343620db4c4cf0d9eb01495a8bf5758924abdd8b1ab4f7ea556dd1b

Observation 3ed62fb8-7d48-44dc-98be-f3aef646042c · outbound

This paper cites Shortcut learning in deep neural networks.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Shortcut learning in deep neural networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.121512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.121512Z digest=sha256:1aad71161d32b74153f5713df9286bc9e3f1cced5ee31e279393ebfdda5cae2f

Observation 9a346905-9ddc-414e-bf1d-1a8080810ceb · outbound

This paper cites Measuring massive multitask language understanding.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Measuring massive multitask language understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.270251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.270251Z digest=sha256:b558398166b6a7823480e7495de46e0cef656523cc45da677b432f44fa101a28

Observation 56a3c550-052d-4348-b875-16894bba678d · outbound

This paper cites Adversarial Examples for Evaluating Reading Comprehension Systems.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Adversarial Examples for Evaluating Reading Comprehension Systems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.361180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.361180Z digest=sha256:deeca60e75799ff6e1b099b870dcd555246904e722556e4a7c45522b4750fa00

Observation a6fa7254-93ae-453e-b732-8cbf73c1e452 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Solving Quantitative Reasoning Problems with Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.458587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.458587Z digest=sha256:ab65340738c36aed2b55af14c0e37aa6f845342a98cb699666c92bdbd0f5e492

Observation 433ff859-2a4e-4de6-b343-f9853686de99 · outbound

This paper cites Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.524173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.524173Z digest=sha256:5cc9024efd78343461d33142c3e553ddb7762df5894fa0b56cb6ff93513b65a2

Observation ac220ade-97f4-47ec-ae94-9b163ee48c29 · outbound

This paper cites Let's verify step by step.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Let's verify step by step

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.590115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.590115Z digest=sha256:fdb0c916df62d25755c0f6dc6fff7b30704029c9b152e77b65f0c3f360f44ed7

Observation 0ec04a5c-12af-419c-a04b-721bdc6b9759 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl, 2025.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.665265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.665265Z digest=sha256:c831392b6e2bfd563e47b8715b3f84127bec88006d8921664031d84b81723a50

Observation cdbde81f-bb83-4466-8f7c-4181ea1a7913 · outbound

This paper cites Training language models to follow instructions with human feedback.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Training language models to follow instructions with human feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.758528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.758528Z digest=sha256:54c17013e10ffec3bf7aa8532f468105e092d4f4f9a398a2c74ff43b669e8976

Observation ebcd2427-c0b6-4803-bf06-5750abe77b37 · outbound

This paper cites Curriculum reinforcement learning from easy to hard tasks improves llm reasoning.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Curriculum reinforcement learning from easy to hard tasks improves llm reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.878037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.878037Z digest=sha256:319d6735e6c31adf830c661a1cb4d606bdccde205a2372a53c548a1660051326

Observation db9d3758-fb1d-40ca-a4a0-5cb6970d7162 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Direct preference optimization: Your language model is secretly a reward model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.993321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.993321Z digest=sha256:f856720f7bdbc4cb4b6b0f4daefc778b92d54db47edba8096b90f28d75ab47ca

Observation 983eeed1-c37b-46be-919d-6036b7315bb1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Proximal Policy Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.123246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.123246Z digest=sha256:1371a2d832e708f1117b1759444b621cf993317752731a497ecf5d6882060c60

Observation 747ce24d-dcb6-4120-835c-625de9a36168 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.245701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.245701Z digest=sha256:2f6a006500b3e75ebea9e98d34aee4791bd1ed5dd65d886c47b84c21b231c2e1

Observation 47e73c85-0fb2-4f3b-a262-063dcc4ff5a1 · outbound

This paper cites HEAD-QA: A Healthcare Dataset for Complex Reasoning.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? HEAD-QA: A Healthcare Dataset for Complex Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.327712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.327712Z digest=sha256:3a1653cd93685ab9c428ee802b93710d877ee181a9123bc88f87830d4fdf8126

Observation e01159f6-6375-4e28-975d-3c8078019ee7 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.428770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.428770Z digest=sha256:5fdde0d3a7d0a17c6855c29204c79ecdba7932920d56eaf28ca23e0313545dca

Observation 5f1baf7b-1d70-40de-a153-53e171474b09 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Chain-of-thought prompting elicits reasoning in large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.588462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.588462Z digest=sha256:555c7b9a2eef039acfac39484935d296eac1d1383526b7275f78cf40e0fcce72

Observation 9c7ad0c6-f9cd-4d50-9f3f-88f64a9e3845 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.721870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.721870Z digest=sha256:a5a82f2f7ee9eead8e64360ddc472154e9b460d660283b508eda47ffe0850a74

Observation e7fc6145-cb8b-4ef3-9e68-4cf5b78681d8 · outbound

This paper cites Frame: Feedback-refined agent methodology for enhancing medical research insights.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Frame: Feedback-refined agent methodology for enhancing medical research insights

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.865669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.865669Z digest=sha256:2fefadd6ec2b09cdb45555003743b9f565ae9eb7f598d84bab1265cdf9dfd6ad

Observation d3795470-4370-46d5-9301-1677b24c1bc2 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.991179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.991179Z digest=sha256:628c1e0a8aeb38f0f3415602e0b9e8dc64c3aff0bca5233968aa86afbb95c8b6

Observation 664256ce-62d2-4b7a-9711-f17a87d68503 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.043540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.043540Z digest=sha256:8fbf53499c8992e313e192897ae81de9cef7fc64fe4663183450f52ce673ddd1

Observation 85245822-8158-4d77-a0bb-7676a8d9b49a · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.091347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.091347Z digest=sha256:aa43a9c2a242fc8fb3f278288606d5744d05324de1e7bdf89195cbd0746bf363

Observation af54b007-148c-4ebc-bc7c-6ce0efd556a9 · outbound

This paper cites write newline.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? write newline

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.207076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.207076Z digest=sha256:72032e787394849d61452290b2c1258e9ec17bc3cba6b596a169f61503dfbd42

Observation 32270c72-0952-4a5f-aea5-d7edcf434404 · outbound

This paper cites @esa (Ref.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? @esa (Ref

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.302554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.302554Z digest=sha256:7090b93ec7b3bf2bd35dea8c2f455aa0a9a8d3741c3b1fa58566ec60ba136a08

Observation 700a9b75-6981-47c6-ba3c-aea8b9ed55e9 · outbound

This paper cites an unresolved cited work.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.377723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.377723Z digest=sha256:613be84b7fbee7f67ef5a1942b38569adeaca12f0ce736f32f61a475a4279241

Observation ce937af4-a397-4290-9587-029f73a6da46 · outbound

This paper cites an unresolved cited work.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.407043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.407043Z digest=sha256:5f65c5fa677d6a1dbc08249bb7d0d0b7cb74d35096f8e972f04ca3dc4f529d70

Pith citing papers

Observation a8fbd985-0893-40d4-a625-a2a3988894ea · inbound

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors cites this paper.

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:30:07.859170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T21:15:07.735336Z digest=sha256:d5d53f5a97abbe48deab58688dfcfda65e2913359ef636531787283d8882badc