Pith. sign in

Paper Citation Record · LEDGER

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

As of 12 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2510.10541.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.10541 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:20:20.407043Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-25T21:15:07.735336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T19:30:07.857950Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b1a2395-b57b-4d82-90e2-ecbfe2981f0c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Training Verifiers to Solve Math Word Problems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.047448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.047448Z digest=sha256:d6ed67b2d4d0c7e0beb4e9ae68a50d66c46118065a2cf7aa39269c1760acdd65

Observation 3ed62fb8-7d48-44dc-98be-f3aef646042c · outbound

This paper cites Shortcut learning in deep neural networks.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Shortcut learning in deep neural networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.121512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.121512Z digest=sha256:e851b295d1c8d5a5e642b18f3f038429c72a016216c00f381b4e6a3aa799a252

Observation 9a346905-9ddc-414e-bf1d-1a8080810ceb · outbound

This paper cites Measuring massive multitask language understanding.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Measuring massive multitask language understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.270251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.270251Z digest=sha256:f948cc4a21f1030d622445b77dd7943d46a78404afe8be5a7d83eaea67d10fc6

Observation 56a3c550-052d-4348-b875-16894bba678d · outbound

This paper cites Adversarial Examples for Evaluating Reading Comprehension Systems.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Adversarial Examples for Evaluating Reading Comprehension Systems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.361180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.361180Z digest=sha256:ffae968e4c8ea4b4fc7118b020510090ea598797d0600f555adc5bea471295f7

Observation a6fa7254-93ae-453e-b732-8cbf73c1e452 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Solving Quantitative Reasoning Problems with Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.458587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.458587Z digest=sha256:1b00a7cfa96db8ae744575086090268054db1641284ceaac99ae27be22ffd6a9

Observation 433ff859-2a4e-4de6-b343-f9853686de99 · outbound

This paper cites Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.524173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.524173Z digest=sha256:3e757a6cfff86dffa9e169647673ebc0976873e11ff7b5e9540bbf169ffb34c0

Observation ac220ade-97f4-47ec-ae94-9b163ee48c29 · outbound

This paper cites Let's verify step by step.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Let's verify step by step

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.590115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.590115Z digest=sha256:7a0076311eb38b84cb44f8179b27e9ed1d14d12cd4578baf6f1613465e444907

Observation 0ec04a5c-12af-419c-a04b-721bdc6b9759 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl, 2025.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.665265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.665265Z digest=sha256:dfa1a6a0ad71e9840611042cbc050e26bcfcd70728dcc218a3d51ae687092e9c

Observation cdbde81f-bb83-4466-8f7c-4181ea1a7913 · outbound

This paper cites Training language models to follow instructions with human feedback.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Training language models to follow instructions with human feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.758528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.758528Z digest=sha256:ab492f452ede6f3aed6aee7cac6dda54d6c2f6e3bf55c63d290efeb4df65ff03

Observation ebcd2427-c0b6-4803-bf06-5750abe77b37 · outbound

This paper cites Curriculum reinforcement learning from easy to hard tasks improves llm reasoning.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Curriculum reinforcement learning from easy to hard tasks improves llm reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.878037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.878037Z digest=sha256:29790de7195a3d926a5a2be529867e15c9966dd50f23c9a66be71c454f3995f5

Observation db9d3758-fb1d-40ca-a4a0-5cb6970d7162 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Direct preference optimization: Your language model is secretly a reward model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.993321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.993321Z digest=sha256:a2f8c41c2cc9716f6b90aecf7375f557cf546bd9d7d9f4b6eebfb9747ea9b570

Observation 983eeed1-c37b-46be-919d-6036b7315bb1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Proximal Policy Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.123246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.123246Z digest=sha256:68ce216bb20cdb3d522de118f9c27cbae815d007007d40328dfbe0b3e158399d

Observation 747ce24d-dcb6-4120-835c-625de9a36168 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.245701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.245701Z digest=sha256:4fb3e616895e516572d0a85e0779fe2c9f4aa502d84f226648f965b92af72d58

Observation 47e73c85-0fb2-4f3b-a262-063dcc4ff5a1 · outbound

This paper cites HEAD-QA: A Healthcare Dataset for Complex Reasoning.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? HEAD-QA: A Healthcare Dataset for Complex Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.327712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.327712Z digest=sha256:21d937d7dbe4cc63de5f294a588af89ec45c60a87ac79d4f7afcf157f48f8932

Observation e01159f6-6375-4e28-975d-3c8078019ee7 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.428770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.428770Z digest=sha256:8ff8e3024e40733908fc70e650f6f5ffb1a31ec06a9ce8118416481407c3cde9

Observation 5f1baf7b-1d70-40de-a153-53e171474b09 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Chain-of-thought prompting elicits reasoning in large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.588462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.588462Z digest=sha256:a77a45cffcf6f8d5655fef76ac36b3d4e3364caa42e18d39de62ca29e3414397

Observation 9c7ad0c6-f9cd-4d50-9f3f-88f64a9e3845 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.721870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.721870Z digest=sha256:e565b48eea2b0b792f4ee874732030925f32bd2fcea601b9367c4a9ad8c938aa

Observation e7fc6145-cb8b-4ef3-9e68-4cf5b78681d8 · outbound

This paper cites Frame: Feedback-refined agent methodology for enhancing medical research insights.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Frame: Feedback-refined agent methodology for enhancing medical research insights

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.865669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.865669Z digest=sha256:d3f14e86bba587f5ea555de0603294c0a2e30fa06b76d752122cf2bba818bb58

Observation d3795470-4370-46d5-9301-1677b24c1bc2 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.991179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.991179Z digest=sha256:f8c84b8edb8f1d855ee0b767202823900e7fb162ea3edde09a7dabed9060d883

Observation 664256ce-62d2-4b7a-9711-f17a87d68503 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.043540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.043540Z digest=sha256:8311edf35aaf699e27069afdb0b1d8836c6c201ba7c76097d9a65ff762b6a75e

Observation 85245822-8158-4d77-a0bb-7676a8d9b49a · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.091347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.091347Z digest=sha256:79b1b502a8871b6f0399fe4c68648fcd00fe7a14e7f56591193036fd74cd9563

Observation af54b007-148c-4ebc-bc7c-6ce0efd556a9 · outbound

This paper cites write newline.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? write newline

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.207076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.207076Z digest=sha256:c3797b2a096f0e88db70d3190e729f855a6395bf3286d1f53bf84d71bdb49dc0

Observation 32270c72-0952-4a5f-aea5-d7edcf434404 · outbound

This paper cites @esa (Ref.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? @esa (Ref

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.302554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.302554Z digest=sha256:a42d6c7e99b2f6d75a24f5d922eed164edc4cfbda729100691a14d0bad6f5c21

Observation 700a9b75-6981-47c6-ba3c-aea8b9ed55e9 · outbound

This paper cites an unresolved cited work.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.377723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.377723Z digest=sha256:cfb45468332826346b9ff8ea34f0a7264e1239e3ac791b91a9735fbe5059267b

Observation ce937af4-a397-4290-9587-029f73a6da46 · outbound

This paper cites an unresolved cited work.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.407043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.407043Z digest=sha256:1629e696ec344f4121d899290c3e66b17966eb52ee1e190af0fa0fa7e512c50e

Pith citing papers

Observation a8fbd985-0893-40d4-a625-a2a3988894ea · inbound

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors cites this paper.

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:30:07.859170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T21:15:07.735336Z digest=sha256:e480e175a3c0277bc29acec238e73ac764cd3d8ed34e602d7a8e1694b5e6c58f