Pith. sign in

Paper Citation Record · LEDGER

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data

As of 5 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 1 inbound Pith citation observation for arXiv:2604.18493.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.18493 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T05:32:23.972335Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:04:51.970049Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T14:28:31.391729Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact15
  • verified fuzzy19
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch42

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d982aff-524c-4a4b-99f9-06491df24d12 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data TTRL: Test-Time Reinforcement Learning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:02.011777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:8928121948ad3c5e0a991482e53a3b49b71973cc2e9439be039c080113d65d3e

Observation f7f0a398-e21b-4adc-858e-2091dc4b56c2 · outbound

This paper cites ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:02.052658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:17c05a0cb8d434ed0fb0265a4afac11ba07bc322fcbf34b7f1c3c379b31374d6

Observation 43f1f110-ab36-4a96-bb50-dd20d000d334 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T05:36:02.059877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:402344516686e28a15cbbffb78b866df65b946b4616b59b5304e06fa8213488b

Observation a77039c5-391b-46f0-961b-c9bd3e6d4e0a · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:2c99d6b5e1377d5880b9a39251d835ed54d58dd3df9440d8e2427df002349565

Observation 95ea86d6-6133-45e6-9987-8b5466e5ed5a · outbound

This paper cites One Token to Fool LLM-as-a-Judge.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data One Token to Fool LLM-as-a-Judge

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-12T02:08:19.458599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:77d38fe964408e3fee09016edd906d260c779e91f850d78aff6051da45ef752b

Observation fdd4da0d-4987-4e9a-b13e-ced8e7368487 · outbound

This paper cites OpenAI o1 System Card.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data OpenAI o1 System Card

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T05:36:02.050328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:2dac921ce477bfe5724d33d3a16e79ed50e3708444e8cf401007a8aa12a52267

Observation 47a0b006-db2a-4212-8b40-51b8c559e048 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T05:36:02.055017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:8cbe147040a21a59b06351606daf1f17b5c156af6fcce3f6c118ed61474575ab

Observation 6b320650-b316-425c-9ddd-a63262182404 · outbound

This paper cites Qwen3 Technical Report.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Qwen3 Technical Report

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T05:36:02.019262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:d3bf3e3fff9ac4cfaedd7344222f9fe747d3c63cf0dc2c5254fed28115d1c9be

Observation 5b18a2b9-e4fb-4e0e-ba7f-197c79a03bbb · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T08:21:05.971096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:b7870220e307450302fe813c0021567f1f36fa6347cfaa2f29787359d09109f5

Observation d67d51a4-c9e1-4ea0-94f3-2bcf7f5312e0 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T12:12:09.004283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:20d458031e5c0a56a21d63173093b07076f537044c95969b6f88007915b3c277

Observation 52f3cb1b-099f-4c53-9e14-867cb07c84b5 · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Reinforcing General Reasoning without Verifiers

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:02.036290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:853d28c8e06c1458c12f5a59b1da70c78f4509deec83e80263fb4fc3f87bee31

Observation b8b20a9b-7f99-42a0-9d14-1775e1bc2563 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T12:31:34.026635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:abfb7bfd001b2abef85d992ac4474fddc1f5de4cb00d3d1eb32d66ac7d5f0b26

Observation c5a719de-c496-4442-952c-34eb2b8251e0 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Maximizing Confidence Alone Improves Reasoning

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:02.040925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:0a7fc99571e8687e5c536991fb7ba71a380de7260de433ed13dd229005ad6694

Observation 21d385bb-39f7-49c0-b23a-75d66eacf57b · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T15:58:33.819070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:b7efe70ede88063f871267e8b3861e5cc422f645031422b8508c78ded594d5ca

Observation 98aa62d3-2ba4-46c1-a96c-2b8cb987701b · outbound

This paper cites Learning to Reason without External Rewards.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Learning to Reason without External Rewards

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:16:57.297836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:c4a890c220b0a6085694ab8ef5df47e1ea8913ed12548f0c47a1735fcbe19a6c

Observation b03f0231-88d5-41b8-972d-ca08f6ba4319 · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:02.057418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:39cc8d014a6bc6d032439a4f6c4d3b9c6571cafae21b94315d2a782f295eee36

Observation 9d56d6ae-3ad8-4349-a497-dbd30086eec0 · outbound

This paper cites Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.038742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:411134eca948125e068db79be94b2881b76dca76a58229aa2007e800de1950a3

Observation 0ebb5673-a95c-46c8-b46b-c61a588a9abb · outbound

This paper cites s1: Simple test-time scaling.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data s1: Simple test-time scaling

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:02.024064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:f2095b1489e7a80b63c32dad208a8aa64a7f68dd5edab957edb7f98d387edb88

Observation 696efccc-7a89-405e-80f4-a3c7a869dc1c · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:02.016993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:2fe580e397a881088675f1d06589d5ccc9bee9b327c761348ba27eb25a1cc233

Observation 9ea96b77-9e51-4be1-b951-b8c924bb42b2 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:23:58.476668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:de8eabfc8e05372deb4f11e63c565b7fb3337615a005049c61d79bbec80d6f15

Observation 8ae0e32a-3386-4927-a0a5-ec71e8f0c24c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T05:36:02.026279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:77d452fed8b7b338f94ca4efd273320e7946a99cb5bca53004d7db7001631550

Observation 5e6d2d1b-a4e5-46e4-9faf-25570ad0a42e · outbound

This paper cites s1: Simple test-time scaling.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data s1: Simple test-time scaling

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.399505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:5af29a1f91376053cdee95044a59f9eecdc237220f11d60da678af731ee29f1a

Observation 97b7fe68-7dfb-4165-bc4e-0cade0795304 · outbound

This paper cites e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.402466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:0c472049a9c9265b4f95b6fbb2ef3998dbca9dabeae3d567ecdccae1b70c4fb1

Observation f9113143-e73d-4457-9c01-23d025e5bb2b · outbound

This paper cites Can large reasoning models self-train?.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Can large reasoning models self-train?

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.968236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:dd6d47979859d5ab171b6a966a33412d3a5f6ce2668ec1b5ca892626ed97a659

Observation 89cda544-1c8e-48f0-a479-3f16ac20cc81 · outbound

This paper cites First Conference on Language Modeling , year=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data First Conference on Language Modeling , year=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.149853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:0a8387df6ff0b2e694bdd0c147c5116b343d3ea132b6a6acb859661fe77efbfc

Observation 4831cb6f-d040-423c-a5b9-d941ebdeb8cf · outbound

This paper cites Hugging Face repository , volume=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Hugging Face repository , volume=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.156222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:0b7c1ba638762876a37bbcc103dafc5f5b24d558547dff68ae1ace9a9da5d043

Observation 98b3b3c6-30d8-4225-bfa1-75cf3791759a · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Measuring Mathematical Problem Solving With the MATH Dataset

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:00:39.484061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:3822d2cf9fa9a839464894f7782dc875ed7364fddda1eda15d19c9decdb16e10

Observation 55898f63-c9ab-42a1-9e93-f14d32339cfc · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.975383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:4cba5fbe62d7b1b58226f8361caf38e77b2a154c9ba4cded3cc40a844c31f0d4

Observation 26110588-8257-45f7-bbd3-bde092241b05 · outbound

This paper cites International conference on machine learning , pages=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data International conference on machine learning , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.139626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:eb739bcc5371d05e28fc561d027d210cc2e9f26976a062f9cf98a7e181712a0d

Observation c97915af-e7eb-4acc-b148-7d9ad74b52a5 · outbound

This paper cites Tent: Fully Test-time Adaptation by Entropy Minimization.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Tent: Fully Test-time Adaptation by Entropy Minimization

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:09:24.402331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:ec7cbcf589c5b97ff1f192978d4e0f4f416bc14dec322059bb9af16ad7f3e700

Observation 534b6c4f-b157-46aa-ba63-5fe0fa65c371 · outbound

This paper cites Advances in neural information processing systems , volume=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Advances in neural information processing systems , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.137596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:334bda3ff595468c8cf0312b5ee32631a1d6d8752157af05c2343c6b7e1f95c9

Observation 7ac4f40b-16d9-4972-8feb-a9c4b7a41e7a · outbound

This paper cites Workshop on challenges in representation learning, ICML , volume=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Workshop on challenges in representation learning, ICML , volume=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.166341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:32019e2df83ce0dd2db508ed4ea04ddf450814ca570aa9b4b3406de8075313c9

Observation 6c1c20d2-4eb6-42f2-9452-72f79511353d · outbound

This paper cites International conference on machine learning , pages=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data International conference on machine learning , pages=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.168441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:a5e60f2a8e35ea8639374c9c5ee5249157f05457e0d6434a75c19e251e239a8f

Observation 63ddf2bd-59e8-48f4-a51e-3321c0340c5d · outbound

This paper cites Nature , volume=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Nature , volume=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.158224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:50e08a5719dc22948fbb9db03942f79e058f556683a74a1b768204c484b3e765

Observation 095e9a13-56fe-4811-9a1c-7278e980bc9c · outbound

This paper cites 1992 , publisher=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 1992 , publisher=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.133855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:89de551aa594145ff3b2dcc218bbf4d698b12edee052594dad08183e81695e29

Observation 91a4765b-38ad-4bf6-b8e1-28fb4649bdc2 · outbound

This paper cites 2015 , publisher=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2015 , publisher=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.146122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:557c1d70fd5d30f499682b12d3eda88ed5ff1904cda0f48ae96ba680286b70bf

Observation eec50442-9be8-410a-9192-680113fb5c04 · outbound

This paper cites Evolutionary computation , volume=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Evolutionary computation , volume=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.162083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:f60fc8312f09bbcb7824cb616493eeb6c16132f276505c05d565ce32a4b9be9f

Observation 3150ed39-a429-4888-822e-2c036bfbba5e · outbound

This paper cites Frontiers in Robotics and AI , volume=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Frontiers in Robotics and AI , volume=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.170575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:7ed6546dbf2c2c422e089d4f0fcb57a31a357b0e7786d9db0f52cdf9d764da1c

Observation d85fd09c-20f7-485f-9cac-9bb1d58d1c65 · outbound

This paper cites Jointly Reinforcing Diversity and Quality in Language Model Generations.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.947611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:159dfaffcab74e3eb9233e655df2a38a5bc8f184ef57ba59330b9ea763eb2b66

Observation 712dbd10-e8fa-4e8c-b391-6cdde66035f1 · outbound

This paper cites Modifying Large Language Model Post-Training for Diverse Creative Writing.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Modifying Large Language Model Post-Training for Diverse Creative Writing

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.965924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:f4cc321f07cb5c2845ec005df6df99f311acf15e898bc8c0d0b4ae140f69524a

Observation 039526f7-c65a-432c-a4db-ffbfe5e0f3b3 · outbound

This paper cites Supervising the search process produces reliable and generalizable information-seeking agents.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Supervising the search process produces reliable and generalizable information-seeking agents

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:02:48.394837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:d67bdb84f2868eabb0cc08987b096040f19555f0b0ec7922bb9316d78c7e23a0

Observation d41f4700-3287-4645-9187-33226ff86958 · outbound

This paper cites R1-RE: Cross-Domain Relation Extraction with RLVR.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data R1-RE: Cross-Domain Relation Extraction with RLVR

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.990762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:a5e4e4b2c95c1df72652c5cfd2e91476969a545bc3191f3c45c609b1959c5862

Observation c50ad9dd-2bad-4881-a5e8-43369204a2e7 · outbound

This paper cites CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.999728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:31341a3b0a4574d608cbf329c50a4b57243b1337eb289cb1cda7e11aef44c484

Observation 4608f550-acd6-40c2-ad73-24abc60bcef4 · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T05:36:02.006979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:35166993c47edfb1aa39152af9e3bce85b2be1faf30beadde4d87c2f5cbd9e53

Observation 4bf022af-6523-4e8d-b7b2-2929d433f2d5 · outbound

This paper cites Parallel-R1: Towards Parallel Thinking via Reinforcement Learning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.009207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:6ed67d2bf79a4757684644a779595e95fc318fe40fa59b4b30409e9934b0c463

Observation 0485d328-d6f1-41f6-99ea-6222fc70d86b · outbound

This paper cites Learning to Reason via Mixture-of-Thought for Logical Reasoning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Learning to Reason via Mixture-of-Thought for Logical Reasoning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.997503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:bf6b4a529b77810d13b374db253e40493bd79bc562eb5739c31f73e4a11bf60e

Observation f95d8c1f-8773-4da6-b3a5-dbcc2a60c164 · outbound

This paper cites arXiv preprint arXiv:2505.17312 , year=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data arXiv preprint arXiv:2505.17312 , year=

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.995170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:5673f53e1ecfdcdea9747add873496327f97ca23e75e95ecfd5e0076f0da2a43

Observation 071fa43e-982e-4eb2-b78f-fd27e1edc040 · outbound

This paper cites In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 3: System Demonstra- tions), Bangkok, Thailand.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 3: System Demonstra- tions), Bangkok, Thailand

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.992928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:650b2c370fada20dcf377a280982a1ef079fd7fc9fe9bb9c7da79bf22e05b3f4

Observation d5e49c82-93f5-4adb-b388-7b1d2a99e7c5 · outbound

This paper cites Defending Jailbreak Prompts via In-Context Adversarial Game.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Defending Jailbreak Prompts via In-Context Adversarial Game

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:02.001919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:22cfa928fbf60c220085608ca2a3fbcc8180093580f902982e512739c9d5d4bb

Observation 698142b3-acf4-4bd0-8187-5ea823064abd · outbound

This paper cites 2023 , publisher =.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2023 , publisher =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.135657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:f293368f4c39a555e31dec193d566cbc94ac2cad70ed26790db2e214c5aa9012

Observation 8e36973e-67df-41e7-9a97-302f3a39317d · outbound

This paper cites 2024 , eprint=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2024 , eprint=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.144077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:f37fe588a80ea965d96e4c0857d5672ee7a194bd3174521742ec57948d2a3982

Observation 678442dc-6f4e-4b81-9c71-cf18cc9545e6 · outbound

This paper cites 2024 , eprint=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2024 , eprint=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.142084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:2f8156cd4976da528eefb50be64522b52e73a2cce50cbdd1ef24dd8518e73156

Observation 82b04643-3a14-492f-813f-e86bb6a40b19 · outbound

This paper cites Forty-first International Conference on Machine Learning , year=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Forty-first International Conference on Machine Learning , year=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.152295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:b13a6071bdf7f91eef8b7a54b6f23d6fbbb213f3f4dbf79c44404954e59aa6b3

Observation 45f50f34-9bc3-4bc8-bc40-73d74df433d9 · outbound

This paper cites Aligning Large Language Models by On-Policy Self-Judgment.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Aligning Large Language Models by On-Policy Self-Judgment

Reference 54

Resolution
verified exact
doi, observed 2026-05-10T05:36:01.404182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:e218f32de6c88decb64b45fb6a258e3bad599b4ff8726dc5fba00ded81086e3a

Observation b657e844-1f88-4e90-bfe9-b4ed64598212 · outbound

This paper cites an unresolved cited work.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-21T20:14:21.154192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:beb2e27a2ff4c10dfbb9416b0feb028ebfd566cb3bce2fa882ec485c3e4c5b38

Observation d77c955d-4b00-40ee-8a6e-d31116646c74 · outbound

This paper cites 2025 , eprint=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2025 , eprint=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.164119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:4e3db1cf8d54681549943ac7e7f07479b110111833e0e1a28e272bef657a2ffd

Observation f71a7038-d81c-4929-b5a1-7fed1804b6cf · outbound

This paper cites 2024 , url=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2024 , url=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.160158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:15ce2f4171b5d0c3b98beadbb9ecadc9c504b4fc5a124e4432e4f411fb457b11

Observation 2c8918ec-5cca-43ce-b873-f66263a5a20c · outbound

This paper cites 2025 , eprint=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2025 , eprint=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.131627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:49ba0d4d6b20306d6d9f140af609001f062d2df18bbaedeb0bb430728f3b7d37

Observation 3d562eac-0960-4867-9192-5a117513b837 · outbound

This paper cites 2025 , eprint=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2025 , eprint=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:14:21.148048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:b0d35ac145366aa445d059b6e951ed51d98f6091c80326df5f5e74dddc51af1a

Observation 3837c19e-f352-4972-aee6-56285219ce18 · outbound

This paper cites Serl: Self-play reinforcement learning for large language models with limited data.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Serl: Self-play reinforcement learning for large language models with limited data

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.977673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:e4a021a859a5ed201797b1abdcf3c47929361410c4bf6b9874a315259ead6fc0

Observation 24083e58-3a7d-4795-ad66-57b4b709d172 · outbound

This paper cites Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.961406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:d70e9a92ed83ac317dcf03d88d9870242d5b5196be4a5255645ac1b157aa6db3

Observation 68a635f4-60e6-45b8-9935-fb20d3e069d7 · outbound

This paper cites arXiv preprint arXiv:2509.15194 , year =.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data arXiv preprint arXiv:2509.15194 , year =

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.949886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:ce3b9da303527c772e12f081cea4cb44c86ac19dad2ebdbe1e1b2ba0bb99eb50

Observation 07669db1-072f-490d-bc54-986566efa132 · outbound

This paper cites arXiv preprint arXiv:2509.23095 , year=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data arXiv preprint arXiv:2509.23095 , year=

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.988613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:3f4261932e2cf7e2c9d8c320011b99365fe0c05d6fafac2ee49fbb67aacc129f

Observation 5162493c-3ae6-416d-8e69-bcf803ca90cc · outbound

This paper cites Clue: Non-parametric verification from experience via hidden-state clustering.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Clue: Non-parametric verification from experience via hidden-state clustering

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.972859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:34af0dc43a23187035cdb623388f7652ed6df4d71717d1f14f58c118e1aa9ac1

Observation 7d2189a3-1cd6-40b3-834c-fc70876c7e91 · outbound

This paper cites Can llms guide their own exploration? gradient-guided reinforcement learning for llm reasoning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Can llms guide their own exploration? gradient-guided reinforcement learning for llm reasoning

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.986460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:87fc21f417a2181fca5769a34c4ff809f8337da3a77cc7d054009c4b13d101f0

Observation ac91009d-f39c-4271-9c98-335f9dd7d980 · outbound

This paper cites Exploring multi-temperature strategies for token-and rollout-level control in rlvr.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Exploring multi-temperature strategies for token-and rollout-level control in rlvr

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.970575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:42c68fa680434cbcf16c0f1f825e2dcdff51a5aeef6ae5f6b14d9c20bab84aab

Observation de6189ef-5623-4637-a986-bdee79e7b202 · outbound

This paper cites V ogue: Guiding exploration with visual uncertainty improves multimodal reasoning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data V ogue: Guiding exploration with visual uncertainty improves multimodal reasoning

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.942930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:f43fc2d08e0770047fa6b172e563e178ade3f4691d7de417b5ca5baf43dff974

Observation 7b8365cc-580e-4ddd-9fb8-b09be38c1ab5 · outbound

This paper cites Visplay: Self-evolving vision-language models from images.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Visplay: Self-evolving vision-language models from images

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.935863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:0001dd9bf58c77564910564350d640eafd1d01bbb855ed29fe57fb1b15d53df8

Observation edfc3ae5-4955-4c93-bb1c-5879db03753e · outbound

This paper cites arXiv preprint arXiv:2510.02172 , year=.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data arXiv preprint arXiv:2510.02172 , year=

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.952217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:38787831d66fb99534b54d79a7443b1d066267f1208cbeeefb1c2180ca67af39

Observation aa87b846-4536-486e-aa05-534c6f69aa0a · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:05:26.778195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:138ccefe810b46be66e422097040c14247d9a3675292394a71d9cae184acc167

Observation 29bfc979-f3e0-4a74-bd01-18b712687633 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Proximal Policy Optimization Algorithms

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T05:36:01.940490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:6b3a92c8a3810fa7c2e1dede82da6f4e52bdff44d34e91c3c6748d252c97fd36

Observation 2c94ed64-a235-4ca7-afe5-59f43d3d5e7d · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.984213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:a45a7333d687984386724eb3adea4906ab9143cd226b867a09ad147ef8d4cceb

Observation 0d3b114b-7072-4f2c-83fe-8eb9bab0386a · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:0b4fd1ffe94502a461e1d6ea112e4823a42c3ba6dd85b5395833b73cb3b91997

Observation 91e51ab2-e4e0-4a61-b9fc-5863ab8ade4b · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.954519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:8117f61126fc08425933f4b190c97aceac0b3259108de4e2e48ff828e005a3ba

Observation 98630689-e10f-4b36-b5a8-5e003b457348 · outbound

This paper cites Save the good prefix: Precise error penalization via process-supervised rl to enhance llm reasoning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Save the good prefix: Precise error penalization via process-supervised rl to enhance llm reasoning

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.956824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:f4d108683ee2d6f544e5a50931faf6c8635a27623220b72ea1b93874ddde8dd2

Observation ae620dac-a917-4846-9c49-b721a7444220 · outbound

This paper cites Alignment Risks from Capability-Seeking RL Training.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Alignment Risks from Capability-Seeking RL Training

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-06-05T02:16:23.300865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:be9c2ae3ca8a2b5dc4bc7842a3e119b8cc041e21484107fdd5347ac9df6eb56a

Observation 64d68eea-30f4-4754-b8df-bf4de29647c5 · outbound

This paper cites Stable and efficient single-rollout rl for multimodal reasoning.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Stable and efficient single-rollout rl for multimodal reasoning

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.982078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:957d5beb1d595286ab99d388ba3d9314ad785d7963bc39c5daf51e2e81f93340

Pith citing papers

Observation 50715b1b-3c85-4a8e-a212-7f03ec207950 · inbound

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents cites this paper.

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:31.392889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T07:04:51.970049Z digest=sha256:419d354a48f63f91f456f0b2c55b1102cb14e12b047ac490e343f2f6212e58e7