Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T05:32:23.972335Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 1 inbound Pith citation observation for arXiv:2604.18493.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T05:32:23.972335Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T07:04:51.970049Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T14:28:31.391729Z
77 of 77 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3d982aff-524c-4a4b-99f9-06491df24d12 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data TTRL: Test-Time Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f7f0a398-e21b-4adc-858e-2091dc4b56c2 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 43f1f110-ab36-4a96-bb50-dd20d000d334 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a77039c5-391b-46f0-961b-c9bd3e6d4e0a · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95ea86d6-6133-45e6-9987-8b5466e5ed5a · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data One Token to Fool LLM-as-a-Judge
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fdd4da0d-4987-4e9a-b13e-ced8e7368487 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data OpenAI o1 System Card
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47a0b006-db2a-4212-8b40-51b8c559e048 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6b320650-b316-425c-9ddd-a63262182404 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Qwen3 Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5b18a2b9-e4fb-4e0e-ba7f-197c79a03bbb · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d67d51a4-c9e1-4ea0-94f3-2bcf7f5312e0 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 52f3cb1b-099f-4c53-9e14-867cb07c84b5 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Reinforcing General Reasoning without Verifiers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b8b20a9b-7f99-42a0-9d14-1775e1bc2563 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c5a719de-c496-4442-952c-34eb2b8251e0 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Maximizing Confidence Alone Improves Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 21d385bb-39f7-49c0-b23a-75d66eacf57b · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98aa62d3-2ba4-46c1-a96c-2b8cb987701b · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Learning to Reason without External Rewards
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b03f0231-88d5-41b8-972d-ca08f6ba4319 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9d56d6ae-3ad8-4349-a497-dbd30086eec0 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0ebb5673-a95c-46c8-b46b-c61a588a9abb · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data s1: Simple test-time scaling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 696efccc-7a89-405e-80f4-a3c7a869dc1c · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ea96b77-9e51-4be1-b951-b8c924bb42b2 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ae0e32a-3386-4927-a0a5-ec71e8f0c24c · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e6d2d1b-a4e5-46e4-9faf-25570ad0a42e · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data s1: Simple test-time scaling
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 97b7fe68-7dfb-4165-bc4e-0cade0795304 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f9113143-e73d-4457-9c01-23d025e5bb2b · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Can large reasoning models self-train?
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 89cda544-1c8e-48f0-a479-3f16ac20cc81 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data First Conference on Language Modeling , year=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4831cb6f-d040-423c-a5b9-d941ebdeb8cf · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Hugging Face repository , volume=
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98b3b3c6-30d8-4225-bfa1-75cf3791759a · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Measuring Mathematical Problem Solving With the MATH Dataset
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 55898f63-c9ab-42a1-9e93-f14d32339cfc · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 26110588-8257-45f7-bbd3-bde092241b05 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data International conference on machine learning , pages=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c97915af-e7eb-4acc-b148-7d9ad74b52a5 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Tent: Fully Test-time Adaptation by Entropy Minimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 534b6c4f-b157-46aa-ba63-5fe0fa65c371 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Advances in neural information processing systems , volume=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7ac4f40b-16d9-4972-8feb-a9c4b7a41e7a · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Workshop on challenges in representation learning, ICML , volume=
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6c1c20d2-4eb6-42f2-9452-72f79511353d · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data International conference on machine learning , pages=
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 63ddf2bd-59e8-48f4-a51e-3321c0340c5d · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Nature , volume=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 095e9a13-56fe-4811-9a1c-7278e980bc9c · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 1992 , publisher=
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 91a4765b-38ad-4bf6-b8e1-28fb4649bdc2 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2015 , publisher=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eec50442-9be8-410a-9192-680113fb5c04 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Evolutionary computation , volume=
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3150ed39-a429-4888-822e-2c036bfbba5e · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Frontiers in Robotics and AI , volume=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d85fd09c-20f7-485f-9cac-9bb1d58d1c65 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Jointly Reinforcing Diversity and Quality in Language Model Generations
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 712dbd10-e8fa-4e8c-b391-6cdde66035f1 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Modifying Large Language Model Post-Training for Diverse Creative Writing
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 039526f7-c65a-432c-a4db-ffbfe5e0f3b3 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Supervising the search process produces reliable and generalizable information-seeking agents
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d41f4700-3287-4645-9187-33226ff86958 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data R1-RE: Cross-Domain Relation Extraction with RLVR
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c50ad9dd-2bad-4881-a5e8-43369204a2e7 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4608f550-acd6-40c2-ad73-24abc60bcef4 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4bf022af-6523-4e8d-b7b2-2929d433f2d5 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0485d328-d6f1-41f6-99ea-6222fc70d86b · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Learning to Reason via Mixture-of-Thought for Logical Reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f95d8c1f-8773-4da6-b3a5-dbcc2a60c164 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data arXiv preprint arXiv:2505.17312 , year=
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 071fa43e-982e-4eb2-b78f-fd27e1edc040 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 3: System Demonstra- tions), Bangkok, Thailand
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5e49c82-93f5-4adb-b388-7b1d2a99e7c5 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Defending Jailbreak Prompts via In-Context Adversarial Game
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 698142b3-acf4-4bd0-8187-5ea823064abd · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2023 , publisher =
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8e36973e-67df-41e7-9a97-302f3a39317d · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2024 , eprint=
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 678442dc-6f4e-4b81-9c71-cf18cc9545e6 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2024 , eprint=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 82b04643-3a14-492f-813f-e86bb6a40b19 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Forty-first International Conference on Machine Learning , year=
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 45f50f34-9bc3-4bc8-bc40-73d74df433d9 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Aligning Large Language Models by On-Policy Self-Judgment
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b657e844-1f88-4e90-bfe9-b4ed64598212 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d77c955d-4b00-40ee-8a6e-d31116646c74 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2025 , eprint=
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f71a7038-d81c-4929-b5a1-7fed1804b6cf · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2024 , url=
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c8918ec-5cca-43ce-b873-f66263a5a20c · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2025 , eprint=
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3d562eac-0960-4867-9192-5a117513b837 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data 2025 , eprint=
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3837c19e-f352-4972-aee6-56285219ce18 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Serl: Self-play reinforcement learning for large language models with limited data
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 24083e58-3a7d-4795-ad66-57b4b709d172 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 68a635f4-60e6-45b8-9935-fb20d3e069d7 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data arXiv preprint arXiv:2509.15194 , year =
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 07669db1-072f-490d-bc54-986566efa132 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data arXiv preprint arXiv:2509.23095 , year=
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5162493c-3ae6-416d-8e69-bcf803ca90cc · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Clue: Non-parametric verification from experience via hidden-state clustering
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d2189a3-1cd6-40b3-834c-fc70876c7e91 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Can llms guide their own exploration? gradient-guided reinforcement learning for llm reasoning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ac91009d-f39c-4271-9c98-335f9dd7d980 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Exploring multi-temperature strategies for token-and rollout-level control in rlvr
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation de6189ef-5623-4637-a986-bdee79e7b202 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data V ogue: Guiding exploration with visual uncertainty improves multimodal reasoning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b8365cc-580e-4ddd-9fb8-b09be38c1ab5 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Visplay: Self-evolving vision-language models from images
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation edfc3ae5-4955-4c93-bb1c-5879db03753e · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data arXiv preprint arXiv:2510.02172 , year=
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aa87b846-4536-486e-aa05-534c6f69aa0a · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 29bfc979-f3e0-4a74-bd01-18b712687633 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Proximal Policy Optimization Algorithms
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c94ed64-a235-4ca7-afe5-59f43d3d5e7d · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d3b114b-7072-4f2c-83fe-8eb9bab0386a · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 91e51ab2-e4e0-4a61-b9fc-5863ab8ade4b · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98630689-e10f-4b36-b5a8-5e003b457348 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Save the good prefix: Precise error penalization via process-supervised rl to enhance llm reasoning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae620dac-a917-4846-9c49-b721a7444220 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Alignment Risks from Capability-Seeking RL Training
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 64d68eea-30f4-4754-b8df-bf4de29647c5 · outbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Stable and efficient single-rollout rl for multimodal reasoning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 50715b1b-3c85-4a8e-a212-7f03ec207950 · inbound
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.