Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T22:45:56.857810Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2605.11461.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T22:45:56.857810Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T08:51:22.327336Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T10:29:44.758886Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 70d21982-717d-46cb-8a73-a099bac0bbdd · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1ee6409a-599f-4fbf-beeb-625ef87d80d0 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8350d31e-0e02-4dae-98d5-3179fb64b761 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Post-training large language models for diverse high-quality responses
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4efd0409-3d61-4fff-90f7-0358f16c0b19 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a070d93c-9384-4c88-afa0-96c44f6c1948 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e955fdf2-5dc6-4b94-bc6d-eb0fd198b68c · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 03444957-c485-4a6e-b590-c17ad6fdeabe · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ea2ef30-4495-4cbf-bc27-c79752a0aa7d · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2cc5077a-38db-4781-ad42-8594a5809f32 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Scaling laws for reward model overoptimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e8eaa5d6-ed76-4b13-a4d6-2d5f0fd87b2b · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eb03bcd3-6957-47a1-a2e4-64716fcd618c · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b0e72bd3-cb95-465d-8f94-3ec23d0f8b01 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation af82c7b1-a355-4249-8089-deb2cf0fbeea · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Diversity-incentivized exploration for versatile reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 30ab1d08-2b81-4c5b-8294-fb3a40483662 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 443734eb-332c-4926-96c8-8add893700f7 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Risk-sensitive rl for alleviating exploration dilemmas in large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 00aa646e-b135-450b-b338-c058bb2559dc · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Determinantal point processes for machine learning.Foundations and Trends® in Machine Learning, 5(2-3):123–286
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71d12950-3c14-472b-a866-16637fcb7633 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Solving quan- titative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 72698e9c-aeeb-4466-9574-642f650d7f34 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Jointly Reinforcing Diversity and Quality in Language Model Generations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ef0cc360-6e4b-4d05-afa4-6b06fb9bea5e · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning DeepSeek-V3 Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c66f6f6f-a551-420e-9bfe-5397881f3863 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 844b31fa-83ab-44a1-8b2e-f7d5ebd5d7fd · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d6b2dc71-89b8-4fdb-9d1e-97c4e4b026dd · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c7075981-c91a-4a0f-a810-57b64ee9d761 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Learning to reason with llms
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 755233fe-efc9-4947-b4c9-10284346a5ad · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Sentence-bert: Sentence embeddings using siamese bert- networks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e2c65282-b7f7-4495-8dd5-93ee85ad2b4f · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a03971a1-5704-4292-8086-ce84dda4f711 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Proximal Policy Optimization Algorithms
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cab2ade1-93c5-4c8d-a98f-7e4710d69843 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8955f1d1-a9a0-40fa-9d01-56b61553ce5c · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd5f24e5-6eff-4771-8935-3a34573bd62b · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Hybridflow: A flexible and efficient rlhf framework
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fcaecc92-9650-4fb4-8d0e-7b3ab5a8e1bd · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning OpenAI GPT-5 System Card
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 466fcc94-82a5-4593-b724-04ae1d33bf5e · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning The many shapley values for model explanation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eb66909f-1c90-42cf-a969-d4655d265920 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5a959122-a490-40ed-9625-6df7f2c79558 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Text Embeddings by Weakly-Supervised Contrastive Pre-training
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2df71e25-67fe-44d9-ba7f-6b8c2a16ca6c · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9d86e4cd-3a7b-438f-b782-2247e5dbad5d · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning The invisible leash: Why rlvr may or may not escape its origin
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d470ffd1-8c8a-4988-b95b-259261b16831 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Progress or Regress? Self-Improvement Reversal in Post-training
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e8d4f60f-6755-490f-b372-f6c88097eba1 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning C- pack: Packed resources for general chinese embeddings
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 12d63714-7461-4126-b065-582ce3ea7b9e · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Qwen3 Technical Report
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 54bd6d96-2304-43ac-9860-3709fcd52c96 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Diversity-aware policy optimization for large language model reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5585099f-ce1f-4d6e-91bc-1b055664ba2d · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bcb43b4b-7ee9-4305-bc98-03103ae8f655 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0de2b83e-5274-4270-9f29-31a86aa97994 · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning k_ i=1 (ri = 1) # =E
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1b255a29-cb82-4e9d-b49b-353f906279cc · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5b0f3a68-188b-435f-a50d-851ac05918bf · outbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning double-counted
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e7956f53-6352-4c8a-8d0e-ad631586c4cd · inbound
Exact Schur-Sylvester Dimensionality Reductions for Non-Smooth Stochastic Complexity and Manifold Sampling Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.