Pith. sign in

Paper Citation Record · LEDGER

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning

As of 5 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2605.11461.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.11461 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T22:45:56.857810Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T08:51:22.327336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T10:29:44.758886Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact30
  • verified fuzzy13
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70d21982-717d-46cb-8a73-a099bac0bbdd · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.670654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:1799c9a392c0dbc1d40a1e9bb3b170f0c2736383d5a13ba2c6e2c6ff75dc4221

Observation 1ee6409a-599f-4fbf-beeb-625ef87d80d0 · outbound

This paper cites EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.533825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:bc826b592c9fe60f406e3503856e3564695e1398c1a0c720f65574780bd47c10

Observation 8350d31e-0e02-4dae-98d5-3179fb64b761 · outbound

This paper cites Post-training large language models for diverse high-quality responses.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Post-training large language models for diverse high-quality responses

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.541535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:410aa7a83a7d0bd228f4ad38b143213b97816ce8b834865ab3b64411c3d21aa2

Observation 4efd0409-3d61-4fff-90f7-0358f16c0b19 · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.563663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:d6e2458cb74294d9b5e7c6f5376a6ff62f14ac720ba10756c7233b189adac7f4

Observation a070d93c-9384-4c88-afa0-96c44f6c1948 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.692002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:796013dfcd612a27e1ffbb1d63fad0125959ad68b40f4be0780e0928826edb74

Observation e955fdf2-5dc6-4b94-bc6d-eb0fd198b68c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.616959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:e4e18e16e1ce91f8128041a77fe2c09711242fd0b1a7008afa811876d25c82d7

Observation 03444957-c485-4a6e-b590-c17ad6fdeabe · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.637728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:a954c1659368b60b4d770e5b987aecc754879585de78c2647d50549b2a51b40f

Observation 9ea2ef30-4495-4cbf-bc27-c79752a0aa7d · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.590794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:2c2d04ee8c9009afcb5f71fd0c8dc3bad2b3c28f487bf0e99a86a1bc5d008535

Observation 2cc5077a-38db-4781-ad42-8594a5809f32 · outbound

This paper cites Scaling laws for reward model overoptimization.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Scaling laws for reward model overoptimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.928668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:1c6c4929fc79ad8299d5309ae79b8000b64ef893454e17f4281a52cd62aec955

Observation e8eaa5d6-ed76-4b13-a4d6-2d5f0fd87b2b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.577261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:f60354716f60dfb2e77ed123efa2b89c60cf44f353787e27340b0ec569f87721

Observation eb03bcd3-6957-47a1-a2e4-64716fcd618c · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.677272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:0ad73cf44920beb1a1da176a4ea7b736483485704cf97aaf93f678c50e8fea1e

Observation b0e72bd3-cb95-465d-8f94-3ec23d0f8b01 · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.497847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:01db2dd49afef96099c4b87a42ecae7f22362b5e89855c6fcf51ca59b3cdcbeb

Observation af82c7b1-a355-4249-8089-deb2cf0fbeea · outbound

This paper cites Diversity-incentivized exploration for versatile reasoning.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Diversity-incentivized exploration for versatile reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.518109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:860371acf53db40d430681b742b8906b4366fe5633cbe935442b26a15fca9013

Observation 30ab1d08-2b81-4c5b-8294-fb3a40483662 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.961511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:ee80d9b54f686b64841299e018e42453cc10b8767c42263b9059eaf5b12218b7

Observation 443734eb-332c-4926-96c8-8add893700f7 · outbound

This paper cites Risk-sensitive rl for alleviating exploration dilemmas in large language models.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Risk-sensitive rl for alleviating exploration dilemmas in large language models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.504480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:bfb54bd0bd18e606f4f2f16530aa36c214347859a856cd7dca14e6ad8f2b7b15

Observation 00aa646e-b135-450b-b338-c058bb2559dc · outbound

This paper cites Determinantal point processes for machine learning.Foundations and Trends® in Machine Learning, 5(2-3):123–286.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Determinantal point processes for machine learning.Foundations and Trends® in Machine Learning, 5(2-3):123–286

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.964645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:69d4e20dbdcc6ea1cc2eec2a6f1b4334d3fa179bd43a5b02d0918b75533b18be

Observation 71d12950-3c14-472b-a866-16637fcb7633 · outbound

This paper cites Solving quan- titative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Solving quan- titative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.967728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:494c8ffc2158b47a5e1e261ada457eb816217ba75be1c59a5e304fcb3fc4df01

Observation 72698e9c-aeeb-4466-9574-642f650d7f34 · outbound

This paper cites Jointly Reinforcing Diversity and Quality in Language Model Generations.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.604622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:6a211808ee306c61cb168dabfbe1fd138498970985509f569f1cebec36916491

Observation ef0cc360-6e4b-4d05-afa4-6b06fb9bea5e · outbound

This paper cites DeepSeek-V3 Technical Report.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning DeepSeek-V3 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.697979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:84a3b6490c2586f559ca1454a684abb5dd00a0bfe1855e18a60108f6a8c99f6e

Observation c66f6f6f-a551-420e-9bfe-5397881f3863 · outbound

This paper cites ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.526067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:03a665b43981f0321c011f57c141ba1a38e19dbccced369237e4431e13087184

Observation 844b31fa-83ab-44a1-8b2e-f7d5ebd5d7fd · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.664024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:08e4a54831958ab0592f2c749e2f337b4744d93256fccd228e1d0c852c0baf10

Observation d6b2dc71-89b8-4fdb-9d1e-97c4e4b026dd · outbound

This paper cites Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.954614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:48f57eaf8ec9a7f83e825e0268b530a604eb51f913f94a8df33d681b95bc8791

Observation c7075981-c91a-4a0f-a810-57b64ee9d761 · outbound

This paper cites Learning to reason with llms.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Learning to reason with llms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.957722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:7ae79a5b5d4dc6da13eaadb869bb8cfe6d2664d5dc6e75c4c105e9c3a517c44c

Observation 755233fe-efc9-4947-b4c9-10284346a5ad · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert- networks.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Sentence-bert: Sentence embeddings using siamese bert- networks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.947952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:3055e5ccfcd1ab61568c2c95f603c676d85445f22d0c43ddbae654003a0516e3

Observation e2c65282-b7f7-4495-8dd5-93ee85ad2b4f · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.584090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:4e906c0569707ae199064f7b4fbdf5188a6ccf4bdafd1586a59fdc69a4d7acb0

Observation a03971a1-5704-4292-8086-ce84dda4f711 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Proximal Policy Optimization Algorithms

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.657573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:2ad23fa93335a1bf88b45ca00ee8dd506585d682a9f5a481d76b774b379b3c02

Observation cab2ade1-93c5-4c8d-a98f-7e4710d69843 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.651666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:d6af4203fe9cf848b40ebf17a8402570bfc429e2b40a3b5d62bf278e8ec22ee8

Observation 8955f1d1-a9a0-40fa-9d01-56b61553ce5c · outbound

This paper cites On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.597759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:bc8d96e88ef32188b6fd4202f9ff6e007884396b05062bbab5a55cca43cc297c

Observation fd5f24e5-6eff-4771-8935-3a34573bd62b · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Hybridflow: A flexible and efficient rlhf framework

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.944851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:33768028f5c73c2a7183e41f19a2731d3499c215759bcc4caeae6a56eb6a6577

Observation fcaecc92-9650-4fb4-8d0e-7b3ab5a8e1bd · outbound

This paper cites OpenAI GPT-5 System Card.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning OpenAI GPT-5 System Card

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.610581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:8ae245e26c5a483dc0729ff0f40ac2c9f3872f5d7a45cb6d8e76ce99e03d3f49

Observation 466fcc94-82a5-4593-b724-04ae1d33bf5e · outbound

This paper cites The many shapley values for model explanation.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning The many shapley values for model explanation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.951199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:f53f3a41c9897f22c93beac84b31b2f9baac9b991cd98062233a132390525f70

Observation eb66909f-1c90-42cf-a969-d4655d265920 · outbound

This paper cites Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-11T02:08:32.357336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:5e612b82f409e66e9ea9a34718de11ba080b270fd4d5ef0e88d11566e6ed387d

Observation 5a959122-a490-40ed-9625-6df7f2c79558 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.554561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:a8abca8f6caf379a3a55a9979a4b687d521a62c3b20f0b3aa51af5acf3d10fe5

Observation 2df71e25-67fe-44d9-ba7f-6b8c2a16ca6c · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.942241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:16d07fb2c0f73c2508e6fd6ef2289487b3b9b5f0ca48c0b8ae3568ecade38dae

Observation 9d86e4cd-3a7b-438f-b782-2247e5dbad5d · outbound

This paper cites The invisible leash: Why rlvr may or may not escape its origin.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning The invisible leash: Why rlvr may or may not escape its origin

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.548282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:fba07c5bf5f4df536c2234e3954f08d60e9577bca6db5f68ca1d737600dd2836

Observation d470ffd1-8c8a-4988-b95b-259261b16831 · outbound

This paper cites Progress or Regress? Self-Improvement Reversal in Post-training.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Progress or Regress? Self-Improvement Reversal in Post-training

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.644947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:2637ac2c7289d717d5ebdac4056468af251ea1a15973a98a86587fc450646d72

Observation e8d4f60f-6755-490f-b372-f6c88097eba1 · outbound

This paper cites C- pack: Packed resources for general chinese embeddings.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning C- pack: Packed resources for general chinese embeddings

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.939099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:e52289f1ace259dadc650ad291b72123f2f0e6543162aa56e2128784d19eff05

Observation 12d63714-7461-4126-b065-582ce3ea7b9e · outbound

This paper cites Qwen3 Technical Report.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Qwen3 Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.624071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:3671284aad06a3bc73480c220f61be752f13de98e93df812cb6e1399df61b27a

Observation 54bd6d96-2304-43ac-9860-3709fcd52c96 · outbound

This paper cites Diversity-aware policy optimization for large language model reasoning.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Diversity-aware policy optimization for large language model reasoning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.631105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:8a3f073d57273e17595abdedd2f6bb6fd4c4590479b68f4acee164a8be4590e3

Observation 5585099f-ce1f-4d6e-91bc-1b055664ba2d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.683745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:774c93eafd61b001fc2830d507c0e1cb88a1f6bdeb8fc36307b5bc2edf2065d1

Observation bcb43b4b-7ee9-4305-bc98-03103ae8f655 · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.570853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:7ddaced12f42165de9d5c74c0f8e15f3d244d14a6c3fce617e9b313c7f23c81f

Observation 0de2b83e-5274-4270-9f29-31a86aa97994 · outbound

This paper cites k_ i=1 (ri = 1) # =E.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning k_ i=1 (ri = 1) # =E

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.936726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:b1322ae891ebfb19ab4e268880087c4896d6cfcbb2506d1019b30f1817e01c11

Observation 1b255a29-cb82-4e9d-b49b-353f906279cc · outbound

This paper cites an unresolved cited work.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-20T22:49:11.931541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:bac9c09aa8d10690f6e2bf8e2963beaf67a45ecc4cf8c507f08dc2df5b507fe2

Observation 5b0f3a68-188b-435f-a50d-851ac05918bf · outbound

This paper cites double-counted.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning double-counted

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:49:11.934089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:ecc160b72ef262f19d361c15043db004c8ef353def7fb06f97a7550f88a58c7e

Pith citing papers

Observation e7956f53-6352-4c8a-8d0e-ad631586c4cd · inbound

Exact Schur-Sylvester Dimensionality Reductions for Non-Smooth Stochastic Complexity and Manifold Sampling cites this paper.

Exact Schur-Sylvester Dimensionality Reductions for Non-Smooth Stochastic Complexity and Manifold Sampling Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.760416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T08:51:22.327336Z digest=sha256:9b504a37186b697819f92717047cf7be82235d0f9fe474ea0c458352ad10224d