Pith. sign in

Paper Citation Record · LEDGER

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2605.00365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.00365 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T19:56:39.465133Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:54:44.345264Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T21:54:44.621660Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact9
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch10

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c63f10a-6c93-4e41-b446-d0fd7b3d0cb2 · outbound

This paper cites Nature , year=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Nature , year=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.336237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:3d51d59e443308993e1874eb9a0ec94a8cb5d6f2a9aaabf2022b3186619bf38b

Observation ed441f2d-4a7c-4f56-8876-1442f660c6dc · outbound

This paper cites 2024 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2024 , eprint=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.344229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:d5eee6a988bba6049dbdd64c245ef4274230b04458de3a94a00da3d028544657

Observation c80f782c-49df-404c-ac90-0055ef923aae · outbound

This paper cites arXiv preprint arXiv:2601.15609 , year=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity arXiv preprint arXiv:2601.15609 , year=

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:16.745333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:e099235a666bb168acc647ea20403e38d62197a22f91380efd0d858571bd8b68

Observation f08b654c-3f9a-47fa-a0a7-ad00e088a2ed · outbound

This paper cites Advances in neural information processing systems , volume=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Advances in neural information processing systems , volume=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.340306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:93dfbdb2f07dd09aa7605305a029aaf000736ad544c8c9f93fdc158afda646a5

Observation c695e834-270d-4a59-82a6-38dbfe8e6e63 · outbound

This paper cites 2026 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2026 , eprint=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.332258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:b330c7f471cf57ecce41823c3ebcc7c2a37fa168ed9a3d493185460722bc1cd6

Observation 537ecb71-4d46-4791-93e6-b26b58eb06c8 · outbound

This paper cites Agentic Reinforced Policy Optimization.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Agentic Reinforced Policy Optimization

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:57:12.173630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:7938e41dd4f7f4178bcd759704ba9c70bb37e27aca195ff01d92f1fcda9ca672

Observation a14abe58-6121-4983-9cbf-3aae2077b979 · outbound

This paper cites Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening

Reference 7

Resolution
verified exact
doi, observed 2026-05-09T20:01:36.736368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:a6c0ad2af9149f0eea91be5b823a3628376e84c63bc70f8a0f491387528c952f

Observation e9ddb812-b7d7-4af8-838d-3c78e5830b14 · outbound

This paper cites arXiv preprint arXiv:2509.25133 , year=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity arXiv preprint arXiv:2509.25133 , year=

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:16.537353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:3cff4bfd3271ca6da0d53a1a1472eae24db71693141033f25fde314ad9a72aa0

Observation d9c90e26-a085-44e3-9e8a-ea9752a63180 · outbound

This paper cites 2021 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2021 , eprint=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.348374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:a2b43ae4cb6c933cb441dc580380bb4a5c93dc0113b3cbcb366b21d6d87d32e4

Observation 24de10b9-817f-4028-82f3-d3a849216c59 · outbound

This paper cites Diversity-incentivized exploration for versatile reasoning.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Diversity-incentivized exploration for versatile reasoning

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:16.815429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:0b4800ad937de8c0b1b6ce13b89c017f03b3baf060191a107d0ee59436aa0305

Observation 54e1a03b-6357-4945-85a8-13093bc99c7f · outbound

This paper cites 2025 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.328843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:9301191a96ba63041a789e53de3d0d8fa91d9acedeefaf09600cb28f61f9af8a

Observation 84330745-ad0b-4b9d-ac2b-3be2bc8f9003 · outbound

This paper cites 2025 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.310540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:2d9ff5447f7254e92ea81df347f3094db076a53212683af60a9ebd8bf6031e22

Observation e32dc371-2dc2-44a8-9eff-409e458b0ba4 · outbound

This paper cites 2025 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.314034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:47a8f68aab6d949e5aa846e4f15d215bd75eab7461686c8c2c58f127dc515dd3

Observation 1a7e5f71-1920-4c91-8646-0bd053605a74 · outbound

This paper cites 2025 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.300557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:d45dd579a2eee8c940beb059775484f67b739a96fd4408dd61ce835f50af4fbc

Observation ffe8d3ad-fd93-423f-9af3-dca52795b780 · outbound

This paper cites Reasoning with Sampling: Your Base Model is Smarter Than You Think.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Reasoning with Sampling: Your Base Model is Smarter Than You Think

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:16:05.617665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:de474c2034b7dbd15747396df93c916c10a171079f218fda9feb5d237508b935

Observation 8622f2a9-dbf4-40ce-a373-990bef82e04f · outbound

This paper cites Enhancing efficiency and exploration in reinforcement learning for llms.arXiv preprint arXiv:2505.18573.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Enhancing efficiency and exploration in reinforcement learning for llms.arXiv preprint arXiv:2505.18573

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:15.845341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:7a9a7c3b93b17424da089fba6e08e8573885c83920d76a8a2f41592560ee25bc

Observation b171ce6b-071d-4c54-8e73-8c5ae6b9eccf · outbound

This paper cites 2025 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.283159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:b30a045ade68e65df275d9459e619409bc9c93b229c7b20aeee0a5a3b48f6c1f

Observation 04663399-ea7e-4ef9-b83e-8582e627ac35 · outbound

This paper cites Notion Blog , year=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Notion Blog , year=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.263225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:50acaf2fc9f11855abf16f46c406c9a07c749605caada0d3db0b986292e83c95

Observation 984540f4-233e-4855-ab98-9c3de7490746 · outbound

This paper cites NeurIPS , year=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity NeurIPS , year=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.279733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:1171b4569b77188d46a974323f287272ddebe358d0c5b6590c4a71cc30c0edd3

Observation 282031cd-6c11-4d82-b87c-0e9cc8b09057 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:2b8021b85e94beeb44804ca108453366ef0713ec564ed930379cd2e38d1985b0

Observation f2bb6b0f-2b72-42da-aa49-42e91c9e71fb · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T15:26:15.514567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:ce1b95dce7bcf59d30edea0faf5ebdbcbe5c9e8358f3251515a0030eae6b2111

Observation eb631d45-676c-41d6-adb8-2f2b66495092 · outbound

This paper cites Bowman , booktitle=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Bowman , booktitle=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.286498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:1e7f9475cfe9e50558bb4c2932d9452a684436b5630c6ac110c3d4a5610ffb7d

Observation bc507de8-a56e-4be9-9afa-962b35aa80f9 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Advances in Neural Information Processing Systems , volume=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.290106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:a65ee5f4cafc8abb020e75354afe8ab7c68d56e1f0e1e42104f49ea34e5a8d90

Observation 8a1c6b22-a6c6-4cc4-bae0-0713e68beac9 · outbound

This paper cites Qwen2.5: A Party of Foundation Models , url =.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Qwen2.5: A Party of Foundation Models , url =

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.293524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:24178aac26c8f19f1ba8a135bf6c66c6b305c694c79180fa6e3dfb2fea7a14a0

Observation cd0037da-7843-47a4-9362-e2b42f68043d · outbound

This paper cites 2025 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.307351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:2fb06ce39572b27c0609b96552a99c77b927b7a6bfe23fce74f9c4ecc1362450

Observation 6f9a438e-d767-480b-9e90-66e45d68f524 · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:15.655822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:22965eacb58dd560ad40741d13edad3053af0d0c083db417e24f07688f16029c

Observation 784e7443-325a-4bcf-9740-4902538e4ec0 · outbound

This paper cites Differential smooth- ing mitigates sharpening and improves llm reasoning.arXiv preprint arXiv:2511.19942.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Differential smooth- ing mitigates sharpening and improves llm reasoning.arXiv preprint arXiv:2511.19942

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:16.118404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:d205f69ca6db472166fb0bdcc88653abca01fac325c300521b43e0fcb522f94d

Observation 820c6cfe-775d-41e3-8e00-a5fdce9977f2 · outbound

This paper cites Transactions on Machine Learning Research , volume=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Transactions on Machine Learning Research , volume=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.304077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:b0b1e8756309c640c63461ac94f8289eea83062ce32d7b0ca5fb2ebe44560a55

Observation de914d03-9ffa-4678-a00e-b646b81acce5 · outbound

This paper cites Hugging Face repository , volume=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Hugging Face repository , volume=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.266979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:137cad2b7cafdeaea8462e20d4cd100ac35765e105de72af43c95ea96e9ccda4

Observation 65d3827a-0af7-4dd5-8a12-8ecfc1a8951b · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:16.344559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:d7dcdaacf5059dfe1b83401630f38f52385832afe6fc252f485751f976696729

Observation 0171a140-eb49-405c-851a-3a052f9101c8 · outbound

This paper cites 2025 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.270368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:4e118db83a718e57cbc00ff88dbd1b093e1b4cfefa8fdcc26d3937a55f65ea36

Observation 7081bfc4-1cd2-4b53-9ab6-7d92e432a443 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T15:26:16.276142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:73e81a5e4efa07d69c40a89cc82107d10a2c72d918a1f056a3239fe3f6efa695

Observation 9ebee529-2a35-42a0-acdf-3447cf0af6e0 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:26:15.994525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:a3915a85caf3f99b7a4e243e6f5b44e600d5a020daf6325fb5095e5a9aec6787

Observation e9ba9776-59af-4dd6-a264-345e83e61ca9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Proximal Policy Optimization Algorithms

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T15:26:15.760359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:7cc744d1732612198a2efedcfc9273064ba00633dead8638269052fd40b44b12

Observation 1c77f848-df2a-4b1f-84ae-2047ba73f338 · outbound

This paper cites RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-28T02:04:10.066146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:abf0e7037c9ad8962b3bc2d1338c6bc276df375db3c5b04497feba00acccbc90

Observation bc8eb1c4-2005-4bcb-8c0a-d00995cf9619 · outbound

This paper cites Outcome-based Exploration for LLM Reasoning.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Outcome-based Exploration for LLM Reasoning

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:15.307690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:efd782c7ca94280f520cf65f6e1cad3dc90a98d9430b3d366de29423a391fe26

Observation 7b16acb1-5208-4a5e-b25c-d926a4b56189 · outbound

This paper cites ACL 2026 , year=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity ACL 2026 , year=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.276015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:f763eacf1e3a7025ea7cc58adb1fa93c96ddbeed21c8fc7bdea3e391b0b3e941

Observation 3633c4fc-9d87-493b-af8a-864014a4dd39 · outbound

This paper cites Advances in neural information processing systems , volume=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Advances in neural information processing systems , volume=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.297164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:a4b2ad3f4a80b9ac7d79dec5e4c08e23d02037c17d4201386b9221b7ba9511e9

Observation ad0f8c17-b2b0-4238-9e09-7a9f03c753de · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.317613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:8b7da99f4d66147f5a45c47e9df791b2a4084b19b64f51a627806f38a0a52793

Observation 1549946c-0990-42cc-9b04-83df2f79463d · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Does Reinforcement Learning Really Incentivize Reasoning Capacity in

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.322214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:a6e69176c40e0c2832e616c4e6b2878a4ded658780c7437199d8a5067954b929

Observation ba8c1a5a-9146-4eaa-97a1-427fb0dae1fc · outbound

This paper cites 2025 , eprint=.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T15:46:16.325558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:2d092e8d5625874608d6ceda2948ec805d8cb5eb68428813c8604d6800c8edd8

Observation 405b115a-8a38-42b3-add9-9e8eb14af53e · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Reasoning with Exploration: An Entropy Perspective

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:20:56.524904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:999928969a0f4a9008d89887d3d21e0b3b3b9141c2e0d99b373b395890845c97

Observation 025d54a3-7039-44ad-a8b4-e51c809af7c7 · outbound

This paper cites F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:11.286378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:d8752e5698c884ee6bbebb04aa716fb00a72b7076f13a6cf939c0cf9d7460a16

Pith citing papers

Observation 490492f5-7b91-4f81-b8db-d1e281bbd6c6 · inbound

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning cites this paper.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.626662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.345264Z digest=sha256:fbf08b88a62e0bf83100332673e4a22b1c8c9863ea50a37ef22edd56fed980b9