Pith. sign in

Paper Citation Record · LEDGER

Group Entropy-Controlled Policy Optimization

As of 4 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.16850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16850 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T19:49:04.365663Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3951752-7444-4b59-8d35-fce8cc369944 · outbound

This paper cites Intern-s1: A scientific multimodal foundation model, 2025.

Group Entropy-Controlled Policy Optimization Intern-s1: A scientific multimodal foundation model, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:01.679794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:01.679794Z digest=sha256:3f36ef831a75b7acdb3f52932c704171ceb8f3138f03e1486dd01ec287a861fd

Observation 3b46324e-edd0-4cfe-bcef-8a8227970069 · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.Advances in neural information processing systems, 29, 2016.

Group Entropy-Controlled Policy Optimization Unifying count-based exploration and intrinsic motivation.Advances in neural information processing systems, 29, 2016

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:01.738473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:01.738473Z digest=sha256:9803cffe3d948dcf0886c4188bfab137683eef8719b57df2ac1b3c8b767b515b

Observation 49e41946-7ab5-47bb-9600-10b3ca458802 · outbound

This paper cites Galaz-Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, Sarina M.

Group Entropy-Controlled Policy Optimization Galaz-Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, Sarina M

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:01.851451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:01.851451Z digest=sha256:073d862b8174fc258d3914e359c3f3dc929850acf7a977d365c9a06689ff3daf

Observation 9fe5b607-b04b-47ea-a9e6-ccaa109ae846 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Group Entropy-Controlled Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:01.949922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:01.949922Z digest=sha256:6f01ed0f4300e4c93a146b0327d5bf9000ff2c2ec675639b02c56cd6c8a45a85

Observation a62d2755-5846-4ce1-b287-335ffddf9c1e · outbound

This paper cites Reasoning with exploration: An entropy perspective.

Group Entropy-Controlled Policy Optimization Reasoning with exploration: An entropy perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.029824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.029824Z digest=sha256:4b0116e000337ddb3ced05912bce79c656a645eb730daa3a008844cc186bd190

Observation cbacc640-2d73-45f0-986b-c457f1abeb3c · outbound

This paper cites Lmdeploy: A toolkit for compressing, deploying, and serving llm.https: //github.com/InternLM/lmdeploy, 2023.

Group Entropy-Controlled Policy Optimization Lmdeploy: A toolkit for compressing, deploying, and serving llm.https: //github.com/InternLM/lmdeploy, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.088843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.088843Z digest=sha256:02850985a5e0a8f62d3f662e33c92c18a969bf9bc2699918789fadfb7464eb1a

Observation 04a0c6c7-b3be-495a-b945-f7751d827f69 · outbound

This paper cites Xtuner: A toolkit for efficiently fine-tuning llm.

Group Entropy-Controlled Policy Optimization Xtuner: A toolkit for efficiently fine-tuning llm

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.149150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.149150Z digest=sha256:5bdf55574de74186c89d260bfd1979e9ace0636a0d2950d601dca0d7a03821b6

Observation 3915b89a-996d-4c6c-9071-91fc0fc6946b · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Group Entropy-Controlled Policy Optimization The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.197871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.197871Z digest=sha256:b0d04fb4b773a124124d6a94a5ab402b499fda8e490e36d367c1a7672ebf5f40

Observation 4461ca0c-bc8c-4814-9e6a-cc35785395ca · outbound

This paper cites Physics: Benchmarking foundation models on university-level physics problem solving, 2025.

Group Entropy-Controlled Policy Optimization Physics: Benchmarking foundation models on university-level physics problem solving, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.255082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.255082Z digest=sha256:586047487c4ef01366f5d83b14f099d28f55b0b57f0db9de4874ffa8b2062460

Observation d7aedcc4-23a7-4909-afa3-fe71e0f4579a · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Group Entropy-Controlled Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.310544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.310544Z digest=sha256:eb8548cde8b1c69bddf0da87bfbb49e85d93baa38820386655afd57a21bb523d

Observation b8e718e4-f2f7-4523-b26b-60ae1d5827eb · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Group Entropy-Controlled Policy Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.455134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.455134Z digest=sha256:80f4c4304ee2dbe56bf837ff24d094e8c65dd6d7761a2fdfcf9726015e13f70d

Observation 297732f6-951e-4213-a9bd-9635c2534e19 · outbound

This paper cites OpenAI o1 System Card.

Group Entropy-Controlled Policy Optimization OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.573280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.573280Z digest=sha256:19cb0fa527bd8536a9ec5f428d08b24fc01321a1762a7f7c927d401a081f0312

Observation b9b1ccc2-fbc0-4dd9-9747-bd436fbbe1ff · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024.

Group Entropy-Controlled Policy Optimization Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.691616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.691616Z digest=sha256:547da38c054c6d363d5c5a6e2e6134c5a7ee92999f45697c8a54ab98f0429164

Observation 1f86e6df-98a4-4f95-8ca5-60eb243347d5 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

Group Entropy-Controlled Policy Optimization Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.784932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.784932Z digest=sha256:a38b0bcc35e5212a10d1b3f43b1a5d3a1b920cc5e07db8854b8c7361c14202bb

Observation b241928e-20ea-4a2c-a831-18038e847613 · outbound

This paper cites Chartqapro: A more diverse and challenging benchmark for chart question answering, 2025.

Group Entropy-Controlled Policy Optimization Chartqapro: A more diverse and challenging benchmark for chart question answering, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:02.899076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:02.899076Z digest=sha256:e67d081df0afb26969784dfef3f6c44990cfd734a0f7b2063295e4d7daa7076d

Observation 46aa9504-718a-4b22-829d-2be7b683ea05 · outbound

This paper cites Generalizing verifiable instruction following, 2025.

Group Entropy-Controlled Policy Optimization Generalizing verifiable instruction following, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:03.014633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:03.014633Z digest=sha256:7c7abba35c92edae0e1cdc9b3d294709e7cbdcdb8e8528d8b4ca1d6c468f7b3a

Observation d17d23f4-238b-49bf-9c12-5d6606564c5f · outbound

This paper cites an unresolved cited work.

Group Entropy-Controlled Policy Optimization Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:03.157878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:03.157878Z digest=sha256:a604b49e9f8dfb384e5380274e141ca0ccf3496860776bffae2f1b39173df987

Observation 201f0c88-98b9-4c20-a6a3-3affea4d4f7e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Group Entropy-Controlled Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:03.268973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:03.268973Z digest=sha256:f7d645fe71d692f81fe77f4e50fa25f2a15d0ea30784bfd3362c2488950858ad

Observation fd8d8ef2-3a78-475b-95c9-fb5b9b1df638 · outbound

This paper cites Qwenlong-l1.

Group Entropy-Controlled Policy Optimization Qwenlong-l1

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:03.394154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:03.394154Z digest=sha256:8e4d901a7914b895c9e588ae566c1e70a6b2392826986f1bfdfc0173243dce0a

Observation eab9fc93-71b2-4d8c-8344-cacc191fd3d1 · outbound

This paper cites MIT press Cambridge, 1998.

Group Entropy-Controlled Policy Optimization MIT press Cambridge, 1998

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:03.485381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:03.485381Z digest=sha256:4a4633fc857b0010f534d74164bfe76c841bdf8eb89bab0d3d3fd6b388fd5e97

Observation 89f03583-36ba-4414-add0-7d01e18afa4c · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Group Entropy-Controlled Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:03.579148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:03.579148Z digest=sha256:43ffd59f9d666e1ff61c3e180fc9ec6cd2a8f1936198774a9071f5b5022a2d8a

Observation 23ed1846-5ab3-416e-9831-df6360788df9 · outbound

This paper cites Tsybakov.Introduction to Nonparametric Estimation.

Group Entropy-Controlled Policy Optimization Tsybakov.Introduction to Nonparametric Estimation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:03.758433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:03.758433Z digest=sha256:486b9aa4f36b7a5b13e10c3fc3c8296bdb913f76c662342f26f51e07ec234ea2

Observation 22122e75-0946-4a36-806e-175ba395cd3a · outbound

This paper cites Dsdr: Dual-scale diversity regularization for exploration in llm reasoning.arXiv preprint arXiv:2602.19895, 2026.

Group Entropy-Controlled Policy Optimization Dsdr: Dual-scale diversity regularization for exploration in llm reasoning.arXiv preprint arXiv:2602.19895, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:03.876845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:03.876845Z digest=sha256:150618c715cfb8f4b291a09d14ab3f8cfefaac28835e58098fb24f1785e14858

Observation d7b49ccf-b3fd-4b4d-b601-a9cb8f659f00 · outbound

This paper cites On the entropy dynamics in reinforcement fine-tuning of large language models.arXiv preprint arXiv:2602.03392, 2026.

Group Entropy-Controlled Policy Optimization On the entropy dynamics in reinforcement fine-tuning of large language models.arXiv preprint arXiv:2602.03392, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:04.006924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:04.006924Z digest=sha256:1e4c8d18225ce27c3ea82482f5faed716dd22c0bd594b9c4e8e0eb350c8b96b0

Observation 3146c057-b560-43ff-bf72-2a9c7f828dd9 · outbound

This paper cites Cmphysbench: A benchmark for evaluating large language models in condensed matter physics, 2025.

Group Entropy-Controlled Policy Optimization Cmphysbench: A benchmark for evaluating large language models in condensed matter physics, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:04.166574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:04.166574Z digest=sha256:4e8d78b60ec3176338a773c18607b35dffc5eb85d61b9c76b413388b7855a65c

Observation 82f2643d-472c-4884-a52b-aafeb3818f53 · outbound

This paper cites Entropic: Towards stable long-term training of llms via entropy stabilization with proportional-integral control.arXiv preprint arXiv:2511.15248, 2025.

Group Entropy-Controlled Policy Optimization Entropic: Towards stable long-term training of llms via entropy stabilization with proportional-integral control.arXiv preprint arXiv:2511.15248, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:04.247560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:04.247560Z digest=sha256:a7313a5a8d1a6efbe76aef54b351bcd4497d7ed7409b498fca52ef4c748cc405

Observation 743f9664-442b-4338-9ff5-41cb55f937b8 · outbound

This paper cites Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark, 2025.

Group Entropy-Controlled Policy Optimization Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:04.312886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:04.312886Z digest=sha256:ed3631cff56e9ae7cde163412fc62ab0aefd159e7fe40f1920a3f949a66973f7

Observation 35cdd4e2-41d7-4449-a764-0dc0ecc016d3 · outbound

This paper cites Scientists’ first exam: Probing cognitive abilities of mllm via perception, understanding, and reasoning, 2025.

Group Entropy-Controlled Policy Optimization Scientists’ first exam: Probing cognitive abilities of mllm via perception, understanding, and reasoning, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T19:49:04.365663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:49:04.365663Z digest=sha256:7f73ca41dc64cbcd60db3962f3da8598945aedd1739571f16d61e9ffa2dff069

Pith citing papers

No inbound Pith citation observations are available.