Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T19:49:04.365663Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.16850.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T19:49:04.365663Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c3951752-7444-4b59-8d35-fce8cc369944 · outbound
Group Entropy-Controlled Policy Optimization Intern-s1: A scientific multimodal foundation model, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b46324e-edd0-4cfe-bcef-8a8227970069 · outbound
Group Entropy-Controlled Policy Optimization Unifying count-based exploration and intrinsic motivation.Advances in neural information processing systems, 29, 2016
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e41946-7ab5-47bb-9600-10b3ca458802 · outbound
Group Entropy-Controlled Policy Optimization Galaz-Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, Sarina M
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe5b607-b04b-47ea-a9e6-ccaa109ae846 · outbound
Group Entropy-Controlled Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a62d2755-5846-4ce1-b287-335ffddf9c1e · outbound
Group Entropy-Controlled Policy Optimization Reasoning with exploration: An entropy perspective
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbacc640-2d73-45f0-986b-c457f1abeb3c · outbound
Group Entropy-Controlled Policy Optimization Lmdeploy: A toolkit for compressing, deploying, and serving llm.https: //github.com/InternLM/lmdeploy, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a0c6c7-b3be-495a-b945-f7751d827f69 · outbound
Group Entropy-Controlled Policy Optimization Xtuner: A toolkit for efficiently fine-tuning llm
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3915b89a-996d-4c6c-9071-91fc0fc6946b · outbound
Group Entropy-Controlled Policy Optimization The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4461ca0c-bc8c-4814-9e6a-cc35785395ca · outbound
Group Entropy-Controlled Policy Optimization Physics: Benchmarking foundation models on university-level physics problem solving, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7aedcc4-23a7-4909-afa3-fe71e0f4579a · outbound
Group Entropy-Controlled Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e718e4-f2f7-4523-b26b-60ae1d5827eb · outbound
Group Entropy-Controlled Policy Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 297732f6-951e-4213-a9bd-9635c2534e19 · outbound
Group Entropy-Controlled Policy Optimization OpenAI o1 System Card
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b1ccc2-fbc0-4dd9-9747-bd436fbbe1ff · outbound
Group Entropy-Controlled Policy Optimization Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f86e6df-98a4-4f95-8ca5-60eb243347d5 · outbound
Group Entropy-Controlled Policy Optimization Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b241928e-20ea-4a2c-a831-18038e847613 · outbound
Group Entropy-Controlled Policy Optimization Chartqapro: A more diverse and challenging benchmark for chart question answering, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46aa9504-718a-4b22-829d-2be7b683ea05 · outbound
Group Entropy-Controlled Policy Optimization Generalizing verifiable instruction following, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d17d23f4-238b-49bf-9c12-5d6606564c5f · outbound
Group Entropy-Controlled Policy Optimization Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 201f0c88-98b9-4c20-a6a3-3affea4d4f7e · outbound
Group Entropy-Controlled Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd8d8ef2-3a78-475b-95c9-fb5b9b1df638 · outbound
Group Entropy-Controlled Policy Optimization Qwenlong-l1
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eab9fc93-71b2-4d8c-8344-cacc191fd3d1 · outbound
Group Entropy-Controlled Policy Optimization MIT press Cambridge, 1998
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f03583-36ba-4414-add0-7d01e18afa4c · outbound
Group Entropy-Controlled Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ed1846-5ab3-416e-9831-df6360788df9 · outbound
Group Entropy-Controlled Policy Optimization Tsybakov.Introduction to Nonparametric Estimation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22122e75-0946-4a36-806e-175ba395cd3a · outbound
Group Entropy-Controlled Policy Optimization Dsdr: Dual-scale diversity regularization for exploration in llm reasoning.arXiv preprint arXiv:2602.19895, 2026
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7b49ccf-b3fd-4b4d-b601-a9cb8f659f00 · outbound
Group Entropy-Controlled Policy Optimization On the entropy dynamics in reinforcement fine-tuning of large language models.arXiv preprint arXiv:2602.03392, 2026
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3146c057-b560-43ff-bf72-2a9c7f828dd9 · outbound
Group Entropy-Controlled Policy Optimization Cmphysbench: A benchmark for evaluating large language models in condensed matter physics, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82f2643d-472c-4884-a52b-aafeb3818f53 · outbound
Group Entropy-Controlled Policy Optimization Entropic: Towards stable long-term training of llms via entropy stabilization with proportional-integral control.arXiv preprint arXiv:2511.15248, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 743f9664-442b-4338-9ff5-41cb55f937b8 · outbound
Group Entropy-Controlled Policy Optimization Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35cdd4e2-41d7-4449-a764-0dc0ecc016d3 · outbound
Group Entropy-Controlled Policy Optimization Scientists’ first exam: Probing cognitive abilities of mllm via perception, understanding, and reasoning, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.