Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T12:35:53.841154Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2606.03021.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T12:35:53.841154Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9cd82d61-b56b-4d8f-9f13-34bd64e80f03 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning It should only elaborate on the high-level strategies and concepts, without going into specific calculations
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 385afc16-3699-4125-8320-8e9d49d717e7 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 899a701b-d902-4f49-be45-ba8587e7d2b1 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 979abb64-bdab-4e0d-b2b6-67b019b00cd8 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Qwen3 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7129faf-0e80-44ad-910e-eebc887b4ae0 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning propose-select-think
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76bdf2ed-787b-4446-a334-e03c072b8e29 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Yes” as a measure of the similarity between the two candi- date solutions. As shown by “HDPO (LLM-Div)
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39731c6f-2d28-4915-b94e-3bc739c303af · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd16bfac-7cb3-42c3-ad1b-fe3cd3f88465 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning ex- plore–evaluate–select
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e99ce3d-65a6-453e-933c-ba1c6eb4c932 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning [1]”, “[2]
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d28b855-43b2-407a-a798-1051481b37e7 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef730c80-1d06-4cc6-9e13-5492557f0045 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Yes", otherwise output
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9fa484b-de64-4e03-9d2c-d9ef92b202b8 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a052d0c-7bf0-476b-92d9-46860eae661c · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d53f7e89-9203-4e0a-8659-4df956b38678 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning We need to find the values ofa,b, andcthat maximize|a|+|b|+|c|while satisfying these constraints
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c57d91c-3805-4377-ba7f-785502f9e381 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Then express a, b, c in terms of these values and use linear programming or symmetry arguments to maximize |a| + |b| + |c|
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9beee6ca-e8d8-48be-b696-a85acde1f562 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Apply the method of Lagrange multipliers to maximize the linear functional |a| + |b| + |c| subject to the quadratic constraint
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1248710d-c51b-4cf2-8ab1-450984e3ebb0 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning However, this may miss the global maximum if the optimal polynomial is not symmetric or has non-zero a and b
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3428ce2f-5445-452a-b208-d0731abe43b4 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Scale and shift the Chebyshev polynomial to satisfy the bound|P(x)| ≤1and compute the coefficients to find the maximum of |a| + |b| + |c|
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4a4c7f8-5cfd-49b1-88d5-6bc9f4879966 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning - PointPis 4 units away from the circle, so the distance fromPto the centerOis6 + 4 = 10
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24ba2916-f45f-4821-9ce0-eadac1e024ca · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning But sincePis 4 units away from the circle and AB is parallel to ← →OP, the perpendicular distance fromOto ABmust be 4 (as 8 would place AB outside the circle)
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2be02ac-7311-420d-92d5-4a9a777144d8 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning - The chord ABis parallel to the x-axis and 2 units below the x-axis (since the distance from Oto ABis 4)
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a7f3fdb-33e5-41a6-acbe-b62e5ad9d97d · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dcc36f0-25ae-4a07-bf5c-4094777d030c · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning 19 Case 2 (Generation Model: Qwen2.5-Math-7B-HDPO) Question: CircleOhas radius 6
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a930312e-37f1-4013-9a76-38a24a41f4a8 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning SinceABis parallel toOP and the distance between them is 2, the perpendicular distance fromOtoABis either4 + 2 = 6 or4−2 = 2
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a760fd0a-70d2-482e-9e34-cfe6b241c814 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d786531-c111-4ad6-bd7c-a381c23304cc · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Since chordABis parallel to ← →OP, it is horizontal, and the distance betweenAB and ← →OPis 2, soABis either aty= 2ory=−2
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb55b2d-a1c6-44a6-b6c4-f3d6fbd63325 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning propose-select-think
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1545a31a-f893-4671-aacc-0b9ec204f355 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Then apply the sum of cosine series formula for angles in arithmetic sequence, simplifying the resulting expression using symmetry and periodicity of the cosine function
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b98ad9-4d0a-47e5-9252-bb491351ac34 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning However, this approach lacks precision and relies on approximation, making it unsuitable for exact computation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77077b0-1efc-4674-897d-4c46714e7023 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83cc1e74-017d-4047-a9aa-072a95d4cd11 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning </Candidate Solutions> <selected>[1]</selected> <thinking> We are given the sum: sin2 4◦ + sin2 8◦ + sin2 12◦ +· · ·+ sin2 176◦ This is a sum ofsin 2 θforθ= 4k ◦ wherek= 1,2,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86cde73d-8e00-4f3c-9df8-f0a67f849b41 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a118376-6d2a-4c97-be81-c204f113c223 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f2bc28b-a410-497d-b18d-f30fbda8d347 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Learning to Reason under Off-Policy Guidance
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db52df6-2413-495b-b2a4-73745f4af1a5 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.