Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T22:18:13.418579Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.06987.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T22:18:13.418579Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6ecaa784-1b7b-45af-a597-0c1b5a8609a2 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2d1d6d8d-9ef5-4661-8a4b-9327ceb38a45 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Constitutional AI: Harmlessness from AI Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 97768206-e2e7-458a-9f90-51011be5ec09 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Efficient reinforcement learning with semantic and token entropy for llm reasoning.arXiv preprint arXiv:2512.04359, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 43d3cb3c-2cbf-48be-a23e-6e0625c73da5 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6b1f59e4-5202-4fe7-986e-3fd57d6a8090 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma American invitational mathematics examination-aime 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 84f76c6a-8d46-4b69-a342-163eef14c5b3 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9c6e3eb3-76f2-41ff-92a6-64d5813f8103 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Hero, and Sijia Liu
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8cd12236-c4cb-407c-8299-b389a8f2607f · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 05abc2b9-07d0-4070-a40b-24c2660f3ef5 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Soft Adaptive Policy Optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e2e5cef2-b1a7-4c72-b972-d78c584c3e88 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 188c66a1-c992-44cb-ba81-5a4763573b8b · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a638b56e-b728-4e8a-beb6-e719adf54287 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ed73da8b-ba1d-4468-bba3-9be437561718 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Efficient memory management for large language model serving with pagedattention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 46b2b7b2-b5f5-41c7-96ad-62e48316a218 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Solving quantitative reasoning problems with language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8210591f-20ee-4bc1-b912-975aaa72e5e7 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Back to basics: Revisiting exploration in reinforcement learning for llm reasoning via generative probabilities.arXiv preprint arXiv:2602.05281
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1af9b955-f292-4b10-9334-8e17eeb058ec · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Bandpo: Bridging trust regions and ratio clipping via probability-aware bounds for llm reinforcement learning.arXiv preprint arXiv:2603.04918, 2026
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c424ffdf-b5d7-4941-8fbf-d34faa72f81e · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Let’s verify step by step
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a73a5e02-b75f-4aa2-9b0f-97f26eabc1bc · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Length-unbiased sequence policy optimization: Revealing and controlling response length variation in rlvr.arXiv preprint arXiv:2602.05261
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1e606b2d-4d75-4a88-9c61-66afb52ae5f8 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Understanding R1-Zero-Like Training: A Critical Perspective
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dad217d4-b78d-4c7e-9ad1-4b06b0080d45 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Inter-GPS: Interpretable geometry problem solving with formal language and symbolic reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 307656d7-c748-419f-9431-0a0caaa19fe3 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5752603f-8b67-48e7-b95e-6c9657a8108a · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Clip your sequences fairly: Enforcing length fairness for sequence-level rl.arXiv preprint arXiv:2509.09177
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 85a27c07-f2b0-4cb9-adb9-bb2e7b5b1be5 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6c9d7f76-ac73-4975-b320-deeb51bd696d · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Rethinking the Trust Region in LLM Reinforcement Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f9bf1274-34d2-4bcd-b657-6a9ea585f631 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6d30ce88-df57-4bcd-b92d-b05c0d44fdb4 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Trustregionpolicyoptimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 81b169db-2a72-4d47-9a77-fcf3382ee7a8 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Proximal Policy Optimization Algorithms
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4e27c231-70d3-4c1d-94d4-086dce0ae3ab · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b0ca443f-67e5-4285-b13c-b834c9c04eb1 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f24699ae-8f49-4063-a8fd-548d4f44039d · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Hybridflow: A flexible and efficient rlhf framework
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation df14d639-5808-4c95-ab37-e8443cd1c7ff · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3c35c6cd-3213-478b-acbb-033b90cb0194 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f62037ce-208b-412f-8689-8761caf75da3 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 30250df8-2189-400b-b35a-791e85d7d501 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Simple statistical gradient-following algorithms for connectionist reinforcement learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 11e7ffd4-62da-4fb0-ad77-3e0b2b1ce165 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5c3659a4-0f46-48e2-8f28-4d8e9753cdd4 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Qwen3 Technical Report
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7baf2c4e-27fe-4c60-bb29-22add431be8f · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 118ea934-f2e3-4e36-90a5-a9c9bb703074 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 802ea673-266f-46f2-80bd-7c5365a9ea4d · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Star: Bootstrapping reasoning with reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b6faac81-0a95-46fa-b6b0-1ed56c35f881 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma George E Uhlenbeck and Leonard S Ornstein
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 73096a43-0a93-48c7-b6ab-9b570dd25d25 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Geometric-mean policy optimization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 31dc8757-8ed5-4869-9478-1e6dd1adb69c · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Group Sequence Policy Optimization
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 70b4d23e-3657-4bc5-bef7-a357d4b22626 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Prosperity before collapse: How far can off-policy rl reach with stale data on llms?
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 22670313-d73a-4f44-9de9-f7a4f3f27c49 · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ea02100c-396a-458b-8896-cb07df26438e · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma If πθ < π lower, the gradient is nullified to prevent the probability from dropping too low (over-penalization)
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c478bb1e-f250-47ed-a885-f0c332ddbdec · outbound
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma 1 G GX i=1 ˆAi 1 sg(πθ(oi|q)) 1 |oi | ∇θ πθ(oi|q) 1 |oi | # =E q,{oi}
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
No inbound Pith citation observations are available.