Pith. sign in

Paper Citation Record · LEDGER

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

As of 21 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.06987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06987 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T22:18:13.418579Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact29
  • verified fuzzy15
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ecaa784-1b7b-45af-a597-0c1b5a8609a2 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.546147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:e63fbf81b85d70df554df2b8ee5afe595edcdd8e0bc31618e63315886974950c

Observation 2d1d6d8d-9ef5-4661-8a4b-9327ceb38a45 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Constitutional AI: Harmlessness from AI Feedback

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.194238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:82d4e4869530aff831e11533bc52e7be9e97bab30a86e8012ee7327db55d75f8

Observation 97768206-e2e7-458a-9f90-51011be5ec09 · outbound

This paper cites Efficient reinforcement learning with semantic and token entropy for llm reasoning.arXiv preprint arXiv:2512.04359, 2025.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Efficient reinforcement learning with semantic and token entropy for llm reasoning.arXiv preprint arXiv:2512.04359, 2025

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:26:37.144898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:44fb5e94106ab0088c708ac2dabf6f65621357816759a2f4087c68f7962286fe

Observation 43d3cb3c-2cbf-48be-a23e-6e0625c73da5 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.205825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:6478f6c856dc873172debdb690e4c0d94d7c0c4c476ae54ab46060e86dae13c4

Observation 6b1f59e4-5202-4fe7-986e-3fd57d6a8090 · outbound

This paper cites American invitational mathematics examination-aime 2024.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma American invitational mathematics examination-aime 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.547848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:7e1b05f2edb59654b5aba92a26da82dbfa4e70dcfb73908005829d652c93febd

Observation 84f76c6a-8d46-4b69-a342-163eef14c5b3 · outbound

This paper cites Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.172379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:ea9302d541b71f143b546ac82c08a4d647eae7cc1b75b9e5ae217e9accfcb6d8

Observation 9c6e3eb3-76f2-41ff-92a6-64d5813f8103 · outbound

This paper cites Hero, and Sijia Liu.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Hero, and Sijia Liu

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.558585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:03e5968b14f593373f8b4db5a9a304e929d4d6eddccf7cfeab7637e5b61ed386

Observation 8cd12236-c4cb-407c-8299-b389a8f2607f · outbound

This paper cites an unresolved cited work.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-07-09T22:26:37.570176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:81469b45c6d87c11b3bb5dabb290f63684c33b15bf47d3833c7a491785cf7de2

Observation 05abc2b9-07d0-4070-a40b-24c2660f3ef5 · outbound

This paper cites Soft Adaptive Policy Optimization.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Soft Adaptive Policy Optimization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.220833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:2dbef7eba6426303f51b923263c1ebdddc66301e39f568cc409c2c68320b1c5f

Observation e2e5cef2-b1a7-4c72-b972-d78c584c3e88 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.178406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:9488b757c6537029ec0ffbdf11f51e980475974b2e57f6221ae319b767250d41

Observation 188c66a1-c992-44cb-ba81-5a4763573b8b · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.556925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:3bfab32cbb0a7fed87b45db07d3d4eb2e981879a0bf2d2a5323961e14c423561

Observation a638b56e-b728-4e8a-beb6-e719adf54287 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.184446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:b610891e6f45f4e42e4aa106c03653af984610e36caea6e98d2eb9c625d72e01

Observation ed73da8b-ba1d-4468-bba3-9be437561718 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Efficient memory management for large language model serving with pagedattention

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.560237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:d35063d325274f23f796fdcb5fc2c3dd1a81e6a9b52781a02eb925209eda822d

Observation 46b2b7b2-b5f5-41c7-96ad-62e48316a218 · outbound

This paper cites Solving quantitative reasoning problems with language models.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Solving quantitative reasoning problems with language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.551171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:330e2b43882f12726d16d98fd31a650d49a688ae69b39206d1baf2590b5c7b3d

Observation 8210591f-20ee-4bc1-b912-975aaa72e5e7 · outbound

This paper cites Back to basics: Revisiting exploration in reinforcement learning for llm reasoning via generative probabilities.arXiv preprint arXiv:2602.05281.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Back to basics: Revisiting exploration in reinforcement learning for llm reasoning via generative probabilities.arXiv preprint arXiv:2602.05281

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:26:37.157612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:705cb806c849b21ce6331bcd64027ee4213f206f74a8ae11fcc626980c5ebaad

Observation 1af9b955-f292-4b10-9334-8e17eeb058ec · outbound

This paper cites Bandpo: Bridging trust regions and ratio clipping via probability-aware bounds for llm reinforcement learning.arXiv preprint arXiv:2603.04918, 2026.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Bandpo: Bridging trust regions and ratio clipping via probability-aware bounds for llm reinforcement learning.arXiv preprint arXiv:2603.04918, 2026

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:26:37.163679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:1e9a5943749df2c77e87b0fc5e37c59cb95b867ebbb053a73cc9f1966a5a4440

Observation c424ffdf-b5d7-4941-8fbf-d34faa72f81e · outbound

This paper cites Let’s verify step by step.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Let’s verify step by step

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.561765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:cb75a59f62fd0a8b86b27c88df5d95b354ece3c267d86bed36e9907551385f31

Observation a73a5e02-b75f-4aa2-9b0f-97f26eabc1bc · outbound

This paper cites Length-unbiased sequence policy optimization: Revealing and controlling response length variation in rlvr.arXiv preprint arXiv:2602.05261.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Length-unbiased sequence policy optimization: Revealing and controlling response length variation in rlvr.arXiv preprint arXiv:2602.05261

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:26:37.208422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:55c636f230e685180ad95e6dc052f3c0654a7266bfaa5f8c1550e3891c836224

Observation 1e606b2d-4d75-4a88-9c61-66afb52ae5f8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Understanding R1-Zero-Like Training: A Critical Perspective

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.169302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:62bd1b421c819290441e32fa0b0dad0eadfe99ca8da9788127e003bc97bfc12b

Observation dad217d4-b78d-4c7e-9ad1-4b06b0080d45 · outbound

This paper cites Inter-GPS: Interpretable geometry problem solving with formal language and symbolic reasoning.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Inter-GPS: Interpretable geometry problem solving with formal language and symbolic reasoning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.549493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:59c161e61b5b17895aa61964a29994c3e478ab1c1c96dcc0927e64bfd4286ca1

Observation 307656d7-c748-419f-9431-0a0caaa19fe3 · outbound

This paper cites Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:26:37.154897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:4af3d3d9596efa4ac6d06f39ee9058542b53ef29430e84bc724d4501659eb4e8

Observation 5752603f-8b67-48e7-b95e-6c9657a8108a · outbound

This paper cites Clip your sequences fairly: Enforcing length fairness for sequence-level rl.arXiv preprint arXiv:2509.09177.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Clip your sequences fairly: Enforcing length fairness for sequence-level rl.arXiv preprint arXiv:2509.09177

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:26:37.152364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:5df04823e4dae04a5d1d241c7d7e303419485b1b1771ef28c38d48d2a6b4ff1c

Observation 85a27c07-f2b0-4cb9-adb9-bb2e7b5b1be5 · outbound

This paper cites Training language models to follow instructions with human feedback.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Training language models to follow instructions with human feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.544379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:a6acae4ac902237b6208b5d71d76b73cd5fbd72b6adac46f8b8483280af8d619

Observation 6c9d7f76-ac73-4975-b320-deeb51bd696d · outbound

This paper cites Rethinking the Trust Region in LLM Reinforcement Learning.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Rethinking the Trust Region in LLM Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.210776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:23dd2b0006e6839da05e2dda552c2f7826f4f3c83b30d53a2802b61c2cd2b791

Observation f9bf1274-34d2-4bcd-b657-6a9ea585f631 · outbound

This paper cites Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.213079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:0ac2e99de929c1144a9fb34438437f58518c62ab57897cfd5c338e31b2a00b61

Observation 6d30ce88-df57-4bcd-b92d-b05c0d44fdb4 · outbound

This paper cites Trustregionpolicyoptimization.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Trustregionpolicyoptimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.566613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:bb67fdb855cce324391b4204cc8f61b9767ef35ff30d6da8c8fc2510a684be54

Observation 81b169db-2a72-4d47-9a77-fcf3382ee7a8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Proximal Policy Optimization Algorithms

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.147239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:b3ddeddae6a73862ca20d3395005f0377fdb3a06f245477ac0ecd63acbbcf195

Observation 4e27c231-70d3-4c1d-94d4-086dce0ae3ab · outbound

This paper cites Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.203495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:0a8438961c32412121eddca51555837bb102a4dca9ce324acad2f2e4635f4a63

Observation b0ca443f-67e5-4285-b13c-b834c9c04eb1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.149556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:ffac0b337b99ce15ab7e0d018768e4cd6dced80a2e7d864ece0aa79a62e2e504

Observation f24699ae-8f49-4063-a8fd-548d4f44039d · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Hybridflow: A flexible and efficient rlhf framework

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.563365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:8c1025a670aebdc0d5918888fac0f580108765598290822feb528954723156de

Observation df14d639-5808-4c95-ab37-e8443cd1c7ff · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.181003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:1168fd8e2c55acea362290ff46f402bde06f04b7909e66edb7536380232dcc12

Observation 3c35c6cd-3213-478b-acbb-033b90cb0194 · outbound

This paper cites Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:26:37.166428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:c800aa3141552ad8ce6f164881ca79f18661aad3504c5571f382a4d2d95fa2cb

Observation f62037ce-208b-412f-8689-8761caf75da3 · outbound

This paper cites When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.218159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:24d4884055c8f63172d4ca81d824947a2624130e9f513d8f17aa8f85675f2b37

Observation 30250df8-2189-400b-b35a-791e85d7d501 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.568414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:8b1ca618a61f4462b08e8eaae3895e38f595e1151b30dc076e30cec1f9debf04

Observation 11e7ffd4-62da-4fb0-ad77-3e0b2b1ce165 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.189407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:e7988dcabf322a6b49945ecb60d6aa12be717941e0b387a1b50784647d3a6f72

Observation 5c3659a4-0f46-48e2-8f28-4d8e9753cdd4 · outbound

This paper cites Qwen3 Technical Report.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Qwen3 Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.196554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:f40fb929d22ef65064bef8741b2f8d53dfdd03ce0c6e4d0e3522d0b30b855d5f

Observation 7baf2c4e-27fe-4c60-bb29-22add431be8f · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.198934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:40c0f649a3dc1aaaf216be081a6e6c8d92cfcc5ad4939b9126b76a6e30a65b03

Observation 118ea934-f2e3-4e36-90a5-a9c9bb703074 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.223302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:54ae613906d7daf1392eb02de22d1ff513df81f4d529e7296cf766b7b33baaa3

Observation 802ea673-266f-46f2-80bd-7c5365a9ea4d · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Star: Bootstrapping reasoning with reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.565043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:fff71d75c505ab82f4715667292beed921104125f16bd8db4cc082b2e9ec9422

Observation b6faac81-0a95-46fa-b6b0-1ed56c35f881 · outbound

This paper cites George E Uhlenbeck and Leonard S Ornstein.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma George E Uhlenbeck and Leonard S Ornstein

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T22:26:37.187015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:4342ebd904c96ee7f47cc9a558d5cc9981578bace57d7fc8002506dfd1066e40

Observation 73096a43-0a93-48c7-b6ab-9b570dd25d25 · outbound

This paper cites Geometric-mean policy optimization.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Geometric-mean policy optimization

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:26:37.192014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:241b9e029b89202dd375d877fe75d562a5ece88ceb93a89cbd89fb654d0fd793

Observation 31dc8757-8ed5-4869-9478-1e6dd1adb69c · outbound

This paper cites Group Sequence Policy Optimization.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Group Sequence Policy Optimization

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:26:37.201213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:13de98a7173ae9c2bdec87e46af20c9b05d65e9dde7d526884217277f9e374af

Observation 70b4d23e-3657-4bc5-bef7-a357d4b22626 · outbound

This paper cites Prosperity before collapse: How far can off-policy rl reach with stale data on llms?.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Prosperity before collapse: How far can off-policy rl reach with stale data on llms?

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:26:37.215591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:c6642f5d21fd2f4d10dd741d557c8ca36b2c15d031968fdd05b463e79e976cce

Observation 22670313-d73a-4f44-9de9-f7a4f3f27c49 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:26:37.175791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:46bac8834dc69cc701ba7758ba86f836ce69729ff442c8a8e2d006a1040b3390

Observation ea02100c-396a-458b-8896-cb07df26438e · outbound

This paper cites If πθ < π lower, the gradient is nullified to prevent the probability from dropping too low (over-penalization).

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma If πθ < π lower, the gradient is nullified to prevent the probability from dropping too low (over-penalization)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.555237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:26b5f43979b2012058f35069616256166d784d3db615c8d498d86f66b7a0a8f5

Observation c478bb1e-f250-47ed-a885-f0c332ddbdec · outbound

This paper cites 1 G GX i=1 ˆAi 1 sg(πθ(oi|q)) 1 |oi | ∇θ πθ(oi|q) 1 |oi | # =E q,{oi}.

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma 1 G GX i=1 ˆAi 1 sg(πθ(oi|q)) 1 |oi | ∇θ πθ(oi|q) 1 |oi | # =E q,{oi}

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:26:37.553040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T22:18:13.418579Z digest=sha256:4b2649a57e74885ebe3210ef43af14a416d90688d0d554479c59a1142b3023bf

Pith citing papers

No inbound Pith citation observations are available.