Pith. sign in

Paper Citation Record · LEDGER

Stable Reinforcement Learning for Efficient Reasoning

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 7 inbound Pith citation observations for arXiv:2505.18086.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18086 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:38:47.682695Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:09:39.636062Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:46:20.251611Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9a7d3d8-c24b-4338-a122-6e5ccd4fa911 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Stable Reinforcement Learning for Efficient Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:42.693276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:42.693276Z digest=sha256:e4c920096dbff2b5bf0483d62024f439604926a9efc3c6d617c37e71017df315

Observation cbe93923-50d0-41cd-9438-741854849ba4 · outbound

This paper cites Scaling Laws for Neural Language Models.

Stable Reinforcement Learning for Efficient Reasoning Scaling Laws for Neural Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:42.839047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:42.839047Z digest=sha256:bb23ebcf3af2a8ce2f7b96457c867f03c14f9a165b77dda98c5c12aeb59f1f8b

Observation 3ca8bd60-b550-43a9-b38e-1fa1042351e7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Stable Reinforcement Learning for Efficient Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:42.993274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:42.993274Z digest=sha256:bc9e606e2d40c52114678e14c2a5eed2d79de7449b7e5b20bb79fee3e666f8ba

Observation ecb8034f-13b8-4298-b5b8-19ee820e5f47 · outbound

This paper cites Qwen3 Technical Report.

Stable Reinforcement Learning for Efficient Reasoning Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.173424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.173424Z digest=sha256:be59c277277b36ff9b6d7454f6e1d0c2d474328824dccb179ce1c324ac1544d9

Observation ad775c01-2cc7-44cc-8f75-7f57e29df208 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Stable Reinforcement Learning for Efficient Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.302144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.302144Z digest=sha256:4ca531502dc23f10994ae0ad5a3c4715eea56e6702b96636ee7d8063fdba5690

Observation 513424c0-8561-4ecf-946a-cf6e63e53b55 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Stable Reinforcement Learning for Efficient Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.477652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.477652Z digest=sha256:2ec7f8b3132dd6fecf6f85a04c3f9e56eade9c14151738304d6fcb51fc50033e

Observation a1f10c50-543c-4187-9aba-4dc66166503e · outbound

This paper cites Training language models to follow instructions with human feedback.

Stable Reinforcement Learning for Efficient Reasoning Training language models to follow instructions with human feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.607566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.607566Z digest=sha256:4bde24912d0d98d55b614c7ff3c59d117dffbf47df83942b2fbad049a4e358e0

Observation 5d224063-e1c8-4345-b429-4bca1e0e5344 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Stable Reinforcement Learning for Efficient Reasoning Proximal Policy Optimization Algorithms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.731606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.731606Z digest=sha256:69c2134825a07137b284b9e85e596a7281d11f8b22307477fe19058ecb3c72cc

Observation 945cf57e-8403-4d7f-b180-3496e2dbec1a · outbound

This paper cites Group robust preference optimization in reward-free RLHF.

Stable Reinforcement Learning for Efficient Reasoning Group robust preference optimization in reward-free RLHF

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:49.661260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:38:43.873464Z digest=sha256:5b3f363c9e9297eb036e5b2c899fdeff2cc97289aa42ae4deb5766f0ab416a5e

Observation 2de9ee18-9ab1-4ab1-a720-4f5ab0da98dd · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Stable Reinforcement Learning for Efficient Reasoning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.107498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.107498Z digest=sha256:35f516a411272b751a56e96b916cf94c6d6778a6edf40db07e9c7d3d83929e9e

Observation 09bed22b-4186-4cc2-9e0c-4b544f4540bf · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Stable Reinforcement Learning for Efficient Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.256624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.256624Z digest=sha256:f86799f21012fbad0dcc31cee0cd8c9617485f430a5de3961354270d019ec4fc

Observation a05546bc-dc99-4b26-8b01-9c0784e78b85 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Stable Reinforcement Learning for Efficient Reasoning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.377434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.377434Z digest=sha256:6d2d0d59c645fb39033c15099214e376bd0aff89562bcaaca683edebe2455e8f

Observation b91f5d9f-a288-40d2-8523-5518f8dd2466 · outbound

This paper cites The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks.

Stable Reinforcement Learning for Efficient Reasoning The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.636364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.636364Z digest=sha256:1d1805ac2a29e6a7276737e492169dc5919e0f42a4bb1e4f9779e19c74f8c6bd

Observation 1c8b41cd-4bc4-4d55-a372-8f305df9f6cb · outbound

This paper cites When More is Less: Understanding Chain-of-Thought Length in LLMs.

Stable Reinforcement Learning for Efficient Reasoning When More is Less: Understanding Chain-of-Thought Length in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.762699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.762699Z digest=sha256:801a35e87fdcec8ba21a36b4098c884d546732aaec8926479394209e0ca4e27e

Observation 88487cad-08c9-4ba8-96d8-177b0f0632c0 · outbound

This paper cites Dynamic early exit in reasoning models, 2025.

Stable Reinforcement Learning for Efficient Reasoning Dynamic early exit in reasoning models, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.901523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.901523Z digest=sha256:59da119867b88e6b959644e68952aee46054a4350d5798b0616c51d9c0abea5e

Observation e6ebe74d-386b-4f74-9ac0-e03020456fba · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Stable Reinforcement Learning for Efficient Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.030627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.030627Z digest=sha256:df287912df03e2d1931e1170c42ae6345df82a4cdf03f3ffd589a677652312ad

Observation aaaf58df-ce8c-448f-a445-8adae971953e · outbound

This paper cites S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models.

Stable Reinforcement Learning for Efficient Reasoning S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.299309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.299309Z digest=sha256:5c0273a04adcfbeb332a3896a91bc5781c72794d09a7035335c1925ccb3e606d

Observation 2ab6bda3-e684-4120-8286-37f5ceb9ed7b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Stable Reinforcement Learning for Efficient Reasoning Training Verifiers to Solve Math Word Problems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.504178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.504178Z digest=sha256:75725676e81879af4fae04718d1e15632c7ac9efb0318d8026a27f57cd9248fd

Observation 43b2069d-6689-4980-978e-fa6673d4508a · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Stable Reinforcement Learning for Efficient Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.622242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.622242Z digest=sha256:ba719b56e1d89b2d95355a115436a64a328b40ed3e4f546dbd74a69937cb78a9

Observation f63de3ea-d04c-4cd4-ac5d-e3a980e3cb62 · outbound

This paper cites Aime problems and solutions.

Stable Reinforcement Learning for Efficient Reasoning Aime problems and solutions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:49.078178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:38:45.747265Z digest=sha256:9600a8ccbf82301aa601741c1d81eb22c769924d7c0fc35789106c5b5658454f

Observation f898c5e0-3baa-4776-a942-9b7e0c2f03a1 · outbound

This paper cites Amc 2023, 2024.

Stable Reinforcement Learning for Efficient Reasoning Amc 2023, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.879870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.879870Z digest=sha256:e80ac312c522ae22a5daf0273dcb4127d652eecd7284dbece165a9e08e3b4327

Observation af2ae48e-fb79-4463-9aec-7e25a8fbac07 · outbound

This paper cites Measuring mathematical problem solving with the math dataset,.

Stable Reinforcement Learning for Efficient Reasoning Measuring mathematical problem solving with the math dataset,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.047459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.047459Z digest=sha256:fc94be79f216440d995cf516e45833c23359b46c1c41c441d5cc63a73c6a5d0f

Observation 37b165fe-3c43-458e-999e-5afc1c52a00f · outbound

This paper cites Learning to reason with llms.

Stable Reinforcement Learning for Efficient Reasoning Learning to reason with llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:48.564499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:38:46.312282Z digest=sha256:cbfcdd72ffad56b06874deae5693bcdb3710dad2a73655ce1de7eaa37de8b6fd

Observation 71d0c005-ec14-442e-885a-eded273fb183 · outbound

This paper cites Training language models to follow instructions with human feedback.

Stable Reinforcement Learning for Efficient Reasoning Training language models to follow instructions with human feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.432494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.432494Z digest=sha256:9a02348128507816732856d70ea8b18c0e4a4b2d167ce844703cd79d4e3c7c9e

Observation 0da1767a-63d3-4bbe-acac-d4186ae8de5c · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Stable Reinforcement Learning for Efficient Reasoning On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.529177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.529177Z digest=sha256:d1884951131dba93ad25ff4f3fadfa769381ef1b3c57bb689081f2d8f0588c9f

Observation 143a71ca-8286-4c91-9cea-81e9229fdd6f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Stable Reinforcement Learning for Efficient Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.709657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.709657Z digest=sha256:b461bc935a150c5a72394b0b8a511fba438e94a90581a902eb127637d058373a

Observation ff791d8f-0479-4963-a508-1360be495822 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Stable Reinforcement Learning for Efficient Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.831392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.831392Z digest=sha256:0211c6f1c46ac3c70d452e7ad8fd80e4ef374db660c19e4ad1f1e3dcd351d867

Observation 8f6d38db-9d42-4123-b1f9-235377011f6b · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Stable Reinforcement Learning for Efficient Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.973811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.973811Z digest=sha256:5f7633cbeebc3ef10c9de1814f63612abc65f5c67204ed457a6decdc9fb6d412

Observation 13461ca3-1766-4edb-a8b9-c4016710b253 · outbound

This paper cites Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025.

Stable Reinforcement Learning for Efficient Reasoning Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.135822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.135822Z digest=sha256:0fb98602aaeca2564f2419a0462e3044c8ca1a453ad4c1ef1575525aadbbd24d

Observation ca3c6a0b-18fe-4b0e-a2cf-36a820f24d15 · outbound

This paper cites Training language models to reason efficiently.

Stable Reinforcement Learning for Efficient Reasoning Training language models to reason efficiently

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.237133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.237133Z digest=sha256:18f7d774c01fe8f363f652f0d5ef20cde26f272aa3cc14e3cf0e167a54d93081

Observation 616be23c-0f78-40a6-8881-8a637368c55a · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Stable Reinforcement Learning for Efficient Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.403222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.403222Z digest=sha256:0cb43419c28ec53b235ac6d601a65e5430d6091d0a60d4292b34f4a97a925c19

Observation 445767f5-5ddd-4cc7-9d3a-4bb0c17866a8 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Stable Reinforcement Learning for Efficient Reasoning Adam: A Method for Stochastic Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.561040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.561040Z digest=sha256:f5dc8794b2a363108afc857de77a63746f5ef460c04f9bcb15f1dbaae4ac3529

Observation 456d2ebb-02e8-4cfb-b296-56ce53d3e3d9 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Stable Reinforcement Learning for Efficient Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.682695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.682695Z digest=sha256:42782fd60f8aa65b83985eb87c3ab369dc8d752fd5e5ed4d7a64ebeb30fd5cff

Observation ac5f1b6b-6f50-4cac-89de-0faf5272030e · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Stable Reinforcement Learning for Efficient Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.155558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.155558Z digest=sha256:79e0b9654daf4a94f1745ec47a55fb75e80105a783d80319dff8f7fe526844cf

Observation df1278cb-0481-40ff-81f4-9e0813881153 · outbound

This paper cites an unresolved cited work.

Stable Reinforcement Learning for Efficient Reasoning Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:38:49.453422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:38:43.956922Z digest=sha256:d8977a58cd7744f706b48247b3cab6d305dd51cb99995a4797842fd16067e0c2

Pith citing papers

Observation d005cb89-ec4f-4219-824f-e9cb83aa415a · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security Stable Reinforcement Learning for Efficient Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.636062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.636062Z digest=sha256:c354f3948e3142b835d025c4948754e5415f91ec969bc7ace9eccb92c88dc5d3

Observation 83c4f67c-abf3-4632-be6f-9ddc3cb4bbdb · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization Stable Reinforcement Learning for Efficient Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.726774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:47bf382b09d99d5c0a854bd6ce713ef40518d9aceae61cdee25892c2fd388dad

Observation 32eafdbe-ed52-49a4-b269-446d843604d6 · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Stable Reinforcement Learning for Efficient Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.263012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:799d6b1b765ed1e37ffcd3e59d59519271e4d1e790b5de010f7422566e848684

Observation 6967ca88-cf8c-4aac-a953-bb8e16f81ff5 · inbound

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning cites this paper.

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning Stable Reinforcement Learning for Efficient Reasoning

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.253672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:59:57.983338Z digest=sha256:6c5aa2e9f5b023cdd9da0705db643683ba1dcc3f5e3302e553119e5f21a8b5ab

Observation 2163f506-be2c-4659-b578-8de3e92f0899 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:bdec82f037a963d879cb6e1227f7b52005c1b41515a52bdef359de4cc4589907

Observation af84df61-620c-4d7d-b0a9-b19de84984bb · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.779080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.779080Z digest=sha256:082f7dab03e7aa2c43b3d48d2f32f81c1f3502ae4ccbc59c7f8a43054dfcaea8

Observation ef30e1dc-c058-4e94-95c1-46f0948119e0 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.449308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.449308Z digest=sha256:becfa077003eaaf8881e9e1c767dc7df262754a9fcecfd0b2a0b57f11da6681e