Pith. sign in

Paper Citation Record · LEDGER

What are Key Factors for Updates in RL for LLM Reasoning?

As of 6 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2606.22570.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.22570 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T10:24:53.245739Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch24

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65b611c5-969b-4456-a04e-a1a4daa120da · outbound

This paper cites Scaling Learning Algorithms Towards.

What are Key Factors for Updates in RL for LLM Reasoning? Scaling Learning Algorithms Towards

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:cc78f5c5196b7c8b88940c3fd1113f32674be1432f417ef5d823a77b379e2d78

Observation e281e20f-6585-427b-90fb-5fba933440e2 · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

What are Key Factors for Updates in RL for LLM Reasoning? and Osindero, Simon and Teh, Yee Whye , journal =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:d1f706b7662bd79d9c6b8697673b8981ebe731924985cae4b33293580c4ec45f

Observation 6dab33e7-8946-4901-9c35-4c4fe40c962a · outbound

This paper cites 2016 , publisher=.

What are Key Factors for Updates in RL for LLM Reasoning? 2016 , publisher=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:c13279c363c9265b3452812c7650822271670d773281ecc2ca1bf379a6307d0d

Observation 5ddd0cda-a0b2-4228-af60-eb7cfa26cef5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

What are Key Factors for Updates in RL for LLM Reasoning? Proximal Policy Optimization Algorithms

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.606398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:4fcc02e9b01d0900954091759c459025c587a77fba055b1e25c050af46e65c57

Observation abfbcc2c-8d86-446d-ad8f-3243bd3a4b89 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

What are Key Factors for Updates in RL for LLM Reasoning? DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.601370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:cae1df5e0b90e1912853537a9fb78c003b063185922667b1a559dbdef5921298

Observation d5688a91-f619-4749-a915-c9df7ffba541 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

What are Key Factors for Updates in RL for LLM Reasoning? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.598925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:096c92c6bb25a4880d99369aa9b8631939cc5608428f3c83502eddafe5070f8c

Observation c5527d84-9d29-4b75-9797-8f0b4cf0a6d0 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

What are Key Factors for Updates in RL for LLM Reasoning? VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.603828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:21a75e4d653003d9cfd45869626acf131e26eb6bd7cf97a17883f67fc86d3aa1

Observation f64637eb-4856-4e27-9f83-a6a2bc0cf1e7 · outbound

This paper cites Advances in neural information processing systems , volume=.

What are Key Factors for Updates in RL for LLM Reasoning? Advances in neural information processing systems , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:1b723dbcb8868db9c4905c1d8782535937d2a7d11c506763ec0a344b7a9d1cba

Observation 94281c76-f5a3-49bb-b08d-d9e94d32efe9 · outbound

This paper cites 2025 , eprint=.

What are Key Factors for Updates in RL for LLM Reasoning? 2025 , eprint=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:b7a2b934e8948de29b9e37897bcb8c2570bc4a0ea972b4f9ea3463da718ad8f7

Observation 67eada85-28c2-4a11-9974-de08626c402f · outbound

This paper cites doi: 10.18653/v1/2022.acl-long.78.

What are Key Factors for Updates in RL for LLM Reasoning? doi: 10.18653/v1/2022.acl-long.78

Reference 10

Resolution
metadata mismatch
doi, observed 2026-06-26T10:29:18.566303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:aa60af45788d652a4928469e0692c5ccb531a6774d2f57a9f902564b6b05839a

Observation 7d6175e9-6795-4dc5-b415-5933968ce487 · outbound

This paper cites 2025 , note =.

What are Key Factors for Updates in RL for LLM Reasoning? 2025 , note =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:aef2ac9f2d42da6384e4835b45186b688d6e80e5e641ee43b5f6b9650ad29196

Observation 3876e9d5-9ae3-4dfe-a9ef-e42e59fe95e4 · outbound

This paper cites 2024 , journal =.

What are Key Factors for Updates in RL for LLM Reasoning? 2024 , journal =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:7d5f63e67d06e544a02f4db4c72487963da42d77b7f8454f0d17c07bf0461224

Observation e8e3d195-0951-4d2f-a6f7-f69c6920f5c1 · outbound

This paper cites 2021 , eprint=.

What are Key Factors for Updates in RL for LLM Reasoning? 2021 , eprint=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:a586a54a8efdd7c79887be4c1a777d3aee622b6e9a90f7dbd75394ec008e5f6b

Observation ffc53ab1-a6cf-4127-8cb4-cfb9803b5d22 · outbound

This paper cites 2024 , eprint=.

What are Key Factors for Updates in RL for LLM Reasoning? 2024 , eprint=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:d9180b5e277332e498e0b3825fce0bfaef5d240cdd268b2eae2f97e6c5f9f3d6

Observation 5baf5c3b-b621-45ea-b8c5-cad94d3c6868 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

What are Key Factors for Updates in RL for LLM Reasoning? Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.596474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:97631d7fa284a34fcadabc42de0589852bc35f75581863b35a4494c46fd036b0

Observation 2a207b72-137d-44d8-a353-f43b8e1a264a · outbound

This paper cites OpenAI o1 System Card.

What are Key Factors for Updates in RL for LLM Reasoning? OpenAI o1 System Card

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.594060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:47de0c3d57af2ff080f9af81718f8633f2a4212874128503b1ed2e9789f4ced6

Observation 78689bf4-fd35-4f36-9e1e-660ffcdbe1b3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

What are Key Factors for Updates in RL for LLM Reasoning? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.589102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:7504dc93ff80f8e314eaec05c6c5a2e619a59e56e546cb924dbeabf3dee79a0b

Observation ccb7237e-c26c-4967-87d8-688763bea325 · outbound

This paper cites 2025 , month =.

What are Key Factors for Updates in RL for LLM Reasoning? 2025 , month =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:fba95141f6852a61abaa2c9ba5d29f4b6b01c7d40eb0a6efb22e6af613401ec8

Observation 4afb7a0a-b48a-4c0a-ab97-a7d8e16d28ed · outbound

This paper cites AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning.

What are Key Factors for Updates in RL for LLM Reasoning? AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.581253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:452925b52346279270157dc94de6ef5236ee69b6dd81e5ef66e9779393795efa

Observation 1b693f92-b6ab-4bdb-b1a9-9c23ed2fd224 · outbound

This paper cites Qwen3 Technical Report.

What are Key Factors for Updates in RL for LLM Reasoning? Qwen3 Technical Report

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.605643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:c94b5d1543082c9916027751ba751c8ae0ea0f679ed390bed1ca8edf4e15aa10

Observation 2ef7407c-c595-4162-b127-8a0e10eadcf3 · outbound

This paper cites Qwen2.5 Technical Report.

What are Key Factors for Updates in RL for LLM Reasoning? Qwen2.5 Technical Report

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.586643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:9c54393d34e60d3df4a3b585f229490a0870496b65b530eca83484245a01d082

Observation bde99d6b-5f49-4a85-a388-c870a2042597 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

What are Key Factors for Updates in RL for LLM Reasoning? Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.591750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:5792ca8bdf1d16e3153361e210fe7b32ce073688a7d30a275a8656c6b3306e7d

Observation af7bf6e4-7168-4bd3-8cb3-3bc0f4d082bd · outbound

This paper cites Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs.

What are Key Factors for Updates in RL for LLM Reasoning? Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.578480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:3c3b0165c0441a8130ab7882bb8dd2c62a4e042669a70991a253e47a3461ae61

Observation 53c442e5-1ab9-45a7-a1ac-a0118e30868b · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

What are Key Factors for Updates in RL for LLM Reasoning? Reasoning with Exploration: An Entropy Perspective

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.563501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:21a00ab8bd06129e4bbf88fb3d5a5dc525568ba5cd6920aaab1721ee63b5208a

Observation 91c6b4c4-7983-4953-b57b-0c2be8bfef38 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

What are Key Factors for Updates in RL for LLM Reasoning? The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.569466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:2305ece616b18f9e6a6ca19012702355c32321c55124067cdbdba6175c501fd8

Observation 2b142ae2-05d1-4a9e-bfb1-53ef1010a7c4 · outbound

This paper cites One-shot Entropy Minimization.

What are Key Factors for Updates in RL for LLM Reasoning? One-shot Entropy Minimization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.565648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:43e9f045388caf7e8e31160c987eb7d2f531fe262b38488a6cdc0e4de4a3f156

Observation 14275d2a-83f5-485b-b2f0-ea9c0e1c32d6 · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

What are Key Factors for Updates in RL for LLM Reasoning? SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.560962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:48fc8f5fcea9dd70d9680f244618d5fcc201f2470658fb9d07dd8aec5095e1fb

Observation 3a096dc8-2bbe-4dce-b337-98b70e7851a6 · outbound

This paper cites Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models.

What are Key Factors for Updates in RL for LLM Reasoning? Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.531992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:fb28d35782eb94bfe932834b80bb7060aafde93e7e85b7b264ddb1eef7760a54

Observation 855d224b-6a44-4a97-ae37-fd7df83fc914 · outbound

This paper cites First Return, Entropy-Eliciting Explore.

What are Key Factors for Updates in RL for LLM Reasoning? First Return, Entropy-Eliciting Explore

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.540812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:09f8bc85408722a0150052e9c0f92a6a5c5c263a8aa81b29d5d2acef217fff12

Observation f39d511a-da72-4122-b12f-9f83743d30b5 · outbound

This paper cites Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs.

What are Key Factors for Updates in RL for LLM Reasoning? Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.534766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:968be9b306b10eaa479426fa37aa8f556a431c3d2d015e393217d23a9d4907a0

Observation dad12c66-8e34-44b0-808b-90fa3ac67f1a · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

What are Key Factors for Updates in RL for LLM Reasoning? MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.527440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:b2492cc4322e489967d45c5a06d8e525ab2f2ee17026a1810d9e0b96a9bfe4eb

Observation 6e09a0a3-3f57-4e4a-8eab-93ad19dd7e0b · outbound

This paper cites Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629.

What are Key Factors for Updates in RL for LLM Reasoning? Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.533016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:a59c9b035bf997583520c462fa10800da41c3b963e27ea4b764719fd99a4403e

Observation ba3141db-5b46-46ac-8b82-e2539b502163 · outbound

This paper cites Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR.

What are Key Factors for Updates in RL for LLM Reasoning? Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.600636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:beab39ef80052977d672c1695de831b8b5d0f99a872e0d56cba7777fe86b12b2

Observation 5f26c37a-d901-47b6-8880-98d8220ec7c4 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

What are Key Factors for Updates in RL for LLM Reasoning? The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.595453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:644d80ef8192d46ce44c65b7161edd4475bd533292584038fdd6242d4e3debb3

Observation 7d607ce7-cda3-4fb0-9d51-f107ba4b1af5 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

What are Key Factors for Updates in RL for LLM Reasoning? A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.575477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:fbcc4d72c470b0d94f9c0ff1b909ff5496eee9af0c02469fa88bde054aa929c3

Observation bc31eff6-7769-4562-ad98-fd5267be0c35 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

What are Key Factors for Updates in RL for LLM Reasoning? Learning to Reason under Off-Policy Guidance

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.587637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:a331cffefce32388515167b9e6581e5b41bccb5206ac330feca37eca38264e9d

Observation 0065986c-10e2-4a9b-aeb9-f431a9b40910 · outbound

This paper cites Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions.

What are Key Factors for Updates in RL for LLM Reasoning? Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.603102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:95e6e13527e771bd1eea217e872ef31ffa583d449bc74944a9782b4c8937b147

Observation 967da951-9120-4295-8a7c-68731a45b4d7 · outbound

This paper cites SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning.

What are Key Factors for Updates in RL for LLM Reasoning? SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.590285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:154ab37fea1b10e182a10a047a2d50a0b6cfaaaca6d01d5ef05cf802baa1e53c

Observation 55e747a9-cc43-4b9b-89f2-d32d4ae255f3 · outbound

This paper cites Arnal, G.

What are Key Factors for Updates in RL for LLM Reasoning? Arnal, G

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.592878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:38ef272d2deed0b54c341ef6ae1ecd0043d245037248256ac2d5559d8f882449

Observation 88f54f18-df8a-458b-8cdb-cfc12a3284ed · outbound

This paper cites Part i: Tricks or traps? a deep dive into rl for llm reasoning.

What are Key Factors for Updates in RL for LLM Reasoning? Part i: Tricks or traps? a deep dive into rl for llm reasoning

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.598162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:405587026f9294cd769bc374416cdb1f65777d2bafc8c729ced86a0401428534

Observation 1f61da7e-d60a-46ae-9b91-ecf5b51a6037 · outbound

This paper cites June , volume=.

What are Key Factors for Updates in RL for LLM Reasoning? June , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:62c57a004a99160e79a696b864aba4f206d7e1324786e2f0d09e0bff1ecb9434

Observation aa319e90-2330-4c63-b119-7e5b25f98221 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

What are Key Factors for Updates in RL for LLM Reasoning? Measuring Mathematical Problem Solving With the MATH Dataset

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.547401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:3e80b76c729f0a1a9f395dd0edede0df402e7b733a2ead8f2cf65885daffe8e4

Observation c620a0c4-3449-4494-aefc-ff0d418fb689 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What are Key Factors for Updates in RL for LLM Reasoning? Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T10:24:53.245739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:8c29414bf1bf731d16f7589097b80e55558805c4e61f21e04713d3d0469343e2

Observation 4640e1dc-0d32-4345-816d-47db5b224200 · outbound

This paper cites arXiv preprint arXiv:2504.02546 , year=.

What are Key Factors for Updates in RL for LLM Reasoning? arXiv preprint arXiv:2504.02546 , year=

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.574548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:1aff80e4e427f35f3901d6bccd706f97d16bb1a4081817cbe15d49a91d5b61b2

Observation c7137edb-6139-4332-bbb6-caed3ca98a75 · outbound

This paper cites Group Sequence Policy Optimization.

What are Key Factors for Updates in RL for LLM Reasoning? Group Sequence Policy Optimization

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:09:43.580238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:75f09968459153362338e7e632356adb23ac1b2b47cbbc375e1d9fa07450d62e

Pith citing papers

No inbound Pith citation observations are available.