Pith. sign in

Paper Citation Record · LEDGER

T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2501.11651.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.11651 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:08:21.609464Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T22:25:39.474855Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 81251f2c-dffb-488b-8f22-598a5490d6d0 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 269

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.589761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:23ca499e75b7cc06ad91be4f7e5f1177f491a880572cb7e77d1b9d88675b9cf0

Observation 3abbd6aa-3f04-4587-8587-14516d4153ce · inbound

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search cites this paper.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.609464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.609464Z digest=sha256:0abcff25b04b3beed4384c632bf856e26930393990527af28656626456f42930

Observation ecd65069-43f6-4e93-b30d-fac8feff78d3 · inbound

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems cites this paper.

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:32:36.538291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:32:36.538291Z digest=sha256:af1139efc0fac6913d566f151b8a37b89945120fd4724c7ba5aaceb66cdea267

Observation c480f1c5-9c68-4f25-9b8e-77c4d3e525d8 · inbound

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories cites this paper.

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.864750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:13:12.864750Z digest=sha256:244d25a84183ee7984a79b5ae887495169cbe8720d7e7277b09b63119cbd3e5f

Observation ac01cefd-0396-434f-892a-51412916c5c1 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.806080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.806080Z digest=sha256:c058085fb1e996d2e5c01f19647e755e9c021ce2e3916b9b90600c91c90e36bc

Observation c16c6a00-d1fe-4f2f-a93f-369528e211ca · inbound

Reinforced Language Models for Sequential Decision Making cites this paper.

Reinforced Language Models for Sequential Decision Making T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:54.083020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:54.083020Z digest=sha256:960dd7f36f51fd06c43e27333d17351d50f2b78426b93e89394c0d08eacc31be

Observation a32a9fd8-f044-4e31-8089-d72f7c94858b · inbound

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems cites this paper.

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:50:10.489926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:50:10.489926Z digest=sha256:8a7a224a4863fbf05261b49d694ceddf437b7b09d75846b4d82953f7450c0f64

Observation a12d51e4-c978-4d5f-8b2b-c7e22818bcdb · inbound

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation cites this paper.

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:39:54.390558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:39:53.995700Z digest=sha256:ea8cb175c67c8213069ff77873fd5d8292c4829e1117c218758e9d8fa1691f5e

Observation d89408a0-8077-47b2-95f1-fd431cf9722b · inbound

EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget cites this paper.

EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:56:08.697739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T08:53:31.803096Z digest=sha256:1348562753ee0c3b429b76ffcb693b8e9a68655fc77bc5d5cceac45ee12d5205

Observation 824f44cf-2477-4107-bf83-d736628ddf0d · inbound

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training cites this paper.

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:30:35.407585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:30:17.581842Z digest=sha256:75c95bf68d90333504d620153e26d06fe71e576e71b3d84230ebe969398c5e22

Observation fbac5259-3219-4704-9dae-3d7634003ab4 · inbound

Beyond the Sampled Token: Preserving Candidate Support in RLVR cites this paper.

Beyond the Sampled Token: Preserving Candidate Support in RLVR T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:07.194227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:34:07.194227Z digest=sha256:c9510e88d621aaacefb68f1bfd71d4434321b067aea129289625927a9484049d

Observation caaa32b5-751e-44d1-b34d-a7d5220f2f0f · inbound

When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs cites this paper.

When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:11:51.451496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:11:51.451496Z digest=sha256:b6bf24a544663ce94446ab524c9e3b70853b12dbc79eaadff0a6715448a08e0a

Observation 68f8d439-e2b5-45f7-a112-384252096432 · inbound

Targeted Exploration via Unified Entropy Control for Reinforcement Learning cites this paper.

Targeted Exploration via Unified Entropy Control for Reinforcement Learning T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:49:56.071198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:48:27.733823Z digest=sha256:f2bf9899983bd0dbd3ac1c39f29b51cda112191b6bdaf252d840fd33dedfe9e2

Observation f39bcd0d-440c-402f-a67a-3420cc2047bf · inbound

Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning cites this paper.

Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:01.943069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:40:25.414166Z digest=sha256:833cd26d8c09159322ce3af449e6828e13ae97856d06cc4ceeacf32668ede1f9

Observation 92b1c155-540f-4788-9ef3-23d07d1f64ca · inbound

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning cites this paper.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:07:09.106701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:57:30.877915Z digest=sha256:7ee219286668e9bae9833016c43e3d90c92e89b6af45297532964c0daa8de9b3

Observation b0e72bd3-cb95-465d-8f94-3ec23d0f8b01 · inbound

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning cites this paper.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.497847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:3b51c8546fd2c789385b738a475088d892274343b099c0ac6d1283bfe6ae48db

Observation 2d79e701-e8c6-4cc7-ba42-410b598735f4 · inbound

SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs cites this paper.

SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:44.141720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T20:09:42.841695Z digest=sha256:4a0f29743a24c0e1d7aa9b3e3ec539b18d462d2838e805856eac59889bde5e33

Observation 212e8307-e81a-479c-a1aa-106cd0950245 · inbound

Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization cites this paper.

Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:25:39.476832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-08T22:25:35.119625Z digest=sha256:4dd2e3cdb85b5d307b690d2be7490cf1f8efa5d9ddafc99f9b894ba06b069059