Pith. sign in

Paper Citation Record · LEDGER

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

As of 7 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 10 inbound Pith citation observations for arXiv:2508.00222.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00222 v5

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T01:15:49.658412Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:34:06.659805Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T22:27:26.448121Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact19
  • verified fuzzy7
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7b22eef4-03c0-4df3-93c4-657d5cecb0a0 · outbound

This paper cites How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:16:56.985752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:f254f8839f0b6354d875ca8f7c789b485165c8c44e913f454aa701f30f4f388e

Observation 6ae68cf3-9cc2-46be-b62b-fa959f8e26cb · outbound

This paper cites Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:16:56.993587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:c155801e2ac7ebd42ea07653fb3bee75e9dafcb57c45f9bc412c731dbfa7fced

Observation b96672c0-0bed-4193-9ae3-4c8ab85aac4f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.060549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:883f35d72f952556a4b477716746beddd6c69b84ab8d0eb1b24f99b97a8a5b41

Observation 1f0cd3ad-c5a0-4f7f-b891-ee41de0ec916 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:56.979210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:35ac202ad399fbb3f10abd4e17747fc2b4dda4a722894209013e9e1660e118cb

Observation c02c08fd-8d9a-4359-bab6-f00a2eeacb8f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:56.999752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:2f6d33b2de7ab3458cfa9372329dcc055a86bba0e7896b53bfb3012ee359bca1

Observation d341d843-23ac-4668-85cc-2a4cd35da58b · outbound

This paper cites Process Reinforcement through Implicit Rewards.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Process Reinforcement through Implicit Rewards

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.015996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:58347ff4c9e7028ac85c0680a5eeeb26b5d13f2529d3a8d500ab4c6ad3350b02

Observation d9588cdb-293f-4ccf-bcee-e4773f677bf4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.007208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:0b45f9c1ac01fd1cf2b80218ecf4f0eba72935e4bf933eac898f23958107bbb3

Observation 4cccdf65-b86f-4eaa-8584-d24c5534aadb · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Teaching Large Language Models to Reason with Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:16:57.097300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:380755bc33a094e213ebc51c01a11ca5c5af0b8d5d4d1a707a75b7cf6d8d3b14

Observation a05552eb-1069-4963-ae17-3dd4728d2867 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.023034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:00d32be630329d7b6c94096bea614c68ccf18457318cf5695c5b62a4dcb23bd3

Observation d23aded9-6e25-426d-8887-42be8eb3e84e · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T01:16:57.066799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:a5f4174674a95c76102c212bfa04c0ad7533fead27c737fb59e72ef0e9562365

Observation e437c7e8-593e-4319-abf4-4ed102200f61 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.045584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:6e5835614e08b106bd0bdf20acf0837418ac9e352e99ced1c8bbe9e451db1816

Observation 90be6e53-f651-415a-b8a1-e2b32ae7e8e1 · outbound

This paper cites SuperRL: Reinforcement Learning with Supervision to Boost Language Model Reasoning.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization SuperRL: Reinforcement Learning with Supervision to Boost Language Model Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:16:56.964027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:a21fc1a4476dcebd2a3d8c0a70e5539a8988687ee2846f6dd647468bbba08c46

Observation 22272cb6-e500-487e-8f2a-9bfa08c6b452 · outbound

This paper cites John Wiley & Sons.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization John Wiley & Sons

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T01:16:57.416280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:601421e14736e8ae596919e0fe4c0db5ebe8e96c9152e70a94c591206232c26e

Observation fab2b4f4-8975-4504-98b3-c6df7177afda · outbound

This paper cites Proximal Policy Optimization Algorithms.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Proximal Policy Optimization Algorithms

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.039204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:854b0ff95f3b13037632993fc3aedab2bae1463ebe62e099444167a3c19b29ff

Observation 3614caa9-715a-41bb-9e26-b7ea0abf370f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.032605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:554654669932bc5a5027bdeec7f4a8170afd96bdee100ad0912a479ebe85c8e6

Observation c62b401a-4ddc-4b1c-8099-2e529cba08b3 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization HybridFlow: A Flexible and Efficient RLHF Framework

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.078802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:7401eba8240b477f8ec8eb85eb709f025dbf7c057b4f1f0611e7b534ea7907d0

Observation 25803ded-892f-4eb0-8dfd-08c9acb5dbfa · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.104689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:8b53ee4813ee4411e96f827aebec0206de1f013391738631da714506342a0b19

Observation 7e1e3c82-2fa3-4b9b-8617-47f0d56dc2e1 · outbound

This paper cites UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T01:16:57.084581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:d6ac7b056e936d0a2bb9e2d4c87f8bb1553361aabea5c05378cfbceae2ef55f3

Observation 7132f0ea-f4be-46f5-b269-58b3540dbda3 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Learning to Reason under Off-Policy Guidance

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.090142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:5fe7ecf7cebe044cdbd4c075c8070e817ffc32453cff45fc4bd20061c5eea908

Observation 022e8f73-c2b8-4372-957f-d47abeb89bf0 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.073079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:dcca7d50352fe1fef94f7d396ec94fe8683c7a15cafcf630cd1d6c5f61d3d300

Observation 179d49d9-5143-4989-a9c2-f99e8d816395 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.054416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:2c74abbe37b507d2385ac4ecf93fea47a43b7c53f0b584058196f5fcb5bff739

Observation 1690ad0a-2a23-4c81-9284-9c8fccd9f666 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:56.971782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:1ea0e634983a8efb7f0ff99a57701b569e950c2563ebbea73876ba0b99bbb535

Observation 724ed7d1-bae9-44e3-b0da-9ea529315ae7 · outbound

This paper cites First, we dissect the bias and variance issues inherent to standard Importance Sampling (IS) when using data from a single behavior policy.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization First, we dissect the bias and variance issues inherent to standard Importance Sampling (IS) when using data from a single behavior policy

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T01:16:57.405776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:7bc7800cd674f8e74566b617f13faef436709d638ad7c153a2bb6bf5df5cb5be

Observation b3195e19-88f0-4e10-a4bc-89576fb6853d · outbound

This paper cites Both theχ 2-divergence and the more commonly known KL-divergence (DKL(πθ∥πω)) are measures of dissimilarity between distributions (both are instances of f-divergences).

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Both theχ 2-divergence and the more commonly known KL-divergence (DKL(πθ∥πω)) are measures of dissimilarity between distributions (both are instances of f-divergences)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T01:16:57.419753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:87aaae9ab26088bbf3055388a1bbba647d6f76a0303a48a11eb2f0b9774343a6

Observation e7d5442c-d63d-492c-86aa-29264ba58b21 · outbound

This paper cites safety net.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization safety net

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T01:16:57.402310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:aa99c3a4cd3474e7f52b00e928545221172ac2a98b2034718bd7920b9a3076cf

Observation 1de1f09c-09e0-4355-983e-fcdd4ee855d9 · outbound

This paper cites an unresolved cited work.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-19T01:16:57.399065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:dca2330360c6b85f3e568c3e44483de8454677a73d3c28ae7573c18f86c113f6

Observation b5dc6cc1-442a-4442-89f9-ab6ee5d81ee2 · outbound

This paper cites For our approach, one of the model-generated rollouts is replaced with a correct reasoning trajectory from the training dataset.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization For our approach, one of the model-generated rollouts is replaced with a correct reasoning trajectory from the training dataset

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T01:16:57.392025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:d882295f6053b3af8f134cb271e07a9866a1f2f56ece8240ef255227169683c7

Observation 44f5a834-1b23-4e05-842f-ab92b39fa250 · outbound

This paper cites an unresolved cited work.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-19T01:16:57.413093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:f194108a0f68157adc258b3d2655802dbacb4b1c1a3848c4cdac8be8ff41ca9d

Observation 732bfdc5-a802-4494-8539-6bbe1f4e0228 · outbound

This paper cites During evaluation, we set the sampling temperature to 0.6 and report the average pass@1 score over 5 runs by default.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization During evaluation, we set the sampling temperature to 0.6 and report the average pass@1 score over 5 runs by default

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T01:16:57.409654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:c6e2bea4155535a9aa0231a0cb5ca1cc4b06dbf633961d4e42ae636b7df13d52

Observation 609a9226-2fbf-49ab-a86f-b939df71e029 · outbound

This paper cites The second category consists of four straightforward baselines: 1)SFT, supervised fine-tuning using external reasoning trajectory data.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization The second category consists of four straightforward baselines: 1)SFT, supervised fine-tuning using external reasoning trajectory data

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T01:16:57.395871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:4cbd84b1c8dda804bec2e0db4e8486e8d840e0ebfd06ddad47a16b84aaa9146b

Pith citing papers

Observation f544093f-ec86-4b34-af8a-850645821afe · inbound

Beyond the Sampled Token: Preserving Candidate Support in RLVR cites this paper.

Beyond the Sampled Token: Preserving Candidate Support in RLVR RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:06.659805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:34:06.659805Z digest=sha256:40425ce7e0faae56b7a8596ff8d7c1d7e1aa066b832c41c56f591cc27b36960e

Observation da083899-7760-4574-bf2b-59638c3afc2a · inbound

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models cites this paper.

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T05:28:11.383416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:28:11.383416Z digest=sha256:cacfcf281d687ea5e7ce32d56bb331bbabed9af77c6a6efde85a6efacaa52191

Observation 7f852c4b-d9b7-4d86-b2d2-d3e785e29c46 · inbound

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy cites this paper.

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:53:11.743641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T19:51:08.304741Z digest=sha256:0c7a0843aa9ce2685320736646619d9689fe5a59d0d52fc9c51b62547e7658f6

Observation 895f2d4e-3db9-4fed-a88e-6bfe5a51724f · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:30.040182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:39fc1ee233e7ce2ab657f2a651191901848809afefdc3fc7ae88fb128ee0d47a

Observation 9b2c95e0-46d3-4174-9cd4-f5bdfe9d76c1 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:11:16.745580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:ffc30d247ba280e5f8c7bacb2dc03c3e29c7d07758279fcef5ed16386790b4fb

Observation 54793fe2-ed12-4d07-9310-e806fa66d86f · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:02:40.643093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:576c97013477b215238943e4cadf816aa4fac63b31941516d45719e27454925c

Observation 988a5c69-d1bf-4da7-a7e1-f9ce65807ce5 · inbound

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR cites this paper.

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:10.539031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T15:39:54.001327Z digest=sha256:705dab493eb6f945f0c57a1d7c8dd422540a3fe1c349f77f0ff7ec29fa501ea7

Observation 026b34fc-4a97-43da-bb62-a9a4b3d59496 · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:06:41.027689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:3d2be40876b23decd48730f42be0f1811160f33ce1f07ed48d88021888e7dce2

Observation 4ee16055-4ec2-4e07-a709-aaf498fc08b5 · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:27:26.449420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:bfd637ffe3ffee5958f61a387d5114f996f4ca9d9ff3f804ded198db4ffa40b6

Observation 637b06e1-7920-48cd-8901-029268f58fa7 · inbound

Preference Tuning as Spectral Update Reorganization cites this paper.

Preference Tuning as Spectral Update Reorganization RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T14:23:19.627366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:23:19.627366Z digest=sha256:23940a49a389f76d9027187f7da5a6de46ca7537c8bb93ca9cce990dd23ec831