Pith. sign in

Paper Citation Record · LEDGER

RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

As of 7 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 5 inbound Pith citation observations for arXiv:2508.16546.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16546 v1

Coverage vector

measured 2 of 2 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:21:25.280778Z

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:49:47.876179Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:08:57.813685Z

Reference resolution

2 of 2 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf3bf01c-0ebf-45bb-ae2d-bdd292bf06fc · outbound

This paper cites Learning Dynamics of LLM Finetuning.

RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs Learning Dynamics of LLM Finetuning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:21:25.280778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:21:25.280778Z digest=sha256:b95746b5050eea565b45d176867b2c0d6007a8d525c09b367260eed6fc5fed88

Observation 25b08014-c2c5-4114-9a8a-66e19e7ab620 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T17:21:25.215188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:21:25.215188Z digest=sha256:07266bf474634a6a70c886591cf0c46e3b8e0f5f331c2708688f8f340677a71f

Pith citing papers

Observation 2bd11310-06fa-447c-b220-ffbeb8dc6875 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 247

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.606031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:fcaea5a7efa762e5df3d7dc182fa079535c118505d09d96f801f1b46804e252f

Observation 8e48779b-a561-4402-84fd-9a89f9097237 · inbound

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability cites this paper.

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.629820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:53:22.553430Z digest=sha256:2f49ca5461f6921ee844c79037d5dcded892ba3b92938b40d93cb2a539021229

Observation 3fe11734-8d35-4637-8da6-aaad2aaf8d73 · inbound

Visual Reasoning through Tool-supervised Reinforcement Learning cites this paper.

Visual Reasoning through Tool-supervised Reinforcement Learning RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:03.023471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:05:21.688216Z digest=sha256:8e6fdf2488dceab06e61baf288294c3fb45208ed59fd4b203995fa1747891399

Observation bd565c28-485f-4cb0-9e7a-9263ae2c2ae4 · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:27:26.418219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:fedd4c25ae7c6b7495e8c3afc88db263ee9c84b4150361fd851047f5b2c1711f

Observation 26b368a4-02fb-414a-8bd5-4864e047040c · inbound

Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging cites this paper.

Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:57.816892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T00:57:02.938997Z digest=sha256:7c9ba4b4dee888dda47c48f1e2a4b83016543040a23de1222eaaa6caef3dff8a