Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:21:25.280778Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 5 inbound Pith citation observations for arXiv:2508.16546.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:21:25.280778Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T18:49:47.876179Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T21:08:57.813685Z
2 of 2 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bf3bf01c-0ebf-45bb-ae2d-bdd292bf06fc · outbound
RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs Learning Dynamics of LLM Finetuning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b08014-c2c5-4114-9a8a-66e19e7ab620 · outbound
RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd11310-06fa-447c-b220-ffbeb8dc6875 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
Reference 247
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e48779b-a561-4402-84fd-9a89f9097237 · inbound
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3fe11734-8d35-4637-8da6-aaad2aaf8d73 · inbound
Visual Reasoning through Tool-supervised Reinforcement Learning RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd565c28-485f-4cb0-9e7a-9263ae2c2ae4 · inbound
When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26b368a4-02fb-414a-8bd5-4864e047040c · inbound
Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.