Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2407.10490.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:41.745767Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T22:06:16.575794Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation cd2a7a68-530a-49df-9b75-a23b2a061db7 · inbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Learning Dynamics of LLM Finetuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb0aa861-d8ce-4d6a-9d29-a536550589d8 · inbound
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Learning Dynamics of LLM Finetuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1990d6db-a32b-4898-9e51-03d125c24718 · inbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Learning Dynamics of LLM Finetuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3bf01c-0ebf-45bb-ae2d-bdd292bf06fc · inbound
RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs Learning Dynamics of LLM Finetuning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a6839e6-4658-44d7-a60d-2cff956cb3ee · inbound
Decoupling Task-Solving and Output Formatting in LLM Generation Learning Dynamics of LLM Finetuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d78c3bd0-1096-4aa4-a769-25da7a428295 · inbound
What Is Preference Optimization Doing, and Why? Learning Dynamics of LLM Finetuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9d916c5-57f6-4e1b-b752-dce11ca7e33d · inbound
RAD-DPO: Robust Adaptive Denoising Direct Preference Optimization for Generative Retrieval in E-commerce Learning Dynamics of LLM Finetuning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e158d72-527a-4118-b4f8-b930dae0cf75 · inbound
Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation Learning Dynamics of LLM Finetuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc835128-3a56-42a2-86bb-8a2e16499593 · inbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Learning Dynamics of LLM Finetuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f7172db-7504-4c33-a776-885f13b0e1b7 · inbound
Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Learning Dynamics of LLM Finetuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66c2ec75-2ba5-422c-b88f-6e502a0d3042 · inbound
Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Learning Dynamics of LLM Finetuning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 665c7c56-9e73-4710-84ed-43ddd4ecf1cb · inbound
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Learning Dynamics of LLM Finetuning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c72d8732-7680-4b17-a552-4319e9f2cd63 · inbound
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits Learning Dynamics of LLM Finetuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98f421c5-036b-410b-960f-f19f950996bc · inbound
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning Learning Dynamics of LLM Finetuning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f47741bd-d9bb-435e-b246-cee0ddd3d33c · inbound
An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? Learning Dynamics of LLM Finetuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf00b194-6eb2-423c-b23f-bb76c5b8a4bf · inbound
Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion Learning Dynamics of LLM Finetuning
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d174774-00f0-4a18-81b0-4d95a0f5b50f · inbound
Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs Learning Dynamics of LLM Finetuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.