Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:54:49.897621Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2507.00018.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:54:49.897621Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-05T02:27:50.369374Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
43 of 43 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 155ecc09-68b1-49d9-a69d-74bd2cdbe62a · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63bbd42e-8d4b-418d-b009-286bb6ee8bcb · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections The Llama 3 Herd of Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45844eb2-a2b8-44ab-97ff-08b651d4a23b · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Qwen2 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740526b3-24a8-467b-8455-7f1896fb89aa · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections DeepSeek-V3 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 299141e7-ae50-4248-b9d6-b70024ebe1a7 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Mistral 7B
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4d41498-8e2e-4cd1-a861-aabeb3829e17 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections LIMO: Less is More for Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f19c3e-faec-4b6a-9a17-ba0818a03dd3 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections LIMA: less is more for alignment
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2527488a-32c2-45bc-af58-26484812d526 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 696e9905-ffea-42e7-b833-fc1a31717ae1 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e284614f-2074-41cb-8341-709bb668a09d · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56f6fa23-9bfa-41a4-abf7-b522dacc86ce · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Deep reinforcement learning from human preferences
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba82bcc-c98a-4bea-91d1-b7bba81c268a · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9a44f64-bdb9-4b72-ab81-bbc7d787c66e · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Self-play fine-tuning converts weak language models to strong language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1990d6db-a32b-4898-9e51-03d125c24718 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Learning Dynamics of LLM Finetuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629305d1-60bf-4403-8943-45ff85560ba6 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections IQ-Learn: Inverse soft-Q Learning for Imitation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07a74808-a6d1-4cb0-b197-fc1a7c6aabf8 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Huang, Artem Sokolov, Matt Barnes, Guillaume Desjardins, Alex Bewley, Sarah Bechtle, Jost Tobias Springenberg, Nikola Momchev, Olivier Bachem, Matthieu Geist, and Martin A
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cad91d96-f262-4d7b-abaf-21967381882a · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c44ace5-62ea-4b09-95ad-04272b92cf4b · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Ng and Stuart Russell
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7489688-7b8b-40ac-a258-ad77e3a60e29 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Preserving Diversity in Supervised Fine-Tuning of Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c96eb0c2-448a-4c4c-99d0-6d47d8a457e5 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections f-GAN: Training Generative Neural Samplers using Variational Divergence Minimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf414451-3f34-4e9e-8547-0d2c55c820d0 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Sequencematch: Imitation learning for autoregressive sequence modelling with backtracking
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a6ec825-10bf-418a-9762-31c5031510ba · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Proximal Policy Optimization Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b37f6627-9498-47ce-bc37-d6434e484211 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections A general theoretical paradigm to understand learning from human preferences
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44bf425c-9e56-45d3-aa8c-48e512fbb21c · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Model alignment as prospect theoretic optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b3bfa17-fbe4-42c6-b82a-29bd522508a8 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Simpo: Simple preference optimization with a reference-free reward
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e98b2b3a-7252-4388-a3a0-77c37cb393d9 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Disentangling length from quality in direct preference optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f97dd5a8-7a18-4ff5-b2f4-fbb033ad99b5 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d29a645-ca11-4196-986c-51f102089c7f · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Andrew Bagnell
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bd585c5-1ff1-46d0-a33a-e316b98163eb · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Ziebart, Andrew L
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5944051-dabb-4ec1-b356-adc6efc21089 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Enhancing chat language models by scaling high-quality instructional conversations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4da6cf93-416a-4c3c-bcfd-37ed4e76961e · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cec943e-8c1a-44a2-95f2-7772676d4f1f · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ecdcd53-2441-4def-a799-4c63dd45c8d1 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ac1d6ed-c92c-4bcb-b685-825f7291cc1d · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Xing, Hao Zhang, Joseph E
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebdac33f-13bf-48bd-87bb-d78afd04c21d · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Gonzalez, Hao Zhang, and Ion Stoica
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4f260b2-1c9f-4368-8840-efb07f890521 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections RRHF: rank responses to align language models with human feedback
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 665991a4-8154-4b0a-a47d-ecd103f3e2eb · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00328f4-f5c6-4374-9e20-16d2f547f190 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Contrastive preference optimization: Pushing the bound- aries of LLM performance in machine translation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b364a22-3f24-487d-bcf7-1c013b6c60f4 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections ORPO: monolithic preference optimization without reference model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a16ef0d9-cc25-4281-a02f-478cd785d2a9 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Rush, and Thomas Wolf
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf15fed3-bf22-4143-b499-94887fc6a6a8 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Let's Verify Step by Step
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf6cfa3c-378f-4d8a-a182-8bad4a1293be · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Generative adversarial imitation learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b227c1bc-9db4-4bdc-bfeb-552bf873ef87 · outbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Zemel, and Shixiang Gu
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ba59577-e2d1-4fba-bb7b-b52cc037427d · inbound
Sample-efficient LLM Optimization with Reset Replay Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b63b8c6b-554a-492a-bb64-920db5a6814c · inbound
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8981378c-8c65-4256-96ea-723f954ccbe0 · inbound
Compatibility-Aware Dynamic Fine-Tuning for Large Language Models Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.