Pith. sign in

Paper Citation Record · LEDGER

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections

As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2507.00018.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00018 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:54:49.897621Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-05T02:27:50.369374Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 155ecc09-68b1-49d9-a69d-74bd2cdbe62a · outbound

This paper cites an unresolved cited work.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:50.577843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:48.676604Z digest=sha256:12f78334e323f0d622d21ebd03b8b51e960cdc5f7fcb6cea47120a99b442ebdc

Observation 63bbd42e-8d4b-418d-b009-286bb6ee8bcb · outbound

This paper cites The Llama 3 Herd of Models.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:48.842036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:48.842036Z digest=sha256:5269c33422aac024da8220fb7b939e604b2dd0be903376a63aaec20a666e2aad

Observation 45844eb2-a2b8-44ab-97ff-08b651d4a23b · outbound

This paper cites Qwen2 Technical Report.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Qwen2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:48.952808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:48.952808Z digest=sha256:012dbaaa9197002d29010b1d280531301928068845c7f62457bd4f0ab3a9e05f

Observation 740526b3-24a8-467b-8455-7f1896fb89aa · outbound

This paper cites DeepSeek-V3 Technical Report.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections DeepSeek-V3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.027509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.027509Z digest=sha256:8001f5b5faea286577c4406c35f09bb6ca95f5d0782ff61d187d6f61ec883877

Observation 299141e7-ae50-4248-b9d6-b70024ebe1a7 · outbound

This paper cites Mistral 7B.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Mistral 7B

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.161443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.161443Z digest=sha256:704624dfb872bc2d0b023bc7bf46a68cd961324ea681564e5c7a317b60c17491

Observation d4d41498-8e2e-4cd1-a861-aabeb3829e17 · outbound

This paper cites LIMO: Less is More for Reasoning.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections LIMO: Less is More for Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.273212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.273212Z digest=sha256:06fe64600a9db61777d68a42ee6f628bad886cb123e0b90172f75d97a976d61c

Observation a6f19c3e-faec-4b6a-9a17-ba0818a03dd3 · outbound

This paper cites LIMA: less is more for alignment.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections LIMA: less is more for alignment

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.565495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.378898Z digest=sha256:6f67b4f81d822999178116b7c1d24536315a4391cb3940b6826154817d878964

Observation 2527488a-32c2-45bc-af58-26484812d526 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.466452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.466452Z digest=sha256:e29ec3671c6e685d4032200991491ec4f0a680c165ccce769e88dc4f2026fb4b

Observation 696e9905-ffea-42e7-b833-fc1a31717ae1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.570674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.570674Z digest=sha256:45b8aed3224dcec27b108cce706a05f0385efee279a938bc217e1d662abd69e8

Observation e284614f-2074-41cb-8341-709bb668a09d · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.652357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.652357Z digest=sha256:e5a36a6cb9c93df0c80cadfe837a6bc8a9c99d4d5757175720ae7272cadbb552

Observation 56f6fa23-9bfa-41a4-abf7-b522dacc86ce · outbound

This paper cites Deep reinforcement learning from human preferences.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Deep reinforcement learning from human preferences

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.759695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.759695Z digest=sha256:9cd0efbf277351c4441de395edb6baedd23622c6bdc4159fe93126d5c443e9c0

Observation 5ba82bcc-c98a-4bea-91d1-b7bba81c268a · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.763796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.763796Z digest=sha256:2b6d53ceb67c46b9d6efe1681887c620a5efb8f4b4a1179afdd0715c5bf09c43

Observation c9a44f64-bdb9-4b72-ab81-bbc7d787c66e · outbound

This paper cites Self-play fine-tuning converts weak language models to strong language models.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Self-play fine-tuning converts weak language models to strong language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.551637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.767381Z digest=sha256:f13ad6c2866ac0a5c9d05ce9b59e019cb3ff6587acd8a3e694d0fc09b8092cee

Observation 1990d6db-a32b-4898-9e51-03d125c24718 · outbound

This paper cites Learning Dynamics of LLM Finetuning.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Learning Dynamics of LLM Finetuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.770445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.770445Z digest=sha256:857e8d9433906c745f952745553c93c8fbaf036b3098336e35c64873a5e221e4

Observation 629305d1-60bf-4403-8943-45ff85560ba6 · outbound

This paper cites IQ-Learn: Inverse soft-Q Learning for Imitation.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections IQ-Learn: Inverse soft-Q Learning for Imitation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:54:50.177850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.774884Z digest=sha256:dd2a1addefd57a9eb4b49daa92a4770ce03235413668533a3bfed161ff48f8bb

Observation 07a74808-a6d1-4cb0-b197-fc1a7c6aabf8 · outbound

This paper cites Huang, Artem Sokolov, Matt Barnes, Guillaume Desjardins, Alex Bewley, Sarah Bechtle, Jost Tobias Springenberg, Nikola Momchev, Olivier Bachem, Matthieu Geist, and Martin A.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Huang, Artem Sokolov, Matt Barnes, Guillaume Desjardins, Alex Bewley, Sarah Bechtle, Jost Tobias Springenberg, Nikola Momchev, Olivier Bachem, Matthieu Geist, and Martin A

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.538820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.778563Z digest=sha256:4b882a7b4743a69d950f07968136a4829ca09f45c764d3e3eee8ab33c6c602b2

Observation cad91d96-f262-4d7b-abaf-21967381882a · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.782191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.782191Z digest=sha256:51f8f6df336b2cfa62e98d33fc2fe6cdf3f0901d5df827b119a94f28833f6a0b

Observation 6c44ace5-62ea-4b09-95ad-04272b92cf4b · outbound

This paper cites Ng and Stuart Russell.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Ng and Stuart Russell

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.525704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.785707Z digest=sha256:1df88448cf3c8dc0a8ebe320db807af9029d72a51d8f5b298484a997245e8fa0

Observation d7489688-7b8b-40ac-a258-ad77e3a60e29 · outbound

This paper cites Preserving Diversity in Supervised Fine-Tuning of Large Language Models.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Preserving Diversity in Supervised Fine-Tuning of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.789457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.789457Z digest=sha256:d220449a46bddf462df98b88f0588ecbd8e571a8441cd99daca92ee910e3ce3a

Observation c96eb0c2-448a-4c4c-99d0-6d47d8a457e5 · outbound

This paper cites f-GAN: Training Generative Neural Samplers using Variational Divergence Minimization.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections f-GAN: Training Generative Neural Samplers using Variational Divergence Minimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.793212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.793212Z digest=sha256:9d71e70238ff6e45e03c59db22eb3587506de1f04ef7f9cb3309a3ca81131b76

Observation cf414451-3f34-4e9e-8547-0d2c55c820d0 · outbound

This paper cites Sequencematch: Imitation learning for autoregressive sequence modelling with backtracking.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Sequencematch: Imitation learning for autoregressive sequence modelling with backtracking

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.513128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.797075Z digest=sha256:70bd6da07b55e7f05b8717519149967a16901fe40ec0e0aa5379fb1f6552672c

Observation 6a6ec825-10bf-418a-9762-31c5031510ba · outbound

This paper cites Proximal Policy Optimization Algorithms.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.801207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.801207Z digest=sha256:0901f6b54e282b6e7f36605313b4da132575bd21583e2f858c984314b8eb029c

Observation b37f6627-9498-47ce-bc37-d6434e484211 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections A general theoretical paradigm to understand learning from human preferences

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.501382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.805953Z digest=sha256:2db1dc858501488c72925802a8f249aa8dec38c54018685e0332735f253deae7

Observation 44bf425c-9e56-45d3-aa8c-48e512fbb21c · outbound

This paper cites Model alignment as prospect theoretic optimization.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Model alignment as prospect theoretic optimization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.489223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.809568Z digest=sha256:a22dc295db48e8239158afac38b6e7fbd87bee0bb0ce8f5b6c0a311578475a9d

Observation 1b3bfa17-fbe4-42c6-b82a-29bd522508a8 · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Simpo: Simple preference optimization with a reference-free reward

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.476399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.813601Z digest=sha256:5a7593dedb01894825a8a84955371bdf6f827b432d90cb6bf19a12d22ea8d239

Observation e98b2b3a-7252-4388-a3a0-77c37cb393d9 · outbound

This paper cites Disentangling length from quality in direct preference optimization.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Disentangling length from quality in direct preference optimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.463217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.818015Z digest=sha256:5836e0af36e8bfc6eafe151dd67c7d97939706af0a29a487f14ad0f9420d1d96

Observation f97dd5a8-7a18-4ff5-b2f4-fbb033ad99b5 · outbound

This paper cites Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.821661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.821661Z digest=sha256:f53f4fd46fb24947dcccc85f6ce2b905baf18d8eeea3573fb7d05a7e82cbe659

Observation 3d29a645-ca11-4196-986c-51f102089c7f · outbound

This paper cites Andrew Bagnell.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Andrew Bagnell

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.825840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.825840Z digest=sha256:3cd825a9b1514f2277514531d61b7986e53156262e61397eac012974350c3d90

Observation 7bd585c5-1ff1-46d0-a33a-e316b98163eb · outbound

This paper cites Ziebart, Andrew L.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Ziebart, Andrew L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.449803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.829770Z digest=sha256:94d34da53879c7167e8fa5d45fbbb620d31b1c830b999d2a1dfaca9c1c1ec56e

Observation d5944051-dabb-4ec1-b356-adc6efc21089 · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Enhancing chat language models by scaling high-quality instructional conversations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.436243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.833846Z digest=sha256:8340fc984079b0a2628ce2baef8e1641ffdbc29660b6186e8882728414809790

Observation 4da6cf93-416a-4c3c-bcfd-37ed4e76961e · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.838199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.838199Z digest=sha256:fd21ab2fb9756dbb1b453438d1f0dc3a6060083836410a9fabf23218567ac616

Observation 9cec943e-8c1a-44a2-95f2-7772676d4f1f · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.842747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.842747Z digest=sha256:35f8ac8261205e62a2ff411fd9c9c93cdc60edc37c1097456925dd081e097777

Observation 8ecdcd53-2441-4def-a799-4c63dd45c8d1 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.851308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.851308Z digest=sha256:f80b979f7841d77a1ffaabbf09fbcf2f2c3bdf6d30a176c21ea9efb753f20d86

Observation 1ac1d6ed-c92c-4bcb-b685-825f7291cc1d · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Xing, Hao Zhang, Joseph E

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.423410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.855476Z digest=sha256:588dcb28f2d519771057b7ab756e62d2920251d0f781341cb2d9201e082e81ca

Observation ebdac33f-13bf-48bd-87bb-d78afd04c21d · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Gonzalez, Hao Zhang, and Ion Stoica

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.409199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.859810Z digest=sha256:a1dd97cacf7d29221658af19cf7896fd973dcb387ce4b81a2bd1f9e62ed87f1e

Observation d4f260b2-1c9f-4368-8840-efb07f890521 · outbound

This paper cites RRHF: rank responses to align language models with human feedback.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections RRHF: rank responses to align language models with human feedback

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.395307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.864623Z digest=sha256:48926644750d195ab7e3dcfb701425c2edaed82842e32e4fa6692d4ea21c7d13

Observation 665991a4-8154-4b0a-a47d-ecd103f3e2eb · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.868787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.868787Z digest=sha256:33ee35adc07db445ea6586cbd57161f23e2916cdca851145169e15c589246295

Observation f00328f4-f5c6-4374-9e20-16d2f547f190 · outbound

This paper cites Contrastive preference optimization: Pushing the bound- aries of LLM performance in machine translation.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Contrastive preference optimization: Pushing the bound- aries of LLM performance in machine translation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.381363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.874260Z digest=sha256:f2be7ddfb7e8416bc279d0c8e4829bec9598ddb2201ed817879e4fbdc7b29367

Observation 8b364a22-3f24-487d-bcf7-1c013b6c60f4 · outbound

This paper cites ORPO: monolithic preference optimization without reference model.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections ORPO: monolithic preference optimization without reference model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.368537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.879242Z digest=sha256:d80650531ce6f89aa0cded0a267942bd509728f9fafab06884c0a7192b383e87

Observation a16ef0d9-cc25-4281-a02f-478cd785d2a9 · outbound

This paper cites Rush, and Thomas Wolf.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Rush, and Thomas Wolf

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.354243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.883972Z digest=sha256:4f36c59bc7ea51b7ccf3eb3893f81c343bd3e4cfc71a77ff40437386af51e947

Observation cf15fed3-bf22-4143-b499-94887fc6a6a8 · outbound

This paper cites Let's Verify Step by Step.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Let's Verify Step by Step

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.888136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.888136Z digest=sha256:e855703f5c3a8c903546e0a473426e305a0ff970aba3c0361f3d652fe7b875de

Observation bf6cfa3c-378f-4d8a-a182-8bad4a1293be · outbound

This paper cites Generative adversarial imitation learning.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Generative adversarial imitation learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.340230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.893394Z digest=sha256:c6ff64ac3fc7c1809652d20d9a086408ba7ecea382584abc09c786ab9c65e411

Observation b227c1bc-9db4-4bdc-bfeb-552bf873ef87 · outbound

This paper cites Zemel, and Shixiang Gu.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Zemel, and Shixiang Gu

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:50.326969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:49.897621Z digest=sha256:f6910aeddc04cecba7dac5aef1c97ccfb9f430b9b423307c6758b9a9b9357cfd

Pith citing papers

Observation 8ba59577-e2d1-4fba-bb7b-b52cc037427d · inbound

Sample-efficient LLM Optimization with Reset Replay cites this paper.

Sample-efficient LLM Optimization with Reset Replay Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:41:54.724211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T23:38:08.044450Z digest=sha256:f78c94421357754271b63995f7d7f944232bc4e4824efb22dcff01b04309074d

Observation b63b8c6b-554a-492a-bb64-920db5a6814c · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:37.903879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:1c2b7788c702ffc422704e9331a51b158e8e6fe69279fb4b3f19528cae0860e1

Observation 8981378c-8c65-4256-96ea-723f954ccbe0 · inbound

Compatibility-Aware Dynamic Fine-Tuning for Large Language Models cites this paper.

Compatibility-Aware Dynamic Fine-Tuning for Large Language Models Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-05T02:30:40.875988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-05T02:27:50.369374Z digest=sha256:329d66988d08c2684dd6752ce977fdfc744eb740afb99a3a8b1bb2aa58430f06