Pith. sign in

Paper Citation Record · LEDGER

SLiC-HF: Sequence Likelihood Calibration with Human Feedback

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 93 inbound Pith citation observations for arXiv:2305.10425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10425 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 93 of 93 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:21.000951Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:46.433136Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1722cd2e-8a89-48d6-ba18-325f40167411 · inbound

Self-Rewarding Language Models cites this paper.

Self-Rewarding Language Models SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.528194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:dd8c856784404065ec5e5a73c6140dda5f65534fab50e05626d494cb6a4002b9

Observation 5e276e74-323e-4467-a46d-42e1c255fd92 · inbound

KTO: Model Alignment as Prospect Theoretic Optimization cites this paper.

KTO: Model Alignment as Prospect Theoretic Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:17:53.546101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T12:17:53.478052Z digest=sha256:2ce55158160210436d9a8deedc0285475f4b1dc1643e351e7a6b0cc5b1b2fccd

Observation 34a278a0-618b-4598-9c74-270390f42fdb · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 237

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:38:37.016336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:d66a2491dbd0335d801e2580b0b30a037fa02f244bb153cce2d1353ae1f5db45

Observation 75b6c38f-8957-4816-bbcb-19d670c555ee · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 210

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T09:42:04.390390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:469d9b60a99f81b9ffc6d8c5904a0b7049db138cdd56627df0d38f21390a86a2

Observation 80e4268d-f709-4dcd-893c-ce3f06ca8c0e · inbound

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution cites this paper.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.961970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.961970Z digest=sha256:f2e49ef15c4196f23a8e3044762200839f70d5d70d17a849b15ce6db75f6ada1

Observation 309eb1cf-e393-4d62-bbec-c2ccdf9dabdf · inbound

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment cites this paper.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.080793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.080793Z digest=sha256:0629c149962c0a367266262720c3508636eb3cdc156d7b7530e7efb98cd926a7

Observation 5918f534-acd5-4a7f-8981-d919afe89c61 · inbound

Multilingual Large Language Models: A Systematic Survey cites this paper.

Multilingual Large Language Models: A Systematic Survey SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:00:14.366934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:00:14.366934Z digest=sha256:aeb05261ba094d04fd3bc353f06ec00dab5b34aeab2e9b83ed3f840c95651fdb

Observation 9a74b634-c129-40ca-b1e4-d533e99c7df5 · inbound

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd cites this paper.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.961445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.961445Z digest=sha256:8cf9d6c0351d2b64477f673d7d084459c08d81be3efdbc73175af7fd3584a1af

Observation 8248a1b0-66cd-4763-8a01-371e17283dbd · inbound

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs cites this paper.

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T17:05:15.227369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:05:15.227369Z digest=sha256:a254d7a916fc48f5571b1e7648c50f775022e045d6164837d211b3e238df6ac6

Observation 908931f4-5777-4768-a9c5-6f632332c655 · inbound

AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward cites this paper.

AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:36:36.413997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:36:36.413997Z digest=sha256:118fdee67ec262bf092f29113cad481fd068f7a164aab982c91791b0d12779ce

Observation 0973c55d-5020-4909-848b-b2cbc975cf0e · inbound

Harnessing Preference Optimisation in Protein LMs for Hit Maturation in Cell Therapy cites this paper.

Harnessing Preference Optimisation in Protein LMs for Hit Maturation in Cell Therapy SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:29:23.724112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:29:23.724112Z digest=sha256:d1d6bd22ff1dca8722371f51f085ce5e4c3ee00a3f9ce740e6b62e6f5b876581

Observation 7a0c075c-9a59-4218-876e-38c6449f733a · inbound

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies cites this paper.

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:22:44.243462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T08:20:05.898025Z digest=sha256:c29073f9d0ab8280501297f58fb8e04ddeee725e7fa791aa51e3274315c2c717

Observation a9a2a8f0-9cbb-43c9-8f70-ee65d0b2afc5 · inbound

Time-Reversal Provides Unsupervised Feedback to LLMs cites this paper.

Time-Reversal Provides Unsupervised Feedback to LLMs SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T23:21:06.753886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:21:06.753886Z digest=sha256:6b1b741e32d7d93a5590534c8b802c100facf83a7bf3548a8684f0074ceea4d1

Observation 0908ca83-8015-4323-929d-41e2f8c91328 · inbound

T-REG: Preference Optimization with Token-Level Reward Regularization cites this paper.

T-REG: Preference Optimization with Token-Level Reward Regularization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:56.458770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:15:56.458770Z digest=sha256:f88e953623dd88459bf91b0f2379d187861aba22926c567948ce7307049e3137

Observation 0f707a22-fc9c-4ed6-90c0-2268f4aafa11 · inbound

Weighted-Reward Preference Optimization for Implicit Model Fusion cites this paper.

Weighted-Reward Preference Optimization for Implicit Model Fusion SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:23.264757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:23.264757Z digest=sha256:f6effb142606747d43075cdda0027231151b4441382f3ab042bddd88bc9e29b6

Observation e1aeeee2-82f4-4f4d-a930-29231ddb085f · inbound

SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs cites this paper.

SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:57:44.387680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:57:44.387680Z digest=sha256:3e3b7e67dc02923f130d8061bb85e79b44a53902ae9ca8a89026e3cbf5a78116

Observation 162e90e5-149f-4d8d-aa28-12054b696c37 · inbound

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration cites this paper.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.422799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.422799Z digest=sha256:a7e66cb3951242bf289ea4e38990b659b5b280a25e0f443413112f0361d7d08b

Observation 3f42c634-3393-4e23-8de7-f4e7bf7442a5 · inbound

VideoDPO: Omni-Preference Alignment for Video Diffusion Generation cites this paper.

VideoDPO: Omni-Preference Alignment for Video Diffusion Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:09.709676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:09.709676Z digest=sha256:1c86a2e823f6e026423266f79fd50bda6f1765c0f699a326e5275355f77c8a47

Observation 252bde81-682b-48cd-b034-cb66fe3413f1 · inbound

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment cites this paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.816513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.816513Z digest=sha256:ae5b0065aaff0dd829253641585ae75b884c6d67e4c29e47cbe42408b995e881

Observation 288e9c2b-27e3-4ba6-a0e5-93f5ba2565a2 · inbound

Understanding the Logic of Direct Preference Alignment through Logic cites this paper.

Understanding the Logic of Direct Preference Alignment through Logic SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:21.081152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:24:21.081152Z digest=sha256:887968c8abc51aeb75d85e6eb17f5a53d4528393f0824eec75a055740bc17625

Observation cf9f3441-5d49-418b-a2bf-8de1a4f6eaa8 · inbound

Efficient Long Context Language Model Retrieval with Compression cites this paper.

Efficient Long Context Language Model Retrieval with Compression SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:37.339895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:37.339895Z digest=sha256:864577cefcba7daa426db3680e8cd0c1a94bb65eb63d23129f6aa44679c344e4

Observation 4feb6fe3-b419-460b-a086-09ad778f9d56 · inbound

DIVE: Diversified Iterative Self-Improvement cites this paper.

DIVE: Diversified Iterative Self-Improvement SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:49:13.518160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:49:13.518160Z digest=sha256:3848dc2b94def3db61c750b720c918704939c0226dff7804c5f9c50d1c35af98

Observation ccd42246-84e3-40b2-ba18-505b2e9a2926 · inbound

Online Preference Alignment for Language Models via Count-based Exploration cites this paper.

Online Preference Alignment for Language Models via Count-based Exploration SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.677210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.677210Z digest=sha256:e195b92dbe2f3dcb252eddc873b504476c12d35ca28b683da15394e234e95496

Observation e5b4aa39-da6c-4166-8853-4901daca3e74 · inbound

PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament cites this paper.

PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:31.305980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:39:31.305980Z digest=sha256:963098d56e44f49dc8b4facdc3abf28fd9e61e3b67d07ebafac41a5714c5443f

Observation 487da223-b5cd-4c8b-b54b-6a13d2a10d39 · inbound

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step cites this paper.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.457214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.457214Z digest=sha256:f7de6f3303ab9611fe7a3f147ed3f9795161052691733d7d634a6acb53c1e514

Observation 4f3d6aaa-03ba-48a7-8dcc-31af6a49cd52 · inbound

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation cites this paper.

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T15:33:49.392420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:33:49.392420Z digest=sha256:8935ad19d568a246a21c377bc27de301d4f64c896ee851f35bbec314b283da8c

Observation fb60df69-5da4-4d2d-bc3d-6c928af3da1a · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.896255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.896255Z digest=sha256:dae5efc31da8127bc22f11aebcb2aaf5e0f19496e765223061c381141eadbce0

Observation 049fc6f9-58a1-4cf3-831d-abe63ed3fee8 · inbound

A Checks-and-Balances Framework for Context-Aware Ethical AI Alignment cites this paper.

A Checks-and-Balances Framework for Context-Aware Ethical AI Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T20:09:54.404741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:09:54.404741Z digest=sha256:dcd29f1474c042da58b3de4625a54b8476444e2dbdece3ee535e527d88d2badb

Observation a6eccc06-70ba-46d1-a3e0-5158e1fbea3c · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.583638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.583638Z digest=sha256:ff51a24e9dcf0b1d68f2114a34791711f4a1fcfbd17d8dc868db80a8983623d3

Observation 52741b9a-7c71-44bf-85a1-fe1cb7775f6b · inbound

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms cites this paper.

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T06:04:01.644525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T06:04:01.644525Z digest=sha256:135f93bb61bb807ae51cddd5ea2ac60e6f3aab665e5493c09b5c73adbfd9b7be

Observation ff55ae22-d549-4c47-bb36-1e8685b45293 · inbound

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective cites this paper.

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T04:08:52.221810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:08:52.221810Z digest=sha256:824559a3acffc3f25766d041a695a54910a348c8e432163ee8dde4737bde9eef

Observation 2380c056-2542-4170-8d4c-6ccb870716e3 · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.564192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:24a306238cc99aa36607ac1de28b337cb919747d0d6a90c5208909b95e64121f

Observation 3c00348f-ff9c-439e-a0ce-a04c3d03c991 · inbound

Design Considerations in Offline Preference-based RL cites this paper.

Design Considerations in Offline Preference-based RL SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.069597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.069597Z digest=sha256:40909f6560ca7593a1304a698a715644569f4813c068d05d602328907dd15d8e

Observation f423684f-2bfc-493e-9f0c-f0ec5536991d · inbound

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples cites this paper.

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T11:57:12.067249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:57:12.067249Z digest=sha256:df08b03fd944ba9dcf9de4522b6e0e7b8f5e103823cdf6ae4277804f02f15039

Observation 59175734-fc8a-42e0-9c75-894c1cea7c81 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 193

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.402266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:f7da27bb03fadf8b55380c3701e38b1ab4688414f5e2cae50168b6b2ef371375

Observation 3beabb32-5b52-4fed-a488-db01caae4647 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 193

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:21.000951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:21.000951Z digest=sha256:03f8c3567d2bc4ff9a06d200dc38b3c2a991aec0376773497ada28d9bbd6d6b9

Observation e9883153-07d2-4036-bb4c-bde0e9152d9f · inbound

Direct Advantage Regression: Aligning LLMs with Online AI Reward cites this paper.

Direct Advantage Regression: Aligning LLMs with Online AI Reward SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.695729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.695729Z digest=sha256:4228d1928eafba3a780c229a39739c8e4340fda8cbe4be407a75e48b9e430d86

Observation dd954a4b-1eda-4f04-9c3a-1af5296727af · inbound

From Evidence to Belief: A Bayesian Epistemology Approach to Language Models cites this paper.

From Evidence to Belief: A Bayesian Epistemology Approach to Language Models SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:20.674454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:52:20.674454Z digest=sha256:bc55b6b72cc453c64875c4e223ea61b33cc15d04d45589dbabae2a7be2f4a97d

Observation a00b376d-be08-4587-9f38-56accb3513c7 · inbound

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL cites this paper.

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T00:59:21.913467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:59:21.913467Z digest=sha256:d5d3809465d08816d42e9107b5e062effb30d525cefa4a33f73d656bb1a37221

Observation 136f0f34-c46f-4444-9f9b-2375ff74aa67 · inbound

Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes cites this paper.

Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:47.226126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:47.226126Z digest=sha256:ca9e58cef2cd412d9faab7ac14859bf8108a2ccf4404035ecdc308062a01c4b5

Observation fd360663-2ac3-415c-af6b-3ab7597baae8 · inbound

InfoPO: On Mutual Information Maximization for Large Language Model Alignment cites this paper.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.152786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.152786Z digest=sha256:3b7ac72bfefb0a8d05539502cd7c3b176fcced0119f04857211dc5893de09a45

Observation ebbe1492-80e1-46c9-90d7-5d7cc68d3708 · inbound

ShiQ: Bringing back Bellman to LLMs cites this paper.

ShiQ: Bringing back Bellman to LLMs SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.640889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.640889Z digest=sha256:af71413584ba6d60649544ad41d35ce1b7af5b9248b104d43df951bdae3fef7b

Observation 9f678ab7-fc80-4070-93ff-4800f562e6a4 · inbound

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization cites this paper.

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T21:06:21.637749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:06:21.637749Z digest=sha256:51a1b002c7c0e29ca1d8dacb863c0330e684d96e6e7918104be0542279e913e3

Observation 63e8c024-b2b2-49eb-bf14-fef4b5bccc41 · inbound

CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design cites this paper.

CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:33.685941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:33.685941Z digest=sha256:5fabc062ec8c469706601c6de1c432e16f165ab1dbf89f437ec3b32e1593df97

Observation f2cf9eaa-5d02-454e-91db-cc41a6877c14 · inbound

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? cites this paper.

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:46.299308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:46.299308Z digest=sha256:4cf97e9cef475bfb025bf264ed988b34b549764e598d812b9c36f827ad08f3ad

Observation 87bd096e-8ccb-4955-abd4-69db5d72d6bc · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:42:19.370731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:1e68c647289e3e7158f1494a4027e0925b59ecf2b7ca89ad99f0e39eb49b1ecf

Observation 0f25a33c-c908-4cf8-a2fd-d6b57a926c3f · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:04.554366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:04.554366Z digest=sha256:48d3c1570a581ee2ee506962a312e860c7b0c7dffbb5051bb1417cb7e431804b

Observation 0b8a432e-e368-4c56-b05f-05153a48c8e2 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.251391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.251391Z digest=sha256:2878c9c53a97835151207a0562eb8c42cd92dab3cb22cbb78d39a2a17be795dc

Observation 821c1dfc-b893-4828-acbd-f8130eb23105 · inbound

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences cites this paper.

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:57.177745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:57.177745Z digest=sha256:2b86b28d6e57628a64287f1f7f046873c4a645a8e8151e3824ca6ed7e2535aa1

Observation faf8bccb-9cd4-44ef-92c4-992109603475 · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.316647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.316647Z digest=sha256:d54a1dfd2304166a5b503c2cc7a65831037d3158e93eeef3712935f07d833a71

Observation e928c41e-609e-4942-909e-de9dc75b6c28 · inbound

BPO: Revisiting Preference Modeling in Direct Preference Optimization cites this paper.

BPO: Revisiting Preference Modeling in Direct Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:07:35.619682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.619682Z digest=sha256:0c1cc7dba27f1007a45cd174aa684745c23bedc69641ee2e16e1a0b0855fb248

Observation 96a4e02e-511b-44ed-9d04-e4057629f929 · inbound

Explicit Preference Optimization: No Need for an Implicit Reward Model cites this paper.

Explicit Preference Optimization: No Need for an Implicit Reward Model SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.036861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.036861Z digest=sha256:3f44b0128af7f6bfba9a9f8a6f486843dbc629ed7a9a512b55e2e0724ec9bb29

Observation 1dc57259-667c-4d60-b5b1-f7b8a9949c6b · inbound

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization cites this paper.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.342641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.342641Z digest=sha256:16627da22002c9399d2a01fee3fc412cde49f9146983efc5d396227135b6fc91

Observation ad212efc-b520-4d1a-86e3-86662d50493f · inbound

On Monotonicity in AI Alignment cites this paper.

On Monotonicity in AI Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:08:05.322977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:08:05.322977Z digest=sha256:ad9bcabbc13da373f9fad13286cec2468ff1b3858fb2b172e98895625830b96e

Observation 9520ea74-8889-43c0-bd1b-372e99a61b44 · inbound

Debiasing Online Preference Learning via Preference Feature Preservation cites this paper.

Debiasing Online Preference Learning via Preference Feature Preservation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.720004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.720004Z digest=sha256:9959af126ac64e2dd2548722ba701f40a5ae2d4177c070b8f18658faf1021dc4

Observation 586052fa-c7b1-42f5-b7d7-8e467d6e9ebc · inbound

Value-Free Policy Optimization via Reward Partitioning cites this paper.

Value-Free Policy Optimization via Reward Partitioning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.293738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.293738Z digest=sha256:67356a66008acf627255e4d483de33bcdb28b7eafe9673689dff2f4d177a795f

Observation 7b97a484-904c-4c66-8d4a-e5b1f170f0c2 · inbound

Optimising Language Models for Downstream Tasks: A Post-Training Perspective cites this paper.

Optimising Language Models for Downstream Tasks: A Post-Training Perspective SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 274

Resolution
unresolved
no resolver link, observed 2026-08-06T22:44:44.816609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:44:44.816609Z digest=sha256:eb008b70706e6e889f6235dea9e76e585e484bf5a0ac6132877186b540774872

Observation 665991a4-8154-4b0a-a47d-ecd103f3e2eb · inbound

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections cites this paper.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.868787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.868787Z digest=sha256:ba0aeeb94686b1017f7fc6a60f1131da9abd2382e008c217281fa1601882a740

Observation 75cfad5b-8491-433a-a1af-4a7ef81d8a73 · inbound

Improving Consistency in Vehicle Trajectory Prediction Through Preference Optimization cites this paper.

Improving Consistency in Vehicle Trajectory Prediction Through Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:34:37.335053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:34:37.335053Z digest=sha256:f3505a11027780bfe0e4d7227ac718aa621cb48c969b51c29251475dc5efa6dd

Observation 1036f3d1-a9c8-4f11-bc1c-117c7058c47c · inbound

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users cites this paper.

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:02:07.792836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T05:58:17.452837Z digest=sha256:a7ba5f1442bb229a68fb867ce5d7cef9b2e42a1322a8707c251e0024f7e9f798

Observation 6b833df4-92a0-4977-88b7-29c850e08e7f · inbound

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) cites this paper.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.808890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.808890Z digest=sha256:74bc9710d8fc913290cf66a11e9b0631239082b974669a43a412ff4a82a7c7f4

Observation 5406e952-4b37-4599-b46d-a01751c6eb7f · inbound

Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning cites this paper.

Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:49:21.572980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:49:21.572980Z digest=sha256:2439678242e8dd8b455304014096d166aa33aab09007e4c18a0e10be93596ee1

Observation d00b723e-d4c4-46eb-9e1d-42352f6ba700 · inbound

Phi-Ground Tech Report: Advancing Perception in GUI Grounding cites this paper.

Phi-Ground Tech Report: Advancing Perception in GUI Grounding SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T10:28:56.119117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:28:56.119117Z digest=sha256:158a6b912dc2bc469deaeb47c26b2e1cde06a6b6b62b738ba284531c19f73939

Observation 5fd91298-c1f0-4ab7-ae4c-bcc8c72fe381 · inbound

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints cites this paper.

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T21:33:14.871864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:33:14.871864Z digest=sha256:41c552d1e63ea4fa251d43f8e6c9ee1447dd2a453403f289f7be95b7cfb86b32

Observation 9727fb8b-240f-40a7-a8a8-5838076e0ff6 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.779960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:6fa8c857081cd5a2102d45245654a4961780ed0646b5c4ed264f0dc3a8f0b6f9

Observation dca28929-3147-44b8-80dc-00812939bbea · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:23.843777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:23.843777Z digest=sha256:5fe4e04432a9f85a140df1ed661adeda4eb5cca331f1915144ee785b3043dc43

Observation 6d45353d-31f5-45da-842c-4dfe6de29502 · inbound

Alignment-Aware Decoding cites this paper.

Alignment-Aware Decoding SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:40:15.787091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:40:15.787091Z digest=sha256:d96eca504e1da985f6e72b66a173cfc6f5aa92a8fe036891609d0a2e330241fb

Observation 1409158b-53b4-4237-b651-722e6ea4b7a9 · inbound

POPI: Personalizing LLMs via Optimized Natural Language Preference Inference cites this paper.

POPI: Personalizing LLMs via Optimized Natural Language Preference Inference SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:42:24.390403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T05:41:58.231139Z digest=sha256:f7a5e4ba80ed65d70332a283f9c21b7f2a300621ddb37bed752c9d5a91714ebb

Observation c4a76026-141a-4769-8298-c6042f584ffc · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 209

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.162389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:4906f9be5d198808340118976503694cf59a705523c0d766e9ecc191cc7fd84e

Observation 0783e4e0-6a46-4245-9168-026877db7356 · inbound

Mind the Gap: Structure-Aware Consistency in Preference Learning cites this paper.

Mind the Gap: Structure-Aware Consistency in Preference Learning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.593236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-07T06:41:18.311986Z digest=sha256:2f99d14bc26fbb7bf11d6bf96b202cbe15bb90798a744a002348af880e3da79e

Observation b7f30dcf-b92d-47e0-9d7e-b27dc150c79b · inbound

Anomaly-Preference Image Generation cites this paper.

Anomaly-Preference Image Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:15:37.409124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T18:45:20.508200Z digest=sha256:c003dbe47c8f752bd99f3b64e3ac6ef05452b6224ee1f8fb324e008e492434b3

Observation 2eb63ced-9368-49ee-a8e6-807028c53ccc · inbound

Anomaly-Preference Image Generation cites this paper.

Anomaly-Preference Image Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T00:43:53.717724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T00:39:31.730684Z digest=sha256:e85a3ede078f25822bac53237b983bc2bc50807dd97be12083ada17baa4328d5

Observation 1e1e8f3b-289b-407d-b2a8-157ed91bcedc · inbound

Anomaly-Preference Image Generation cites this paper.

Anomaly-Preference Image Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:12.283055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T00:36:26.957598Z digest=sha256:2164456678b335b367c4a44d355ee32bf868df8fad3bd1973b7ef42684111335

Observation 7fa8adaa-2ecf-40ef-bb7e-38081db26d93 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.182781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T17:01:04.571087Z digest=sha256:77684d7e6c449ef4024087178f6bec1659cfdd316c96bf5090d31ff3095eecb4

Observation 754f1ef1-1628-4e0b-b4f3-c6b36a820e0c · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.407105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:966e8083dd906b85fc4d1550644d9c4d576f5aeef29bac39b389f8077788b1e0

Observation 75f08799-ea08-4a97-9d8d-68b5a3dcbaff · inbound

CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference cites this paper.

CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.841961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T21:36:26.958957Z digest=sha256:e89dba6d7ee171afa336b685111808eeb9bb07a8a87c1b80fc9f2ac555beff9f

Observation 188984cc-6ede-4346-90f0-ba835b7ebe66 · inbound

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment cites this paper.

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:03:58.124603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T04:59:41.187000Z digest=sha256:69cbc431570e3af5053ea144a8974ab676158b1d28e33e70de4ca3d4d635fe6a

Observation fde0b802-f605-407f-921a-0821d188ab28 · inbound

Token-weighted Direct Preference Optimization with Attention cites this paper.

Token-weighted Direct Preference Optimization with Attention SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:06:11.750006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-22T07:05:23.189975Z digest=sha256:64bad1e644936f7acb2a9bd87cde7cc0549d322d268a9de0431ad6b06e6fa989

Observation b6439581-fc26-4740-9979-7b9337f814fe · inbound

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization cites this paper.

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.435524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T14:58:33.167891Z digest=sha256:0b88d0b5d42823648f49a2664956cab25943ead74a7a5af8373263f96310c8db

Observation 496bb23e-2733-4db9-a00b-e336c57b3c5e · inbound

P$^2$-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization cites this paper.

P$^2$-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.135124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T11:13:37.291834Z digest=sha256:95d303a4a95d74a03265b0e9651baf128bb39e7062f245c147d1e31c502d5294

Observation 72de1279-8562-4fca-b409-83c5916d3e1e · inbound

MARD: Mirror-Augmented Reasoning Distillation for Mechanism-Level Drug-Drug Interaction Prediction cites this paper.

MARD: Mirror-Augmented Reasoning Distillation for Mechanism-Level Drug-Drug Interaction Prediction SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:58:03.338464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T09:44:38.871250Z digest=sha256:12adebe5b178f1be8d68aaccd1dcf49b3112ec9578f107971fea5b11767bad7c

Observation 06574a1f-879b-48d7-b34f-95e68bfb9f58 · inbound

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech cites this paper.

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:08:37.812263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T06:01:06.123832Z digest=sha256:0fb9a7a4ca0403a7cf50799a25d884608cd744dc70d8b8eed848dd14b1469030

Observation e3deae7e-03de-46fb-9a72-0ded20e4c8be · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.434801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:37da4656dc7a0b48304b0a498355690bd8a7bef67a387910adfd0c35eef3201d

Observation b1c3db5a-96df-4011-a5a9-6bf7ac7609b3 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 194

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.422406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.422406Z digest=sha256:e5894f1be13c7d54c36d5a4b0fb64880ba11841fb0c5426cb7945dc4b8ad2750

Observation 8fc7b84b-5fce-481d-8dc9-a489f7d6da51 · inbound

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs cites this paper.

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:35:47.677633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T01:22:16.176398Z digest=sha256:125d69a841b28ab20111b266402066bc6f7e90fbee90c0930632885d6a11058c

Observation 8051384f-8274-450f-b576-80cc0349ec02 · inbound

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation cites this paper.

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:14:26.741596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T06:07:44.343539Z digest=sha256:acea6a2194afba1b83feaf8f05d568ebd1e3a75a2646aecc77d4bf17488d1112

Observation cccff79b-33e6-4cc0-bb99-7ab96de8cd7b · inbound

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation cites this paper.

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T09:37:24.616723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:37:24.616723Z digest=sha256:4878dabaff15580050da66ab06f2161a4e3545a9596843053f719e1854527b67

Observation 1470d54a-8ab3-4f29-b8b4-1ffcaa06c8cf · inbound

Unbiased Alignment for Large Language Models with Noisy Preferences cites this paper.

Unbiased Alignment for Large Language Models with Noisy Preferences SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:e5a62540c988a303487451ef224528807a0cb60ce6921832317e416ebcf0382e

Observation c5def73b-34c0-4807-b171-8bde50c6a6f6 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 254

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:896718b0f2950570be48b2e6bf3fff102ece76b9dd0337c40d8715c6228a4b3a

Observation af3df9e8-9307-4cec-bf01-c72357f13696 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 255

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:01.968095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:01.968095Z digest=sha256:1cdbfb2c02ab0c0308613b6d4da139503c76f3e58f87f13982beb33e7201ba99

Observation fe93fbcb-ec48-420e-a923-cb85d810674c · inbound

Every Sample Counts: Supervised Fine-Tuning of Language Models with Pointwise Constraints cites this paper.

Every Sample Counts: Supervised Fine-Tuning of Language Models with Pointwise Constraints SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T05:26:06.829382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:26:06.829382Z digest=sha256:173c2d0b91b918ad8b469f858c2659b4b7b748b5afdb19712dbe199619f5c7a8

Observation 8921611f-4095-440e-86e1-14e4d0f270dd · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:59.868237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:59.868237Z digest=sha256:f01ecc69ba46050fcce7a060288f140a6f566d9e541ef0c6af1129e2c83b5dcd

Observation bf7d0bd4-da96-4800-b1c7-887d73d5295c · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 173

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:35.893789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:35.893789Z digest=sha256:2c40882ee690a4ea751361829cdeced05fb6c0809f2c7de50f722632caa6fabe