Pith. sign in

Paper Citation Record · LEDGER

SimPO: Simple Preference Optimization with a Reference-Free Reward

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2405.14734.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.14734 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 100 of 127 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.525929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T20:05:34.173711Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 63664023-8e42-4bd5-875a-d32642e75721 · inbound

Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing cites this paper.

Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:58:36.882351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-16T06:58:36.684583Z digest=sha256:bb07228fd115668485f5ef677d13eae0c99bf80e9c2bca6680ee8b073b817bfa

Observation baad56b5-f9b2-45f3-81fc-d90330cbd6a9 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:58:17.123554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:81dd3d266a34b23b802d1c952d175ccc708f53c03cca3daa635a7eaef115fa8c

Observation 22771146-9716-4e82-99a5-eb2e0f3b1d5b · inbound

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution cites this paper.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.757342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.757342Z digest=sha256:5a866d42e2b64f1b3fc35336e21da98e28293699c17e7aa3fc6aabf6e2e7296c

Observation 18980d1f-0ac9-4ffb-9fa1-b91d882aea71 · inbound

Adaptive Decoding via Latent Preference Optimization cites this paper.

Adaptive Decoding via Latent Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T20:31:53.174774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:31:53.174774Z digest=sha256:b21cc32e80918b8f9ada0167be852032b0bfad6f131464858432ce7724b146a4

Observation 70e7b513-cee4-435a-b139-692ff894730e · inbound

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment cites this paper.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.999314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.999314Z digest=sha256:a7187f799f7ea7508fd96c949df5750cbb15ba6431539a254e381ec668ea1d81

Observation 47c5ec60-efd9-4ef5-a91c-7304256c0429 · inbound

ProSec: Fortifying Code LLMs with Proactive Security Alignment cites this paper.

ProSec: Fortifying Code LLMs with Proactive Security Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:10:55.524976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:10:55.524976Z digest=sha256:4a799c249dd5255b30ffe8e1e879e9a428b8babd4b77a4b4061c3dbd3efa93f0

Observation 8c071df3-41bb-493c-a256-f81fcce82abc · inbound

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models cites this paper.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.621710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.621710Z digest=sha256:9ff825bb00e647a69aa0acc815d316a93e42d62d084988e307c5ce13bb656f2b

Observation c46ffce3-d532-42ed-a13c-dea5659f2527 · inbound

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning cites this paper.

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T05:13:41.509518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:13:41.509518Z digest=sha256:bbb933c3dc845900c5b9b0e7cecb89fb7f724184cb190292aab9ddef4749d4e4

Observation 265d9257-8f0d-4d7f-bb49-075d279d7e15 · inbound

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies cites this paper.

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:22:44.274185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-23T08:20:05.898025Z digest=sha256:1f28945807ec3a9aba212b1a7b552e5075e5321082c18181c25f08d243d7702f

Observation eb18100c-eed6-430b-a3fd-c97b21c86cb0 · inbound

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? cites this paper.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.742162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.742162Z digest=sha256:eb53a113ed10f6f54637a4e44190f4116a2b820b181bf4436e9825a50f21b8a1

Observation 4c1eeab3-330e-41aa-b0d6-4a8c02f8e333 · inbound

T-REG: Preference Optimization with Token-Level Reward Regularization cites this paper.

T-REG: Preference Optimization with Token-Level Reward Regularization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:56.151799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:15:56.151799Z digest=sha256:5e09d54440b9474e8952b75b22e5dd85541a69ac0a870ac327c0a1faf1559832

Observation 802c7432-bfa4-4aa4-bf4b-795b6f5085a5 · inbound

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization cites this paper.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.083942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.083942Z digest=sha256:3170b5d19329b796d7a7f07843b033b9646d5834960e1c5c2e5a2c759955040a

Observation b0e2d84c-c9d5-4597-acfe-580aae8fcb3f · inbound

ALMA: Alignment with Minimal Annotation cites this paper.

ALMA: Alignment with Minimal Annotation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:38:06.767913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:38:06.767913Z digest=sha256:01a38e7692bc9fc71837413667e95e091f488167c51d027677e0e456706cee66

Observation 3c91929f-c41c-443e-bb96-11fdea8b1140 · inbound

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation cites this paper.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.364993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.364993Z digest=sha256:3bc73e4f70ac5f5dfa7de36a98bb17e4bab2cba01aa3c8c64fd2bc3d0a5d09e7

Observation 026e6fdb-1a4c-42a8-9602-03a242430c1b · inbound

Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models cites this paper.

Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:42:46.559026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:42:46.559026Z digest=sha256:e3c598cb93cf002ecc1910563fbbb338f4fb8c717a7f14d7b30e2a136a2b223c

Observation 2942322d-55b3-4fc2-bc61-621f284992ed · inbound

CleanComedy: Creating Friendly Humor through Generative Techniques cites this paper.

CleanComedy: Creating Friendly Humor through Generative Techniques SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:00.459083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:00.459083Z digest=sha256:834d01fe3fce833b4d91d2aba39c71b417b27187745b71864785037c2e7b719c

Observation f16a6eec-9670-4ec8-95b9-68c29bf7a740 · inbound

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration cites this paper.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.375425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.375425Z digest=sha256:678cb192529581ba0e42db7f3feaa50f7de96987622b57358fe5d8cfed64204b

Observation b24f63fe-024a-40dc-b521-c3f89c9a7e50 · inbound

WEPO: Web Element Preference Optimization for LLM-based Web Navigation cites this paper.

WEPO: Web Element Preference Optimization for LLM-based Web Navigation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:48.837869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:48.837869Z digest=sha256:36ac1b3a08f301f20d938f27a7f926b9f4944e9f0f9d2adf1ae5e5d09f6d843c

Observation 3ada8865-7075-4e2d-ac37-d98dea9d8925 · inbound

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning cites this paper.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.154835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.154835Z digest=sha256:69021a5f97a8aa3588317bee4af85373a74f63cbfdd20217a0e5d37a9b6a55ef

Observation 11b68d6a-a06d-47dc-9a3e-a1628785fa47 · inbound

Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models cites this paper.

Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:15.481684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:44:15.481684Z digest=sha256:a6a3f5874a5f97d40432a6072f04231b490945ae4dcf0ae677249994f5bc4cca

Observation 5b26f855-819d-4343-aa5e-9c62415b335f · inbound

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model cites this paper.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.966786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.966786Z digest=sha256:a9480960173363016d4f92321f0620a3cffa6d38ea3892188abae0725796e5a0

Observation 8489d4b5-97ab-47f5-9741-966797a17cf8 · inbound

Hansel: Output Length Controlling Framework for Large Language Models cites this paper.

Hansel: Output Length Controlling Framework for Large Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:37:34.261999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:37:34.261999Z digest=sha256:8ade6c819caab7462242f9482a44e5b463d9515f47b5e39ed74d76da688bde31

Observation 0943c1fb-e0db-48ed-9cf5-9f167b4771ef · inbound

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment cites this paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.937213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.937213Z digest=sha256:3c27482844a00b39eda3cd3b5ef4097d8e80f9600eadb3b8d1fd84f714046192

Observation 22304670-4771-4e4c-b3e7-b02ca9d3d064 · inbound

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples cites this paper.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.407893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.407893Z digest=sha256:5478d94e27ebaf71328e4e1c498046b9406f1fdac2a2ce7a8cd3b663b45e6ad8

Observation 7e34296a-7701-469c-a51c-ec0dffb8eec0 · inbound

JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs cites this paper.

JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:17:49.032269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:17:49.032269Z digest=sha256:9506955a33254f8627f5f27c03dfa2c7d650fa6f83e32a016bd3e5504a8a8bea

Observation a48d06eb-8497-4b7b-899f-0363ca074d8b · inbound

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback cites this paper.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.309116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.309116Z digest=sha256:a7860099ce85134ee707615711e073592686a7cd1f67dd6183be24ab2d97ebcd

Observation 52e84489-14ee-4766-9d38-dfb01f313e87 · inbound

GAS: Generative Auto-bidding with Post-training Search cites this paper.

GAS: Generative Auto-bidding with Post-training Search SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:56:02.095376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:56:02.095376Z digest=sha256:03e5fe4ee184f45294827def70430b4b573b79f4ab51570c8bd3377ff2ee0144

Observation 1395c49a-5a41-4d11-bc37-eb62a3e543cd · inbound

Multi-Agent Sampling: Scaling Inference Compute for Data Synthesis with Tree Search-Based Agentic Collaboration cites this paper.

Multi-Agent Sampling: Scaling Inference Compute for Data Synthesis with Tree Search-Based Agentic Collaboration SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:53:01.299517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:53:01.299517Z digest=sha256:32cb44ed037e3d42682d5da841919f15d76ef322e9fab7ce76ed1a3db5d94fe1

Observation bfb8b9a2-dfad-482d-80c0-6fe951c4ab8f · inbound

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking cites this paper.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.365253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.365253Z digest=sha256:2eaf543412c9bef9e741cfbaffa18f6c616c6579557ab874d558e81a83c6a2fa

Observation 3e8f2b64-e06a-4f87-b95e-65de95d3453b · inbound

Plug-and-Play Training Framework for Preference Optimization cites this paper.

Plug-and-Play Training Framework for Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.130362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.130362Z digest=sha256:a284cda9637613863214a206b4330c64d886bf23766b370d7198e2e6b7c54f73

Observation fd3a5cdb-fec5-4abc-ab22-a51ad16235c2 · inbound

Prune 'n Predict: Optimizing LLM Decision-making with Conformal Prediction cites this paper.

Prune 'n Predict: Optimizing LLM Decision-making with Conformal Prediction SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:54:43.781667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:54:43.781667Z digest=sha256:7ffdc98f7d1631c662abcfc53be0b2d734991282e2ff886189530328a321dbcf

Observation 87e5657a-153e-49d4-8ab5-6b8a297d1cb5 · inbound

From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning cites this paper.

From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:11.073589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:11.073589Z digest=sha256:125b1625ffb4df24f951246779b61ceb59914b205e7f3826c333013268ff7abb

Observation b26a325b-b3e5-4786-b35b-a585af4cfed3 · inbound

O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning cites this paper.

O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:09:06.805981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:09:06.805981Z digest=sha256:4f9d1f31e0aab2b35227494b99c874df8ced08a86ac3a62804f54e3d9aee7d46

Observation 0aa02464-c6fb-408b-9cb4-382731f3dc7e · inbound

Online Preference Alignment for Language Models via Count-based Exploration cites this paper.

Online Preference Alignment for Language Models via Count-based Exploration SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.508364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.508364Z digest=sha256:fb58416fa3787ea9db652dede3664c4a1925139c742c239852f5431abe214e6c

Observation 378b2aeb-276e-4b79-a9fd-fc4fb01b0e8e · inbound

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation cites this paper.

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:33:49.329189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:33:49.329189Z digest=sha256:f17329dea7184f71f4c4d4e76537013888e68de296e4fb5f77ca6b3111daa9ba

Observation 9972eedc-141b-4155-9a06-839c421d7be7 · inbound

Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game cites this paper.

Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:20:07.170523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:20:07.170523Z digest=sha256:c0f7e63af6a09e48cd1ef213404c257c2ed35b7a9f7836a49f05f360b5a99d63

Observation 465f07ad-38d8-4371-99b1-cdbee1363ed9 · inbound

Controllable Protein Sequence Generation with LLM Preference Optimization cites this paper.

Controllable Protein Sequence Generation with LLM Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:51.828305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:51.828305Z digest=sha256:606c33f229e6144f5007c452558971d1d628b9e92522c028c616c978f6929dc9

Observation d7dacf60-509c-4f74-a326-58e84b29d054 · inbound

Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning cites this paper.

Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:39:22.645558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:39:22.645558Z digest=sha256:820ac98de8bbfb3b9abb5b76ea98c5927f4ff601a29d65c13538b11de8c2ec15

Observation e2c1ff15-ee4f-41cd-ab9d-1f8ae7849bfa · inbound

R.I.P.: Better Models by Survival of the Fittest Prompts cites this paper.

R.I.P.: Better Models by Survival of the Fittest Prompts SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T23:02:23.364165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T23:02:23.364165Z digest=sha256:e2ff33b0016d9557606a6a226ff37ced2b115aff13bdc79536b05207a3614695

Observation 458f3f4d-8e64-4e03-a712-f8cf3b0258cb · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.758075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.758075Z digest=sha256:6edb7630e7303e5418dbc04f0e3447bac11978b724a386a033aa51a3c5c8e623

Observation 27b6e860-e0be-41eb-967f-f367df79c73a · inbound

Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment cites this paper.

Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T19:56:28.740814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:56:28.740814Z digest=sha256:6e4ebc8eb659fb9a7ffd63326f26ba2d1b4cf6967df52112cc3a3497d79bfb41

Observation 6904886a-658d-48c7-b0a1-f449ed55e4b3 · inbound

Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? cites this paper.

Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T18:11:28.152093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:11:28.152093Z digest=sha256:714c779d19f4fd60806ce4d0d35fcd0c0430d984039e5aa604d7f493199619b0

Observation 3a31901d-1717-48a5-9922-8abf75f6bc12 · inbound

The Differences Between Direct Alignment Algorithms are a Blur cites this paper.

The Differences Between Direct Alignment Algorithms are a Blur SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:52:29.437601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-23T03:50:03.720389Z digest=sha256:f20b7caef5ad361257e3975421c4357d9cc42dc550384aca4202d9320339273a

Observation 5ff434a2-881c-4014-822f-0c9e403eee70 · inbound

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing cites this paper.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.070762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.070762Z digest=sha256:ce1b1520370e04dd724e734796d5a3263360754682abfa22b4a2dcb8324df79c

Observation f691c706-5b33-4605-8d5a-35b8db6c267b · inbound

Mol-LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization cites this paper.

Mol-LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:06:14.348926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:06:14.348926Z digest=sha256:4cc729e7bf608659ec3dc0b6e3f1f3c3faafdd333f5458d7c823e9be00c80475

Observation 6970cff7-c592-4727-9b48-150a65477d59 · inbound

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms cites this paper.

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T06:04:01.620696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T06:04:01.620696Z digest=sha256:84d42d0379140b37fb7821e03e39757c1ef819e6e237df0aaeb246d596fd282a

Observation e389116f-3b2e-4729-a77e-c34c32e31f75 · inbound

On Fairness of Unified Multimodal Large Language Model for Image Generation cites this paper.

On Fairness of Unified Multimodal Large Language Model for Image Generation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T04:52:40.587630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:52:40.587630Z digest=sha256:bde194b6b23d413e631ec8008b0fc92aa5da6ec7ed5c9b5112fac635a8605560

Observation bde4c4a1-91de-42ed-ac10-3695cbf60e06 · inbound

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective cites this paper.

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T04:08:51.741471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:08:51.741471Z digest=sha256:f155c5a0ba9666fcd2dce0473f6a9ace3ec6b4d1b5b20d343b1b2a89e705a744

Observation afa58b2c-fcc2-4aad-b839-c1d5bb82114a · inbound

ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization cites this paper.

ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:57:12.721732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:57:12.721732Z digest=sha256:553d2ed042e28d3c366b907ee942a63d02781b38926cc1073683a9192897a4ed

Observation b3f3a684-9bf5-43ce-957c-1282e2e0a2ad · inbound

DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization cites this paper.

DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T06:04:10.837435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T06:04:10.837435Z digest=sha256:3c2cfeab64e105d33b19e3ecbff70c530a1531026cf691bb8c19b56ee38b1798

Observation 62278bdc-0074-46f9-b013-fb0fbe958c65 · inbound

PerPO: Perceptual Preference Optimization via Discriminative Rewarding cites this paper.

PerPO: Perceptual Preference Optimization via Discriminative Rewarding SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T06:01:13.315930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T06:01:13.315930Z digest=sha256:1708eac6dcae5bba3f43cb4a5dd74eb5c3a389ec0966ee35f7a70e2fa188de72

Observation 092edd43-0811-4cc9-81c1-8e01fbf5081a · inbound

Verifiable Format Control for Large Language Model Generations cites this paper.

Verifiable Format Control for Large Language Model Generations SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T22:36:35.280241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:36:35.280241Z digest=sha256:bb8e9ac21c22c2787a29a1773251e82b1309f59a4dccc44252d8230f8393a597

Observation dad8a705-5d0e-4ce1-bad6-2c9cfea9a834 · inbound

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator cites this paper.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.484390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.484390Z digest=sha256:d69510ce2c13537597c221f20a79b02a2fbeaef2e59d290a468d6b4b3baab58c

Observation 01f88fa3-c77b-44a4-9d1a-f45e326b74e3 · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.406513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.406513Z digest=sha256:d4f60a8332ce37640c84afe625c41184cff6cd5e0ccf386c0f7dd27336d878f7

Observation caabea31-3e40-41bf-b61e-2ab71ce38f96 · inbound

Design Considerations in Offline Preference-based RL cites this paper.

Design Considerations in Offline Preference-based RL SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.014455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.014455Z digest=sha256:3d7d918fdae4b17966b40362f93944d2a68c906238ee2a791f87d2a48fb825d5

Observation 5aa3f7ec-4432-493e-98b6-60b849829c99 · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.928601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.928601Z digest=sha256:b6365eb51ad582cd849d32832f89b335e40482145561bdd13ba4ddf49bda9aa6

Observation 153236cf-b82e-4d69-88b7-d80305e92d13 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.525929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.525929Z digest=sha256:ae9908945f244ffb882f3d500bb1e862f313df7700e26b8c98d68d80d9d3549c

Observation 7316f5ab-76ef-45b4-a826-b3a597abcc83 · inbound

A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents cites this paper.

A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:47:11.196787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:47:11.196787Z digest=sha256:7f81c1afc55d3b7a385418cb946ff96a1589d88865d67c334f489bb8fb6531e1

Observation 125a8a78-1c1b-4aac-af6b-4f56387f5b13 · inbound

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization cites this paper.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.859960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.859960Z digest=sha256:b50e389501bc4fef68577ababf4513c8cbddda2f66e854d680e4210fb745fc1f

Observation a3ab10ac-f381-42a0-9f3a-b2659842781f · inbound

Parameter-Efficient Checkpoint Merging via Metrics-Weighted Averaging cites this paper.

Parameter-Efficient Checkpoint Merging via Metrics-Weighted Averaging SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:33.628621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:08:33.628621Z digest=sha256:4c5ce2b408ae15a122f837229240efbd6cb8a0cca00d0b7ebbe17c9d214a9a52

Observation f6f4e601-c873-47e2-aba7-6faa4dfdabbf · inbound

Group Relative Knowledge Distillation: Learning from Teacher's Relational Inductive Bias cites this paper.

Group Relative Knowledge Distillation: Learning from Teacher's Relational Inductive Bias SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:32:08.009827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:32:08.009827Z digest=sha256:ad7f173a0d721c59b29b4a8c3de06aba548d93b99552d80b9a0dea5da7685a67

Observation bde7a4df-48c9-47f2-b78e-fd2143ecb272 · inbound

Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes cites this paper.

Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:47.084128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:47.084128Z digest=sha256:d8f6d21f28155ebccbda5bb597da9513c844c99d97c54299531ccbf34927a143

Observation e416e5e9-a6b0-4958-b35c-d82664df16a0 · inbound

Policy-labeled Preference Learning: Is Preference Enough for RLHF? cites this paper.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.584826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.584826Z digest=sha256:21fe0279af234c161418e58b959224e8c88b9512afb08aff00861a7d7d2163ba

Observation dd1d3041-7ae2-49bc-8346-b91dad26bed4 · inbound

InfoPO: On Mutual Information Maximization for Large Language Model Alignment cites this paper.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.027666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.027666Z digest=sha256:090387caf619a51ae135793fe0b297f430f806b66a6fb6bcf3fa9212918d982e

Observation 3d20fae7-511d-4bec-9a89-981bd0db9232 · inbound

Preference Optimization for Combinatorial Optimization Problems cites this paper.

Preference Optimization for Combinatorial Optimization Problems SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:55:41.000172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:55:41.000172Z digest=sha256:c5fbf07775a7c7689e9ae9fc5971252199e9d89aafbfb4de09fc61c0b20a913b

Observation ee8ee31d-38fc-4f4e-a2a0-dde7bd755e87 · inbound

ShiQ: Bringing back Bellman to LLMs cites this paper.

ShiQ: Bringing back Bellman to LLMs SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.547830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.547830Z digest=sha256:67156b9f8481d8ea4621bc54098501c9a5abbb20b79ecf8fadc601b8440db1dc

Observation 6756cf52-6e94-4b84-9ade-83efaeb08cc7 · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:29.123942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:29.123942Z digest=sha256:1c2948e18eeab6cdac9ba4213771c97e80339189e216ae1b0cd75176146c25e4

Observation 0b3a32fe-32a4-497d-8e53-5d3eeef18190 · inbound

Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals cites this paper.

Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:39:51.729438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:39:51.729438Z digest=sha256:97777dba256ace4920544fde65dcfd90ef56545f50cd3adb27742831f6b61609

Observation 6748e5cc-8e5e-46e6-8f2f-fd9d81909661 · inbound

Risk-aware Direct Preference Optimization under Nested Risk Measure cites this paper.

Risk-aware Direct Preference Optimization under Nested Risk Measure SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.460292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.460292Z digest=sha256:56ef058d42884a4fe76241c2e84e8eee1db7bde2cc9c8bdc4ee7556c021325eb

Observation 83135dab-87ed-40d9-ab41-8082ed5c4dbc · inbound

Improved Representation Steering for Language Models cites this paper.

Improved Representation Steering for Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:30.846708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:30.846708Z digest=sha256:561d4b9e454a0a20fd67d9d7c0c1ef1b507ec7d600ecf00a82324682c7e3a0db

Observation 351d1a80-924b-4edd-b9c1-fef3035bda62 · inbound

LPOI: Listwise Preference Optimization for Vision Language Models cites this paper.

LPOI: Listwise Preference Optimization for Vision Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:55.878822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:55.878822Z digest=sha256:d483882cca21a22501a334fbb7daa5f123d207289fcd0f3eb4c05ded6236328f

Observation 506e6805-bec1-42ef-8741-9021d8de2142 · inbound

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning cites this paper.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:30.952675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:30.952675Z digest=sha256:a711303a290ed1082ae68bdeb07cbc71429a8246c6dad42f8513c184e0439ea0

Observation 3dd5bdc4-8a04-4b06-b302-602038d8cdcf · inbound

K-order Ranking Preference Optimization for Large Language Models cites this paper.

K-order Ranking Preference Optimization for Large Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.152851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.152851Z digest=sha256:e0711923b62db322dd72e75feab42d363542c71ce311e352db54c1cc900c1ff1

Observation 6f96841a-92ea-459f-9b7d-fa100a4db63b · inbound

Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising cites this paper.

Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:57:42.311077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:57:42.311077Z digest=sha256:118a50d74cbf557009bd695936f7364c1c70e69e5fc9d51f1c9c6dd2fd84618b

Observation e1b6d9c6-e1fe-4a0a-a63f-f42a2a325baa · inbound

Aligning Large Language Models with Implicit Preferences from User-Generated Content cites this paper.

Aligning Large Language Models with Implicit Preferences from User-Generated Content SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.094128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.094128Z digest=sha256:2827006f8181f5e573bb9f50a6df14608d94212717451cc1d5319c009fa93c89

Observation 7e9011be-a790-4f63-91ec-4a4a6320fcfe · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:48.465974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:48.465974Z digest=sha256:c0aef5c584ac961a3113c57520000da8ec9784b04fecaab63136e44703dceed6

Observation 3eaad74c-b477-4f61-ad2d-2514ce7af7ab · inbound

Explicit Preference Optimization: No Need for an Implicit Reward Model cites this paper.

Explicit Preference Optimization: No Need for an Implicit Reward Model SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.949816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.949816Z digest=sha256:5cd6b480e20dbdef0276623a4d476f7b7a910e63a30bb666af3b328d95853df7

Observation 17e0d335-3124-4d7f-94a1-32fe35b845e5 · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.373279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.373279Z digest=sha256:f06bd30248c93fca9b23b8934e51592c3f3643a930a690866a4cf19d99418d70

Observation 50ebd32b-b888-4368-9177-dd1c1602f3a5 · inbound

Improved Supervised Fine-Tuning for Large Language Models to Mitigate Catastrophic Forgetting cites this paper.

Improved Supervised Fine-Tuning for Large Language Models to Mitigate Catastrophic Forgetting SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:52:56.039373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:52:56.039373Z digest=sha256:8dc27b4e588efac7dd01b3f556624a9def669b227dba59f42bb282271d703840

Observation a5119672-470d-427a-b83d-82a4addb89ff · inbound

Data Diversification Methods In Alignment Enhance Math Performance In LLMs cites this paper.

Data Diversification Methods In Alignment Enhance Math Performance In LLMs SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:27.665291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:43:27.665291Z digest=sha256:d8e24434748a505676cd27d1cf809beca209731f5fc8dbb8cd0c99f072801000

Observation 3c2eb424-7789-4bbc-bfe9-5ac4c44c0405 · inbound

CTR-Guided Generative Query Suggestion in Conversational Search cites this paper.

CTR-Guided Generative Query Suggestion in Conversational Search SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:54.261257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:54.261257Z digest=sha256:2ca50645233ad5f3b8c6ba66ec54c379791d224df6fb1f39967f454ba7636506

Observation 3fd08862-22af-4123-86d5-2c1132d1fd50 · inbound

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization cites this paper.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.166687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.166687Z digest=sha256:d50265aa5a19455cac7d7fcb1c9cae122a8744497fd2e8bfc2d3005fb540fef0

Observation c272cd5d-6371-4d41-b9ee-39ecfea18ff8 · inbound

A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities cites this paper.

A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.818055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.818055Z digest=sha256:db447651242c455e78108301cc1ab31414f4f1a2a7924a3130d3d384e2afa938

Observation 42a7accf-0fdd-4f7d-904e-d23cf86670f1 · inbound

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems cites this paper.

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.881585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T04:37:33.928616Z digest=sha256:3ad799a473750f2b688e4fca65455b6c6d516d16a18523005e96815d4255bdff

Observation 155970b4-1ee4-40bf-8aac-aaf1bc4814e9 · inbound

Unlearning of Knowledge Graph Embedding via Preference Optimization cites this paper.

Unlearning of Knowledge Graph Embedding via Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:03.154242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:03.154242Z digest=sha256:24482d80dd1fc437efcb6486957e782e5997c873b502a50cea956c34453ea3e3

Observation d0b3af2a-d5aa-4a63-8a1e-37716d4fdfa1 · inbound

SDD: Self-Degraded Defense against Malicious Fine-tuning cites this paper.

SDD: Self-Degraded Defense against Malicious Fine-tuning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.881021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.881021Z digest=sha256:b56c3254df3001c33b46afeda7ee3704740de600bab2aed8d51fe56ffec89484

Observation 4e249eb3-adc2-4642-ab29-e7e659477e76 · inbound

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap cites this paper.

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.560183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T23:46:24.208438Z digest=sha256:5cfd9c9bbb7d0ec8cdf88171dde653eb7a010b69f0fde4a53a1c94eba74ff0f8

Observation 423725fe-c6b9-420e-8742-329ebdf318c6 · inbound

FormaRL: Enhancing Autoformalization with no Labeled Data cites this paper.

FormaRL: Enhancing Autoformalization with no Labeled Data SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T16:11:23.181357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:11:23.181357Z digest=sha256:b32f09376108642320f228e8f6f3dad362ef411be91b3797533ae4d3f800b929

Observation 08de6af2-252b-4bd2-ac16-6ea088802325 · inbound

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework cites this paper.

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:35.355296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:35.355296Z digest=sha256:c5815f1acb891647a2b87a3d5dc20a9bfd2f013c23917ffa43016d726bce5f35

Observation 2e92be31-7c7f-4e0c-8035-366addc3ef44 · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.432883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.432883Z digest=sha256:d4ac579f40331882c33b786ad3718f0d6443eeb9c6b67114aeb7afa3bfb18e48

Observation 844276d1-41a8-416a-9280-a9e2de82f034 · inbound

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation cites this paper.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.197031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.197031Z digest=sha256:0feae2a3e7ae261d5dca98bc129ca0124f6decb62714f2260cde03bcb2bc20e6

Observation 47391bd2-a653-478f-81d3-1b46faa0368a · inbound

Adaptive Preference Optimization with Uncertainty-aware Utility Anchor cites this paper.

Adaptive Preference Optimization with Uncertainty-aware Utility Anchor SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:56.685115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:56.685115Z digest=sha256:15a44f53e9516a615d713032b6c016ee9c969c12dbdd0776d9729f34c2fdbe48

Observation 06e79e23-3c1a-47d9-94e9-dc33394d7499 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.905080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:4dc0ae2fddc71f655e03da391d7336bef65403c36244aeaa1c2a401ee14e22a5

Observation 79534ca0-fd56-4a18-bcc4-64041dc21158 · inbound

Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts cites this paper.

Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T18:50:00.510534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:50:00.510534Z digest=sha256:637365df25bebc73fdf0a88322e7c028b74d94d7da797f34b9f59039a1cf1674

Observation 009329e7-8c10-4483-aa89-bd2d50c491fa · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 150

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:25.513667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:25.513667Z digest=sha256:4b4939929bb56330a70800bac731390c298c3de6352ba3d6f7fd38012380950f

Observation cb1bff44-c5c2-4ff6-9103-d2dacdc9e0a6 · inbound

DDO-RM: Distribution-Level Policy Improvement after Reward Learning cites this paper.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:03.613104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:baa5e7637fd2a970b13c8a84d154f69d83d8bead2d40d6c19c84118553c9ea67

Observation 77c539dc-b714-4a85-8f06-965ef7ee6cbf · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.274742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:c2d26c72e89f9b7bf6e9aed3c6dec71752ea4e293f64f98dfa0977a4e6b16254

Observation 250db052-d982-4bac-9648-c592bb09cc85 · inbound

Bayesian Rate Inference for Sequence Motif Dynamics in Systems of Reactive Nucleic Acids cites this paper.

Bayesian Rate Inference for Sequence Motif Dynamics in Systems of Reactive Nucleic Acids SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.414498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T08:44:22.449452Z digest=sha256:8be37a2466bb4c798e87a0d7f65bf84e525e8aa6ac4aabdabb819d79f01eaa58

Observation d709202f-d817-4d29-a000-a1ef074ccb9f · inbound

Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation cites this paper.

Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:16.638859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T16:30:11.398851Z digest=sha256:8b15a67477980efc53db5f85b7fcbb0a653c51fde941886a149b863bb3215459

Observation 69cfce6b-6145-4de9-b2d8-65af09009508 · inbound

Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation cites this paper.

Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T05:23:57.330115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:23:57.330115Z digest=sha256:de94c1adf0c9bf0011b37d614e801dd307fef4e54ab5c6dc241beb6558b609b9