Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2405.14734.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.525929Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T20:05:34.173711Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 63664023-8e42-4bd5-875a-d32642e75721 · inbound
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation baad56b5-f9b2-45f3-81fc-d90330cbd6a9 · inbound
DataComp-LM: In search of the next generation of training sets for language models SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 22771146-9716-4e82-99a5-eb2e0f3b1d5b · inbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18980d1f-0ac9-4ffb-9fa1-b91d882aea71 · inbound
Adaptive Decoding via Latent Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e7b513-cee4-435a-b139-692ff894730e · inbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c5ec60-efd9-4ef5-a91c-7304256c0429 · inbound
ProSec: Fortifying Code LLMs with Proactive Security Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c071df3-41bb-493c-a256-f81fcce82abc · inbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c46ffce3-d532-42ed-a13c-dea5659f2527 · inbound
ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 265d9257-8f0d-4d7f-bb49-075d279d7e15 · inbound
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eb18100c-eed6-430b-a3fd-c97b21c86cb0 · inbound
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c1eeab3-330e-41aa-b0d6-4a8c02f8e333 · inbound
T-REG: Preference Optimization with Token-Level Reward Regularization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 802c7432-bfa4-4aa4-bf4b-795b6f5085a5 · inbound
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e2d84c-c9d5-4597-acfe-580aae8fcb3f · inbound
ALMA: Alignment with Minimal Annotation SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c91929f-c41c-443e-bb96-11fdea8b1140 · inbound
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 026e6fdb-1a4c-42a8-9602-03a242430c1b · inbound
Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2942322d-55b3-4fc2-bc61-621f284992ed · inbound
CleanComedy: Creating Friendly Humor through Generative Techniques SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f16a6eec-9670-4ec8-95b9-68c29bf7a740 · inbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24f63fe-024a-40dc-b521-c3f89c9a7e50 · inbound
WEPO: Web Element Preference Optimization for LLM-based Web Navigation SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ada8865-7075-4e2d-ac37-d98dea9d8925 · inbound
NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b68d6a-a06d-47dc-9a3e-a1628785fa47 · inbound
Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b26f855-819d-4343-aa5e-9c62415b335f · inbound
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8489d4b5-97ab-47f5-9741-966797a17cf8 · inbound
Hansel: Output Length Controlling Framework for Large Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0943c1fb-e0db-48ed-9cf5-9f167b4771ef · inbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22304670-4771-4e4c-b3e7-b02ca9d3d064 · inbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e34296a-7701-469c-a51c-ec0dffb8eec0 · inbound
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a48d06eb-8497-4b7b-899f-0363ca074d8b · inbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e84489-14ee-4766-9d38-dfb01f313e87 · inbound
GAS: Generative Auto-bidding with Post-training Search SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1395c49a-5a41-4d11-bc37-eb62a3e543cd · inbound
Multi-Agent Sampling: Scaling Inference Compute for Data Synthesis with Tree Search-Based Agentic Collaboration SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb8b9a2-dfad-482d-80c0-6fe951c4ab8f · inbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e8f2b64-e06a-4f87-b95e-65de95d3453b · inbound
Plug-and-Play Training Framework for Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3a5cdb-fec5-4abc-ab22-a51ad16235c2 · inbound
Prune 'n Predict: Optimizing LLM Decision-making with Conformal Prediction SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e5657a-153e-49d4-8ab5-6b8a297d1cb5 · inbound
From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26a325b-b3e5-4786-b35b-a585af4cfed3 · inbound
O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aa02464-c6fb-408b-9cb4-382731f3dc7e · inbound
Online Preference Alignment for Language Models via Count-based Exploration SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378b2aeb-276e-4b79-a9fd-fc4fb01b0e8e · inbound
CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9972eedc-141b-4155-9a06-839c421d7be7 · inbound
Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 465f07ad-38d8-4371-99b1-cdbee1363ed9 · inbound
Controllable Protein Sequence Generation with LLM Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7dacf60-509c-4f74-a326-58e84b29d054 · inbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2c1ff15-ee4f-41cd-ab9d-1f8ae7849bfa · inbound
R.I.P.: Better Models by Survival of the Fittest Prompts SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 458f3f4d-8e64-4e03-a712-f8cf3b0258cb · inbound
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b6e860-e0be-41eb-967f-f367df79c73a · inbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6904886a-658d-48c7-b0a1-f449ed55e4b3 · inbound
Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a31901d-1717-48a5-9922-8abf75f6bc12 · inbound
The Differences Between Direct Alignment Algorithms are a Blur SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5ff434a2-881c-4014-822f-0c9e403eee70 · inbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f691c706-5b33-4605-8d5a-35b8db6c267b · inbound
Mol-LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6970cff7-c592-4727-9b48-150a65477d59 · inbound
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e389116f-3b2e-4729-a77e-c34c32e31f75 · inbound
On Fairness of Unified Multimodal Large Language Model for Image Generation SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde4c4a1-91de-42ed-ac10-3695cbf60e06 · inbound
LLM Alignment as Retriever Optimization: An Information Retrieval Perspective SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa58b2c-fcc2-4aad-b839-c1d5bb82114a · inbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3f3a684-9bf5-43ce-957c-1282e2e0a2ad · inbound
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62278bdc-0074-46f9-b013-fb0fbe958c65 · inbound
PerPO: Perceptual Preference Optimization via Discriminative Rewarding SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092edd43-0811-4cc9-81c1-8e01fbf5081a · inbound
Verifiable Format Control for Large Language Model Generations SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dad8a705-5d0e-4ce1-bad6-2c9cfea9a834 · inbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f88fa3-c77b-44a4-9d1a-f45e326b74e3 · inbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caabea31-3e40-41bf-b61e-2ab71ce38f96 · inbound
Design Considerations in Offline Preference-based RL SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa3f7ec-4432-493e-98b6-60b849829c99 · inbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 153236cf-b82e-4d69-88b7-d80305e92d13 · inbound
Generative AI Act II: Test Time Scaling Drives Cognition Engineering SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 231
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7316f5ab-76ef-45b4-a826-b3a597abcc83 · inbound
A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125a8a78-1c1b-4aac-af6b-4f56387f5b13 · inbound
AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ab10ac-f381-42a0-9f3a-b2659842781f · inbound
Parameter-Efficient Checkpoint Merging via Metrics-Weighted Averaging SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6f4e601-c873-47e2-aba7-6faa4dfdabbf · inbound
Group Relative Knowledge Distillation: Learning from Teacher's Relational Inductive Bias SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde7a4df-48c9-47f2-b78e-fd2143ecb272 · inbound
Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e416e5e9-a6b0-4958-b35c-d82664df16a0 · inbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd1d3041-7ae2-49bc-8346-b91dad26bed4 · inbound
InfoPO: On Mutual Information Maximization for Large Language Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d20fae7-511d-4bec-9a89-981bd0db9232 · inbound
Preference Optimization for Combinatorial Optimization Problems SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee8ee31d-38fc-4f4e-a2a0-dde7bd755e87 · inbound
ShiQ: Bringing back Bellman to LLMs SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6756cf52-6e94-4b84-9ade-83efaeb08cc7 · inbound
MPO: Multilingual Safety Alignment via Reward Gap Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b3a32fe-32a4-497d-8e53-5d3eeef18190 · inbound
Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6748e5cc-8e5e-46e6-8f2f-fd9d81909661 · inbound
Risk-aware Direct Preference Optimization under Nested Risk Measure SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83135dab-87ed-40d9-ab41-8082ed5c4dbc · inbound
Improved Representation Steering for Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 351d1a80-924b-4edd-b9c1-fef3035bda62 · inbound
LPOI: Listwise Preference Optimization for Vision Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 506e6805-bec1-42ef-8741-9021d8de2142 · inbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dd5bdc4-8a04-4b06-b302-602038d8cdcf · inbound
K-order Ranking Preference Optimization for Large Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f96841a-92ea-459f-9b7d-fa100a4db63b · inbound
Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b6d9c6-e1fe-4a0a-a63f-f42a2a325baa · inbound
Aligning Large Language Models with Implicit Preferences from User-Generated Content SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e9011be-a790-4f63-91ec-4a4a6320fcfe · inbound
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eaad74c-b477-4f61-ad2d-2514ce7af7ab · inbound
Explicit Preference Optimization: No Need for an Implicit Reward Model SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17e0d335-3124-4d7f-94a1-32fe35b845e5 · inbound
A Survey on Large Language Models for Mathematical Reasoning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50ebd32b-b888-4368-9177-dd1c1602f3a5 · inbound
Improved Supervised Fine-Tuning for Large Language Models to Mitigate Catastrophic Forgetting SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5119672-470d-427a-b83d-82a4addb89ff · inbound
Data Diversification Methods In Alignment Enhance Math Performance In LLMs SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c2eb424-7789-4bbc-bfe9-5ac4c44c0405 · inbound
CTR-Guided Generative Query Suggestion in Conversational Search SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd08862-22af-4123-86d5-2c1132d1fd50 · inbound
Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c272cd5d-6371-4d41-b9ee-39ecfea18ff8 · inbound
A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42a7accf-0fdd-4f7d-904e-d23cf86670f1 · inbound
Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 155970b4-1ee4-40bf-8aac-aaf1bc4814e9 · inbound
Unlearning of Knowledge Graph Embedding via Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b3af2a-d5aa-4a63-8a1e-37716d4fdfa1 · inbound
SDD: Self-Degraded Defense against Malicious Fine-tuning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e249eb3-adc2-4642-ab29-e7e659477e76 · inbound
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 423725fe-c6b9-420e-8742-329ebdf318c6 · inbound
FormaRL: Enhancing Autoformalization with no Labeled Data SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08de6af2-252b-4bd2-ac16-6ea088802325 · inbound
HEAL: A Hypothesis-Based Preference-Aware Analysis Framework SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e92be31-7c7f-4e0c-8035-366addc3ef44 · inbound
Improving Large Vision and Language Models by Learning from a Panel of Peers SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 844276d1-41a8-416a-9280-a9e2de82f034 · inbound
Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47391bd2-a653-478f-81d3-1b46faa0368a · inbound
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e79e23-3c1a-47d9-94e9-dc33394d7499 · inbound
Failure Modes of Maximum Entropy RLHF SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 79534ca0-fd56-4a18-bcc4-64041dc21158 · inbound
Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009329e7-8c10-4483-aa89-bd2d50c491fa · inbound
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 150
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1bff44-c5c2-4ff6-9103-d2dacdc9e0a6 · inbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 77c539dc-b714-4a85-8f06-965ef7ee6cbf · inbound
Representation-Guided Parameter-Efficient LLM Unlearning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 250db052-d982-4bac-9648-c592bb09cc85 · inbound
Bayesian Rate Inference for Sequence Motif Dynamics in Systems of Reactive Nucleic Acids SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d709202f-d817-4d29-a000-a1ef074ccb9f · inbound
Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 69cfce6b-6145-4de9-b2d8-65af09009508 · inbound
Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.