Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:48:23.200455Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 10 of 10 outbound references and 0 inbound Pith citation observations for arXiv:2509.09055.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:48:23.200455Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
10 of 10 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e05c4997-8370-41ed-bce8-eefa7f93d9e5 · outbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 713eddfd-fe4b-429b-a5e1-ddcdc75da528 · outbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M OPT: Open Pre-trained Transformer Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d66f49-ccf2-4d08-be9e-eaccbf368904 · outbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a73cc95-c9ce-4d0c-a535-f8894b8213b0 · outbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7740676d-a93a-4607-ba59-16987e2cb33d · outbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Realistic Evaluation of Toxicity in Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40651299-5700-4120-a563-dfeaa452eef7 · outbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Crafting Tomorrow's Evaluations: Assessment Design Strategies in the Era of Generative AI
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91451675-098c-4626-8d54-ab661566828f · outbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Dynabench: Rethinking Benchmarking in NLP
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced9a9ae-fd1e-4e8f-b6bd-6c8b837f7175 · outbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Fine-Tuning Language Models from Human Preferences
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4adc5a6b-4463-4caf-831a-8f4ac4bf5faa · outbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Red Teaming Language Models with Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b8aa4c-6825-4719-a95a-f90b96b1c0e0 · outbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.