Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:07:20.866024Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2412.07812.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:07:20.866024Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9caeb701-2330-4e05-a341-e754f1a175f0 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Attention is all you need
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 613624ee-1b39-4b9a-92a7-0420c29e434b · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Language Models are Few-Shot Learners
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2b02f2b-94b1-49cd-95c0-355a794cfb9e · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset LLaMA: Open and Efficient Foundation Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97c807b5-264a-4f60-b736-2aa27392028f · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8644ba4c-c802-4924-a782-106bcfe14283 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Hashimoto
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 358f491b-c243-4e2a-a117-9e52b4fa6aae · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Training language models to follow instructions with human feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd713223-8f1a-4ff4-8894-678e4f6aff7d · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Direct preference optimization: Your language model is secretly a reward model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bb7314b-e3fb-4ee2-aa82-879df548e19a · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e1d9e01a-83f7-4a64-a874-e49cf931686a · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Rank analysis of incomplete block designs: I
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c50258-0391-49d3-8171-609534ff4ff2 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Aligning Language Models with Preferences through f-divergence Minimization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d717c9d-f1e8-43d0-9949-8dfa75233328 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 91585a72-b80e-4fff-8fb1-2eb24feebaee · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8548cd89-6def-42b2-8c7e-436a66d4991f · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Reinforcement learning by reward-weighted regression for operational space control
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 85d301e9-ea8d-4477-a9b3-ee73db5c6136 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 313c4c64-c3f5-4086-a3c6-a8544adced1f · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Self-Rewarding Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7426cb66-84b2-4892-804f-ce2e0bec140a · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Preference ranking optimization for human alignment
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7dcfaf01-ec37-4599-bcd5-aab0accc7604 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Aligning large language model with direct multi-preference optimization for recommendation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation db3b36ae-6137-450a-877b-b7b68b31c6b8 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset orca_dpo_pairs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e6ccc5b0-3c7e-4723-b0a0-bf3b355a0096 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Orca: Progressive learning from complex explanation traces of gpt-4, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06d4da0-5c14-436f-b6d9-f9cae2dfed13 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6727376-ea2f-4f43-9981-774ad46a2789 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Mistral 7B
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 641784d1-7b41-4cb6-a44b-aa5b6b7b7b1b · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f906b717-4657-461c-a7ad-4ebd4ae96a25 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3dc527c-7859-4971-bd79-27d21c33c307 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Hashimoto
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2179e772-8dac-4f21-89c2-3939ea704843 · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset Hashimoto
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f34206cd-5e88-44d1-aa28-d8f76d6ac1cc · outbound
Multi-Response Preference Optimization with Augmented Ranking Dataset sample instruction,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.