Pith. sign in

Paper Citation Record · LEDGER

Multi-Response Preference Optimization with Augmented Ranking Dataset

As of 22 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2412.07812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07812 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:07:20.866024Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9caeb701-2330-4e05-a341-e754f1a175f0 · outbound

This paper cites Attention is all you need.

Multi-Response Preference Optimization with Augmented Ranking Dataset Attention is all you need

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.771312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.771312Z digest=sha256:68defa95cfc45473597bd409105d0fc96f530b4805f06c72f5cbc981bcf73a13

Observation 613624ee-1b39-4b9a-92a7-0420c29e434b · outbound

This paper cites Language Models are Few-Shot Learners.

Multi-Response Preference Optimization with Augmented Ranking Dataset Language Models are Few-Shot Learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.775640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.775640Z digest=sha256:c196c4b14f8610c2beb094deb161d45cf392b1943c00a6bfac37efa7a7a87812

Observation c2b02f2b-94b1-49cd-95c0-355a794cfb9e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-Response Preference Optimization with Augmented Ranking Dataset LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.779754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.779754Z digest=sha256:75d072994e0f16a72d627b5619e3579dcee3cc581b1a30e20165f2ecfdc73f3b

Observation 97c807b5-264a-4f60-b736-2aa27392028f · outbound

This paper cites an unresolved cited work.

Multi-Response Preference Optimization with Augmented Ranking Dataset Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.783894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.783894Z digest=sha256:d8f2b999ec4b32d738717dbab94356343bbc5993853a1e10c2d2be30ce00c80b

Observation 8644ba4c-c802-4924-a782-106bcfe14283 · outbound

This paper cites Hashimoto.

Multi-Response Preference Optimization with Augmented Ranking Dataset Hashimoto

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.787649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.787649Z digest=sha256:081bdadd5e0bd2a8ed83ed498eb6d78b7bd6b4c51262576558c8f942617e4e36

Observation 358f491b-c243-4e2a-a117-9e52b4fa6aae · outbound

This paper cites Training language models to follow instructions with human feedback.

Multi-Response Preference Optimization with Augmented Ranking Dataset Training language models to follow instructions with human feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.791296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.791296Z digest=sha256:bfd4e01ca9dbe87a95ac871f1f32c8c012297d8cf2acd2577a36bc70bf1c18d5

Observation bd713223-8f1a-4ff4-8894-678e4f6aff7d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Multi-Response Preference Optimization with Augmented Ranking Dataset Direct preference optimization: Your language model is secretly a reward model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.795123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.795123Z digest=sha256:70a359f5bf451d512cd5ce4e80476ce17420b3b68f958ccca961c5415de18c0c

Observation 9bb7314b-e3fb-4ee2-aa82-879df548e19a · outbound

This paper cites an unresolved cited work.

Multi-Response Preference Optimization with Augmented Ranking Dataset Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:07:21.110860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T19:07:20.798359Z digest=sha256:b666e37fdf3d13e8f3254dceb375065381f0fbd567902b208fa46b4b35bea15e

Observation e1d9e01a-83f7-4a64-a874-e49cf931686a · outbound

This paper cites Rank analysis of incomplete block designs: I.

Multi-Response Preference Optimization with Augmented Ranking Dataset Rank analysis of incomplete block designs: I

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.802864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.802864Z digest=sha256:ca340a59c52eebe321fe159f39fab7961082ef5a065e26195e4b0f28cb82d60e

Observation 77c50258-0391-49d3-8171-609534ff4ff2 · outbound

This paper cites Aligning Language Models with Preferences through f-divergence Minimization.

Multi-Response Preference Optimization with Augmented Ranking Dataset Aligning Language Models with Preferences through f-divergence Minimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.806229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.806229Z digest=sha256:bd59b003aad7e3bf92b85d4a208e9dc239974ec3d997f4b28cd4cc3a47c26a71

Observation 6d717c9d-f1e8-43d0-9949-8dfa75233328 · outbound

This paper cites an unresolved cited work.

Multi-Response Preference Optimization with Augmented Ranking Dataset Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:07:21.092716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T19:07:20.809938Z digest=sha256:8c9da34e2efb3c5cc0d741a9e68f77687f50cd0b20ccd6b091718fb35d1c239b

Observation 91585a72-b80e-4fff-8fb1-2eb24feebaee · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Multi-Response Preference Optimization with Augmented Ranking Dataset Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.813062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.813062Z digest=sha256:f91e378895f36faca624925d3e6736cb80452f8057ae89b3babe0b302123d31a

Observation 8548cd89-6def-42b2-8c7e-436a66d4991f · outbound

This paper cites Reinforcement learning by reward-weighted regression for operational space control.

Multi-Response Preference Optimization with Augmented Ranking Dataset Reinforcement learning by reward-weighted regression for operational space control

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:21.082436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T19:07:20.816806Z digest=sha256:519e6feb879c948d5f92ada41314039c74fafa0345683fa082c02bf8098996ce

Observation 85d301e9-ea8d-4477-a9b3-ee73db5c6136 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Multi-Response Preference Optimization with Augmented Ranking Dataset RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.819971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.819971Z digest=sha256:61acdf21db60adbed196078ee5d9a32b703bfe8fe7581cffd873d50bb13e9593

Observation 313c4c64-c3f5-4086-a3c6-a8544adced1f · outbound

This paper cites Self-Rewarding Language Models.

Multi-Response Preference Optimization with Augmented Ranking Dataset Self-Rewarding Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.823371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.823371Z digest=sha256:1a0041b22ac2495b04547a183eb581230df27ff2abba97e67426b137ed31b040

Observation 7426cb66-84b2-4892-804f-ce2e0bec140a · outbound

This paper cites Preference ranking optimization for human alignment.

Multi-Response Preference Optimization with Augmented Ranking Dataset Preference ranking optimization for human alignment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:21.071639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T19:07:20.826942Z digest=sha256:adba20894a301ffc67868db736abf6398ea725f18df721dba01e6c37602fa274

Observation 7dcfaf01-ec37-4599-bcd5-aab0accc7604 · outbound

This paper cites Aligning large language model with direct multi-preference optimization for recommendation.

Multi-Response Preference Optimization with Augmented Ranking Dataset Aligning large language model with direct multi-preference optimization for recommendation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:21.060397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T19:07:20.830611Z digest=sha256:ee216cf7dceb36816432b49333b9e4c39fc3bdb96c3762f749968b537ccacc1f

Observation db3b36ae-6137-450a-877b-b7b68b31c6b8 · outbound

This paper cites orca_dpo_pairs.

Multi-Response Preference Optimization with Augmented Ranking Dataset orca_dpo_pairs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:21.047920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T19:07:20.834311Z digest=sha256:f68c646a50dc2b9b0d46e560709c911981d2cac56d76a700d14596852a5fb204

Observation e6ccc5b0-3c7e-4723-b0a0-bf3b355a0096 · outbound

This paper cites Orca: Progressive learning from complex explanation traces of gpt-4, 2023.

Multi-Response Preference Optimization with Augmented Ranking Dataset Orca: Progressive learning from complex explanation traces of gpt-4, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.837707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.837707Z digest=sha256:5759b5076752f96b7ed0605138e717bf47af06bf69ddc5aefbf6caf42156bd6f

Observation a06d4da0-5c14-436f-b6d9-f9cae2dfed13 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Multi-Response Preference Optimization with Augmented Ranking Dataset RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.841083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.841083Z digest=sha256:f352f1c25b67fda8b424f477dec4b7586fb0a3911bcf915a4f0d76582c7fcb0c

Observation e6727376-ea2f-4f43-9981-774ad46a2789 · outbound

This paper cites Mistral 7B.

Multi-Response Preference Optimization with Augmented Ranking Dataset Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.845171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.845171Z digest=sha256:0c76f0da7919d1e54f890f45af34ca25bbd08fa5413007eca1243fce6ce708a5

Observation 641784d1-7b41-4cb6-a44b-aa5b6b7b7b1b · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

Multi-Response Preference Optimization with Augmented Ranking Dataset Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.849412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.849412Z digest=sha256:a8394d5b5a4607b604a5e705e69542782959dcb9ada8f19107ce54ccbc7946f8

Observation f906b717-4657-461c-a7ad-4ebd4ae96a25 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Multi-Response Preference Optimization with Augmented Ranking Dataset Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.853679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.853679Z digest=sha256:c199f65205b03cdcefa2fb4eb310ed7511c8d4519a846c79daabe222ab7ca8aa

Observation f3dc527c-7859-4971-bd79-27d21c33c307 · outbound

This paper cites Hashimoto.

Multi-Response Preference Optimization with Augmented Ranking Dataset Hashimoto

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:21.021275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T19:07:20.857852Z digest=sha256:a7c324cbedaa3a21063cccf22549fbf9dc5a4e4d7bd174b3bf6e83cf2ce9d758

Observation 2179e772-8dac-4f21-89c2-3939ea704843 · outbound

This paper cites Hashimoto.

Multi-Response Preference Optimization with Augmented Ranking Dataset Hashimoto

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:20.861953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:07:20.861953Z digest=sha256:b1e5e33c915bb486a231dfa23560987b84d39dd90b50f15ee0535bfec456d92f

Observation f34206cd-5e88-44d1-aa28-d8f76d6ac1cc · outbound

This paper cites sample instruction,.

Multi-Response Preference Optimization with Augmented Ranking Dataset sample instruction,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:21.002341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T19:07:20.866024Z digest=sha256:5c6ecffac393da5af0bd3fa736caacf7f30ab6f56a047ee1ce98b9cb55812b7c

Pith citing papers

No inbound Pith citation observations are available.