Pith. sign in

Paper Citation Record · LEDGER

Human Alignment of Large Language Models through Online Preference Optimisation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2403.08635.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.08635 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:37.284453Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T13:11:24.028811Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e57a0052-665a-4d6a-a9bf-e207e09a7be3 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models Human Alignment of Large Language Models through Online Preference Optimisation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:58:16.898151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:7f1978777f8ea093174c91000c966f81b103a4be71e2e34af888abf72c89e84b

Observation 15b86385-82d0-4368-b077-1c142b2587e1 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought Human Alignment of Large Language Models through Online Preference Optimisation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:37.284453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:37.284453Z digest=sha256:5a37e06ab9aeab7921d082edbbba34bda15d87537935a0fb06fda9c94cbc03ee

Observation 90faa925-f9ee-4c16-8373-4b82a5ff6928 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Human Alignment of Large Language Models through Online Preference Optimisation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.031523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:493d61db17f95339b52c716ae48df525d1f38927961a4313384f23d7ac3a4c6b

Observation 95ae58a7-83aa-40ae-bf8c-5e961fcbf0ce · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Human Alignment of Large Language Models through Online Preference Optimisation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:24.182186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:24.182186Z digest=sha256:ec172695d79029a0b3b8eee7ff53e40479d54c261f3b2e07fb875dbc1f40debc

Observation 7e042000-3c1a-486f-a564-7ec1890f3e78 · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning Human Alignment of Large Language Models through Online Preference Optimisation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:26.503032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T04:15:24.919355Z digest=sha256:224639b87589ff864a353fcca38c49b3231f36b40a372229d00e0b5f4939cac2

Observation 0f7ab43a-c49f-4647-a192-6e4db481869c · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning Human Alignment of Large Language Models through Online Preference Optimisation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:05.096769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T22:03:05.102274Z digest=sha256:6a4419f6170211ffdf6731a8881156b8b0801736c2f542520831c15e179518b0