Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2312.00886.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T22:08:56.617012Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T12:06:55.872775Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 148b4597-b560-4a14-9dee-189fc89bff1f · inbound
KTO: Model Alignment as Prospect Theoretic Optimization Nash Learning from Human Feedback
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 526fd740-085f-4992-a9f9-b3d9e4eebd52 · inbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Nash Learning from Human Feedback
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 15b61f5f-8df9-4c0e-8730-d495dde7959e · inbound
How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Nash Learning from Human Feedback
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 30aadae7-2443-4c38-9558-65ab51d5d6dd · inbound
Incentivizing High-Quality Human Annotations with Golden Questions Nash Learning from Human Feedback
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a3cd3657-fa64-4176-b4ce-dc9a7c65ebfc · inbound
Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs Nash Learning from Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b45b80-0383-4fa7-8acb-c871fe844295 · inbound
Multiplayer Nash Preference Optimization Nash Learning from Human Feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0c981e19-615a-494e-ace3-f7193d967e00 · inbound
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization Nash Learning from Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1af94db-53b0-4dd8-9598-0ae13b266a4f · inbound
Safety Alignment of LMs via Non-cooperative Games Nash Learning from Human Feedback
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53a01c0a-9734-478d-ab79-e241af13b989 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Nash Learning from Human Feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0bde6246-f87d-4b37-868a-439a6fc25875 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Nash Learning from Human Feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9897e36d-a351-40ec-a2b0-45164887c80c · inbound
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment Nash Learning from Human Feedback
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b7285d33-29b8-44e5-908d-fe7e2778d833 · inbound
Response Time Enhances Alignment with Heterogeneous Preferences Nash Learning from Human Feedback
Reference 272
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8342c205-b8dd-4cb3-8264-27178e3552fb · inbound
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Nash Learning from Human Feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04463d32-a51b-422d-b73c-6a22e62e7938 · inbound
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Nash Learning from Human Feedback
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d0ebe70e-8901-4f02-8446-3fcfbe42e772 · inbound
Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Nash Learning from Human Feedback
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19cc1741-87fe-4ae5-8d02-0f82ff8b90a8 · inbound
Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Nash Learning from Human Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8106f321-7694-4ef0-89b1-a7aacf6a5055 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Nash Learning from Human Feedback
Reference 186
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fed52fb-1125-4a62-ab4c-a71d9d2971f5 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Nash Learning from Human Feedback
Reference 187
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59004412-9117-4cb3-bc73-ab0efe768ff1 · inbound
Probably Correct Optimal Stable Matching under Two-Sided Uncertainty Nash Learning from Human Feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.