Pith. sign in

Paper Citation Record · LEDGER

Nash Learning from Human Feedback

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2312.00886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.00886 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:08:56.617012Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:06:55.872775Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 148b4597-b560-4a14-9dee-189fc89bff1f · inbound

KTO: Model Alignment as Prospect Theoretic Optimization cites this paper.

KTO: Model Alignment as Prospect Theoretic Optimization Nash Learning from Human Feedback

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T12:17:53.556390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T12:17:53.478052Z digest=sha256:4da7c9d70581d01913fb3e9076db77e6e5dc483c4588f36ebad9ae65aa8beb41

Observation 526fd740-085f-4992-a9f9-b3d9e4eebd52 · inbound

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types cites this paper.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Nash Learning from Human Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.384210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:77de6e4972a5fce49af9deec0a820fdda768364f3a6fb3bd853ed7f45b96cdd0

Observation 15b61f5f-8df9-4c0e-8730-d495dde7959e · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Nash Learning from Human Feedback

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.556527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:adf5ac6761619462ea89c1bbf7bad82253db86894e4e703d6fe95dbef9f73993

Observation 30aadae7-2443-4c38-9558-65ab51d5d6dd · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions Nash Learning from Human Feedback

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T13:42:19.361562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:4544aa744ef3d8a0f9e4dc4dbf7a2b173455c186c965fedf12dbe9de04b5af57

Observation a3cd3657-fa64-4176-b4ce-dc9a7c65ebfc · inbound

Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs cites this paper.

Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs Nash Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:08:56.617012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:08:56.617012Z digest=sha256:ae3fd0fee986800509ad55a750294ae2cb28d1b2889e49f98742ce29cfc5100f

Observation 60b45b80-0383-4fa7-8acb-c871fe844295 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Nash Learning from Human Feedback

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:11:23.958548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:0956cc37ab718c5dddf44ba63c2a8350bf8c263fd9bf8020a771cc80baea7286

Observation 0c981e19-615a-494e-ace3-f7193d967e00 · inbound

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization cites this paper.

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization Nash Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:57.669889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:57.669889Z digest=sha256:5bfb96e8143ea308d0dc94f2314a09834f6494d6a38690d110d23c0de08b00bd

Observation e1af94db-53b0-4dd8-9598-0ae13b266a4f · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Nash Learning from Human Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.615297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.615297Z digest=sha256:6735d21e9eba1bf6e69bcfc174908e4598d50ab4ff1110c7db838bdf7170a10c

Observation 53a01c0a-9734-478d-ab79-e241af13b989 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Nash Learning from Human Feedback

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.624270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:50443c2fc86ea11ac1d844c5ec6ccd68be06114f73a0633e76ba2167e0b405af

Observation 0bde6246-f87d-4b37-868a-439a6fc25875 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Nash Learning from Human Feedback

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.465569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:2588322baae1492396d618cbfe73f866fcea4ab509ac640bc4756e3753005bad

Observation 9897e36d-a351-40ec-a2b0-45164887c80c · inbound

Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment cites this paper.

Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment Nash Learning from Human Feedback

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:35:49.320786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:43:51.965882Z digest=sha256:67c302000462445d04e775afa3cc0e2c28a755ae8e586bf97061779adece0dbb

Observation b7285d33-29b8-44e5-908d-fe7e2778d833 · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Nash Learning from Human Feedback

Reference 272

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:59.579682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:b8b5cd9f62ccb487a5dbf85d6d999a8445f5b2d4ec6c62c6c7e49dc86711c4ab

Observation 8342c205-b8dd-4cb3-8264-27178e3552fb · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Nash Learning from Human Feedback

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:55.196686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:922ba8ee8c8e7bb4141b344839cdf6606802d6d9d67542da093ba606bb1ed21f

Observation 04463d32-a51b-422d-b73c-6a22e62e7938 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Nash Learning from Human Feedback

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.010410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:d71541fd70257f1aab5d0e5158285b9cbd6a17b19644fe4d1db15e565c81e675

Observation d0ebe70e-8901-4f02-8446-3fcfbe42e772 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Nash Learning from Human Feedback

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:55.874192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T02:31:11.200818Z digest=sha256:81ef1e2db1bdf1a73c60637f103c6e76b73713c39f8b77a29d3a17b11f654490

Observation 19cc1741-87fe-4ae5-8d02-0f82ff8b90a8 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Nash Learning from Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:fdbcb094c8114b2660d8ac4d0c3bc0c5e7d87195fd8d9e307b1b0ca9617008f4

Observation 8106f321-7694-4ef0-89b1-a7aacf6a5055 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Nash Learning from Human Feedback

Reference 186

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:b7dead9405521f55cb5459cf02e89b37367a232fbdf052dbd5f7c5d9e47e9236

Observation 8fed52fb-1125-4a62-ab4c-a71d9d2971f5 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Nash Learning from Human Feedback

Reference 187

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:54.271474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:54.271474Z digest=sha256:781148f2e3af4f5527b2ff5c473593bb0ba3efedf74c36e3d7a774b9320f8c45

Observation 59004412-9117-4cb3-bc73-ab0efe768ff1 · inbound

Probably Correct Optimal Stable Matching under Two-Sided Uncertainty cites this paper.

Probably Correct Optimal Stable Matching under Two-Sided Uncertainty Nash Learning from Human Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T12:55:57.279598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T12:55:57.279598Z digest=sha256:f4254f5c64d6982b64634da06f984230e61cf61e5fa9019c47c7925d2944a5a1