Pith. sign in

Paper Citation Record · LEDGER

A Minimaximalist Approach to Reinforcement Learning from Human Feedback

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2401.04056.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.04056 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:40:42.044924Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:06.225176Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e858087c-4f0c-4fdb-a226-b81725be73f4 · inbound

KTO: Model Alignment as Prospect Theoretic Optimization cites this paper.

KTO: Model Alignment as Prospect Theoretic Optimization A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:17:53.574466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T12:17:53.478052Z digest=sha256:b46ffb5ffaace1c9bbb31dd6b510fbca13e3a64731b58eef69192a0bce5bb5d4

Observation 1daba942-17dd-4a26-bfe4-8731dcd449fd · inbound

Design Considerations in Offline Preference-based RL cites this paper.

Design Considerations in Offline Preference-based RL A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.044924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.044924Z digest=sha256:537f9fc1b8f3ae2b4a40bdeddf25b408538c64e62c51dccb5d4af0d420e0853c

Observation ba2d89d1-1192-451c-b3d4-1c5f0e0aab1b · inbound

Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games cites this paper.

Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T20:38:32.922229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:38:32.922229Z digest=sha256:0844d8a3cae6d35d56836081852bacea50861d66b1472ad37003472c58c0f21c

Observation 528e43c5-8c35-4cc0-81c2-71cccdb87003 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.173114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.173114Z digest=sha256:2cd4582858ceccc5a8398ad0639ee7ccdb8cc46924c890007543b535066301ca

Observation cd7f7b4b-28b6-43e8-b9aa-3ae3317ff2c8 · inbound

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? cites this paper.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.847272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.847272Z digest=sha256:567f4e3f88069ec819be692207907342fd5c7c926df9a5490d81658e3075fe87

Observation d722aab7-75a3-4e2c-b6c7-a7e5fb01397f · inbound

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs cites this paper.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.528670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.528670Z digest=sha256:b8bd088d8e09be2a2b0392ec1dda3b0e95e7dce5c2a86afb005803fbdf212f5c

Observation 16a5bab6-50fb-4ee9-8ab7-5debbd946e2c · inbound

Game Theory Meets LLM and Agentic AI: Reimagining Cybersecurity for the Age of Intelligent Threats cites this paper.

Game Theory Meets LLM and Agentic AI: Reimagining Cybersecurity for the Age of Intelligent Threats A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:53:19.732489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:53:19.732489Z digest=sha256:3dbcdff6ee23baef7e2ca0750435869a68745808732c40bdd3546503649da66e

Observation bc2264f5-4b18-4e15-8750-cca5d7658c02 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:40.529374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:40.529374Z digest=sha256:c526910cac5398a14296065e5e2f2e019d7073a25c45e00bfbd8980df3304964

Observation ee96b49f-3cc8-47e0-9d31-f77649e93a17 · inbound

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective cites this paper.

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:23.941662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:23.941662Z digest=sha256:b976e48cb4571b6c59c1d4b05014c86b16c83b46a924a5ee0ba33e81c8b921dc

Observation b11854d4-cec3-48d3-a19b-07aaa8f37f4c · inbound

Why Does Agentic Safety Fail to Generalize Across Tasks? cites this paper.

Why Does Agentic Safety Fail to Generalize Across Tasks? A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:06:00.273725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:55:38.554161Z digest=sha256:1eca734ff5a04a3e2de9ad7f1c34bc0910fb97bfb03389fc8c587d57318046a2

Observation 98ceb732-b196-4dd3-a71f-c97c274c37f3 · inbound

Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs cites this paper.

Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:31.191814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:56:01.953853Z digest=sha256:cfe1fd130240af2a34c81d8d30ab7e2e85e4aaeace6125b61bbb06ab87092ed0

Observation 43f040d0-a16a-423f-9405-13456cb68b72 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:17.360108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:5403756317c6bdb8723dfef1e19c8e90b39d241a829f7dbe1122779203100008

Observation 055acd68-30f0-470a-b6b4-38d503007950 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:06.540244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:b8a489b0027d3082bce43dcb54f5254401766037d2e510bf3f7830107d11561b

Observation 0789e877-01c2-4fa2-9f50-2a28bbbd37dc · inbound

MAPL: Multi-Objective Preference Learning for Robot Locomotion cites this paper.

MAPL: Multi-Objective Preference Learning for Robot Locomotion A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:40:06.226832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T21:11:00.753949Z digest=sha256:abb24700d3f1c2188256e07e59904284551ecaa13cf526ad6793726e777b6326

Observation d35339d3-9b4b-4251-a48b-9d7aa3a6ede3 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 187

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:0835c0e5209fc91ecb15e44542d3177568efc5b9a9ad55030e7a0a736b4b5a06

Observation a5a3a907-f792-423e-8385-c14807685fc2 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 188

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:54.387397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:54.387397Z digest=sha256:147b15a1a8d1ad7e21235049328b7402ec298b465695b939e1c5cb6f04a50c91

Observation 28d11439-d28d-4cf4-9178-6e6ffaba6eb9 · inbound

Visual Token Compression Enhances Robustness of MLLMs cites this paper.

Visual Token Compression Enhances Robustness of MLLMs A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T13:10:37.604919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:10:37.604919Z digest=sha256:c6f545ad44321a2bf301b1c73a00b692aad01e0f2e93e57153357d8d87ef9ac7