Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2401.04056.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T19:40:42.044924Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:40:06.225176Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e858087c-4f0c-4fdb-a226-b81725be73f4 · inbound
KTO: Model Alignment as Prospect Theoretic Optimization A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1daba942-17dd-4a26-bfe4-8731dcd449fd · inbound
Design Considerations in Offline Preference-based RL A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba2d89d1-1192-451c-b3d4-1c5f0e0aab1b · inbound
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 528e43c5-8c35-4cc0-81c2-71cccdb87003 · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd7f7b4b-28b6-43e8-b9aa-3ae3317ff2c8 · inbound
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d722aab7-75a3-4e2c-b6c7-a7e5fb01397f · inbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a5bab6-50fb-4ee9-8ab7-5debbd946e2c · inbound
Game Theory Meets LLM and Agentic AI: Reimagining Cybersecurity for the Age of Intelligent Threats A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc2264f5-4b18-4e15-8750-cca5d7658c02 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee96b49f-3cc8-47e0-9d31-f77649e93a17 · inbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b11854d4-cec3-48d3-a19b-07aaa8f37f4c · inbound
Why Does Agentic Safety Fail to Generalize Across Tasks? A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98ceb732-b196-4dd3-a71f-c97c274c37f3 · inbound
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43f040d0-a16a-423f-9405-13456cb68b72 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 055acd68-30f0-470a-b6b4-38d503007950 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0789e877-01c2-4fa2-9f50-2a28bbbd37dc · inbound
MAPL: Multi-Objective Preference Learning for Robot Locomotion A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d35339d3-9b4b-4251-a48b-9d7aa3a6ede3 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 187
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5a3a907-f792-423e-8385-c14807685fc2 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 188
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28d11439-d28d-4cf4-9178-6e6ffaba6eb9 · inbound
Visual Token Compression Enhances Robustness of MLLMs A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.