Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2402.08925.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:46.346809Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 096646df-cbfe-410b-b49b-a168e2289500 · inbound
Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce76eea6-79d0-441b-acdb-e17c7eb94076 · inbound
AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 124
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b22f4d-898f-4f46-9fba-5bcad3ffbc6f · inbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5faf15b-93b8-4fd3-ae42-d75431351b0a · inbound
Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceee655c-706a-4bed-9b54-bddb236a2c17 · inbound
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 074dd878-1664-4c9d-88dc-fa80b3d653da · inbound
Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e120f7f4-43cc-4699-8c0a-44fc1c404824 · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb46bd4-3130-4a4c-8d0c-9357b37ad186 · inbound
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4587229b-30a2-4f7f-bbc2-331ef119362b · inbound
Routing Sensitivity Without Controllability: A Diagnostic Study of Fairness in MoE Language Models MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d32f6137-2840-45a6-baa6-75fd3da89908 · inbound
Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 611edaa1-0145-4bd2-9645-e1a3494810c1 · inbound
Three Models of RLHF Annotation: Extension, Evidence, and Authority MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e820d6a1-ae8a-4936-8432-c780f7c99abe · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 897c0801-edca-4526-891b-9f9aab4e11b2 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67124bfd-3103-4e17-bac8-cb7791b09bc8 · inbound
Common-agency Games for Multi-Objective Test-Time Alignment MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 186
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ffe8c78-39bb-4d15-8aff-ac38ad70fead · inbound
Spectral Souping: A Unified Framework for Online Preference Alignment MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a008025-17e7-46c0-92e4-3797e5b5484b · inbound
In-Context Reward Adaptation for Robust Preference Modeling MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2ccfd2b-6502-4400-9d0a-2d42449f3549 · inbound
Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159df242-8eb4-4c45-9569-a202f71cc470 · inbound
Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 170
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33884791-2692-4c9a-9c63-1b500d432db3 · inbound
Hidden Consensus:Preference-Validity Compression in Human Feedback MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.