Pith. sign in

Paper Citation Record · LEDGER

MaxMin-RLHF: Alignment with Diverse Human Preferences

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2402.08925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.08925 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:46.346809Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 096646df-cbfe-410b-b49b-a168e2289500 · inbound

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? cites this paper.

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:46.346809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:46.346809Z digest=sha256:ef8882939aec99e79041f5ad15d975995df87b612e0ee94254f4ec33f383a0ab

Observation ce76eea6-79d0-441b-acdb-e17c7eb94076 · inbound

AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing cites this paper.

AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:55.785963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:55.785963Z digest=sha256:3e4a505ab1d1e8a66627a70c7a779f89cae4c5dbf64eb024c4de92e3aff2ebb3

Observation e4b22f4d-898f-4f46-9fba-5bcad3ffbc6f · inbound

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning cites this paper.

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:32.906989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:32.906989Z digest=sha256:96bf8ab58daeec46a90304f99f2b0b7047cfe95a49dea36ec9378d6f8a6fcce0

Observation a5faf15b-93b8-4fd3-ae42-d75431351b0a · inbound

Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching cites this paper.

Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:37.813949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:37.813949Z digest=sha256:20493f23febdd4a0254416ac1306091988f18562052edccab14028706a6037e3

Observation ceee655c-706a-4bed-9b54-bddb236a2c17 · inbound

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? cites this paper.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:05.774124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:05.774124Z digest=sha256:c786249c0dc465db58f37b710f7f9568a550d13f9e965b78f8b717adad708aaf

Observation 074dd878-1664-4c9d-88dc-fa80b3d653da · inbound

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory cites this paper.

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T01:04:30.349050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:04:30.349050Z digest=sha256:e5344592239941473decc370e696e7239fac61ee5d9d1be6bf951abcfff30f94

Observation e120f7f4-43cc-4699-8c0a-44fc1c404824 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:24.955343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:24.955343Z digest=sha256:89cd682d595bde3e04e6121b63586263f0d5a8abfbd5b9a63c082c04ae85d960

Observation 1fb46bd4-3130-4a4c-8d0c-9357b37ad186 · inbound

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences cites this paper.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.144244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.144244Z digest=sha256:e9d1685108159e183c78d3c8b4d6381d85d1f71e0aece83710385fe566bfb2d7

Observation 4587229b-30a2-4f7f-bbc2-331ef119362b · inbound

Routing Sensitivity Without Controllability: A Diagnostic Study of Fairness in MoE Language Models cites this paper.

Routing Sensitivity Without Controllability: A Diagnostic Study of Fairness in MoE Language Models MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:08:14.502735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T23:06:53.369244Z digest=sha256:24af1a00271a79eb7252bf8d4c1ff82194d94a879fcb3fd2b97b757f01b2183f

Observation d32f6137-2840-45a6-baa6-75fd3da89908 · inbound

Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration cites this paper.

Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T10:27:28.560210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:27:28.560210Z digest=sha256:69042c81fad749b21dc90ac54ce5d9b62058d1829d534076968a9b9b34b10f72

Observation 611edaa1-0145-4bd2-9645-e1a3494810c1 · inbound

Three Models of RLHF Annotation: Extension, Evidence, and Authority cites this paper.

Three Models of RLHF Annotation: Extension, Evidence, and Authority MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:12.376743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T14:28:31.451466Z digest=sha256:aaa07afddb7193f0113ddf89df484bb038d62fad8a353c2f449c2c3c889f213e

Observation e820d6a1-ae8a-4936-8432-c780f7c99abe · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 145

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.363434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:5de439593e05dd3c5e502fda78003e1c98a5bd55795f44f81897046adbcb5258

Observation 897c0801-edca-4526-891b-9f9aab4e11b2 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 145

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.687669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:8d5f2f6e43a66adda88db347ffb81f46be78b81cfdb379e130f8f85d413f9971

Observation 67124bfd-3103-4e17-bac8-cb7791b09bc8 · inbound

Common-agency Games for Multi-Objective Test-Time Alignment cites this paper.

Common-agency Games for Multi-Objective Test-Time Alignment MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 186

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:15:06.670446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:14:53.685486Z digest=sha256:451f8f6622d8a4041673987f9f89d828c93948c07af5dd334f15324670d19e93

Observation 3ffe8c78-39bb-4d15-8aff-ac38ad70fead · inbound

Spectral Souping: A Unified Framework for Online Preference Alignment cites this paper.

Spectral Souping: A Unified Framework for Online Preference Alignment MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.933384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:54:56.356555Z digest=sha256:5723d9044ef788778889994f3e32e5815b169144daae75aa4847a6e9d4bc5253

Observation 2a008025-17e7-46c0-92e4-3797e5b5484b · inbound

In-Context Reward Adaptation for Robust Preference Modeling cites this paper.

In-Context Reward Adaptation for Robust Preference Modeling MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.199143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:19:55.177440Z digest=sha256:c1f1834b40219d878b7c8b5622195f27d12705236fb3f26594db34c850a35275

Observation a2ccfd2b-6502-4400-9d0a-2d42449f3549 · inbound

Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis cites this paper.

Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T17:56:39.720418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:56:39.720418Z digest=sha256:622a1b5ca2b08bd2793415b327640de940dc242f1bd3f75b515faf5c126e2279

Observation 159df242-8eb4-4c45-9569-a202f71cc470 · inbound

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs cites this paper.

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 170

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T16:51:06.023631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:49:14.243931Z digest=sha256:8fbd4f0f633e9fe5256fb7714293dfcc5fe43c01127c8a96965575d87eb52bc2

Observation 33884791-2692-4c9a-9c63-1b500d432db3 · inbound

Hidden Consensus:Preference-Validity Compression in Human Feedback cites this paper.

Hidden Consensus:Preference-Validity Compression in Human Feedback MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:07:39.034042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T13:25:24.761133Z digest=sha256:f84a7e022ddb1df2bfbe0224bbe4e48e14ae35f4b27bbdbf3ed6eee505d28dcd