Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:21:33.386906Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2412.16834.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:21:33.386906Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T16:38:06.631594Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T03:57:29.547083Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5581eb0b-618e-43bc-b1fc-f2c35210ed43 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Open and ef- ficient foundation language models,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 080bf1ba-c99c-4b07-8e54-c0fd5fbf53e7 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Openassistant conversations-democratizing large langu age model align- ment,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5d2cb5f1-e569-4123-ab5e-b8eb35a4e966 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Mechanism design for llm fine-tuning with multiple reward models,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 659ffb1b-6c6c-4db9-88ae-034280e88d3e · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Deep reinforcement learning from human preferences,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 13afe7f0-dfec-4ad9-804a-edc1f1c76c0e · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Training language models to follow instructions with human feedback,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2201e4b3-4a35-4481-b6e3-61ef869d996b · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d64fcd5f-ca12-4676-be30-df721c4eb9aa · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9fa04d5-8af8-470f-86fe-f46f49dce922 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Iterative preference learning from human fee dback: Bridging theory and practice for rlhf under kl-constraint,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 84f3f08e-0452-40bb-8c30-f96e7e88307a · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Truthful Aggregation of LLMs with an Application to Online Advertising
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aa31b4f-9400-49ad-8cf5-4aa5b510fff6 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9a329d-d042-488a-a3e7-8775e9c19de6 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Online prediction w ith selfish experts,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1ceec05b-4d48-4569-ba79-91279ec29a70 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Direct preference optimization: Y our language mo del is secretly a reward model,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bf8c9c6e-1d9c-4cd5-856b-2f461b19f5b9 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22f9dbe2-1194-4ba6-a6ce-6bcf58055ccf · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 770af17c-b057-4d2b-8d01-158888e17854 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b23332-c366-4c73-87c2-e633713626e4 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning OPTune: Efficient Online Preference Tuning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57010b3c-4180-449f-97b4-8e19fc0d0e30 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Rl hf from heterogeneous feedback via personalization and preferenc e aggregation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0065d0fc-4288-4c1f-bc2d-dcd572a60ec3 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Auctions with LLM Summaries
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a1be47e-1fa3-4e51-b83a-ef1b41116068 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Co llaborative algorithms for online personalized mean estimation,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6417fda4-3096-4ec6-acf2-616faec6f01b · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Mechanism design for col- laborative normal mean estimation,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eb98fa7f-171c-44ed-b445-fa521143fb8f · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Strategyproof mechanisms for group- fair obnoxious facility location problems,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f59b7e01-752d-4ea8-88fa-29e8daedd497 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Positive intra-group exter nalities in facility location,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 102e7f13-4d9e-4e7b-b86e-24dd6689a6e8 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2834b20f-dec2-43fe-9d11-f6a871fb107c · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Human-i n- the-loop: Provably efficient preference-based reinforcem ent learning with general function approximation,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cac1b905-4782-4861-851b-c2214bd75009 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a179afb-9f1c-45ea-8f62-8831a933f466 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning E fficient com- petitions and online learning with strategic forecasters,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e64388f0-e0a5-4448-a6ad-1238f6fb6213 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning No-regret and incentive-compatible online learning,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f2b7db93-697f-4409-934e-07d819ea4fd0 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning trlx: A framework for large s cale rein- forcement learning from human feedback,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d03a433a-d317-4c59-b926-19421549c927 · outbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Tl; dr: M ining reddit to learn automatic summarization,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e4c26c1d-ba06-405a-9602-c300644a5c25 · inbound
The Battling Influencers Game: Nash Equilibria Structure of a Potential Game and Implications to Value Alignment Online Learning from Strategic Human Feedback in LLM Fine-Tuning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e86a58fa-c6b6-4bea-968b-51222439eda9 · inbound
How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Online Learning from Strategic Human Feedback in LLM Fine-Tuning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a8d5d594-0c3c-4105-980d-49db6e563d74 · inbound
Incentivizing High-Quality Human Annotations with Golden Questions Online Learning from Strategic Human Feedback in LLM Fine-Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.