Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:50.639923Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.23927.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:50.639923Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e4249897-9122-4b45-8413-d0da2bfbad51 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Analysis of thompson sampling for the multi-armed bandit problem
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20d9d1ba-99c3-4248-8eb1-27b873063592 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Thompson sampling for contextual bandits with linear payoffs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 123a3047-6d7c-4587-b00f-bc2f1b16ed18 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Near-optimal regret bounds for thompson sampling.J
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ffa5cb5e-80ae-48f0-a949-f154c1aeaa4c · outbound
Thompson Sampling in Online RLHF with General Function Approximation Preference-based online learning with dueling bandits: a survey.J
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 068d17e4-c7b3-451c-b836-a752a0dd49bd · outbound
Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1c0b73b2-aa5c-49b1-990e-9906ac6bb7ed · outbound
Thompson Sampling in Online RLHF with General Function Approximation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf60b172-58a5-4b1c-a562-17d5db2b54ab · outbound
Thompson Sampling in Online RLHF with General Function Approximation Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df94777c-9a6c-42ee-a834-a528c133d4c5 · outbound
Thompson Sampling in Online RLHF with General Function Approximation On the Weaknesses of Reinforcement Learning for Neural Machine Translation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec97d09e-1422-4c8e-beb9-24e7242a2e90 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Christiano, Jan Leike, Tom B
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c3b6e898-2535-472e-bb8e-1f0b4efb17ce · outbound
Thompson Sampling in Online RLHF with General Function Approximation RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b05ccd0-1758-4e94-898e-5e4cf8905e2f · outbound
Thompson Sampling in Online RLHF with General Function Approximation Schapire, Aleksandrs Slivkins, and Masrour Zoghi
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19324a54-6076-42a7-9974-d420050cd73f · outbound
Thompson Sampling in Online RLHF with General Function Approximation Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b0a6cb4-d055-45c1-963e-d178d68d1d07 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Foster and Alexander Rakhlin
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e17df10-3712-4376-ba25-69acb9567738 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Scaling laws for reward model overoptimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f93d914f-0c4c-473a-82e8-2c013645e3ae · outbound
Thompson Sampling in Online RLHF with General Function Approximation A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e83e8437-ccdd-42d5-b84d-f82d4a42da7e · outbound
Thompson Sampling in Online RLHF with General Function Approximation Reinforced Self-Training (ReST) for Language Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577cacf1-c0d0-4a3d-81be-5207f551a71d · outbound
Thompson Sampling in Online RLHF with General Function Approximation Randomized Exploration for Reinforcement Learning with General Value Function Approximation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f297a784-efaf-4ffe-89e3-7dcd0598597c · outbound
Thompson Sampling in Online RLHF with General Function Approximation Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42b36f4-36c4-48c0-91b7-2a61f39d6dd5 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Learning trajectory preferences for manipulators via iterative improvement
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8fcfabfa-f0f1-4928-b963-d64cda7e98fd · outbound
Thompson Sampling in Online RLHF with General Function Approximation Schapire
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d7c2828-6114-47ee-9b1c-75a721be65d2 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Bellman eluder dimension: new rich classes of rl problems, and sample-efficient algorithms
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c2c094a-4035-4108-932e-73172d5f9e74 · outbound
Thompson Sampling in Online RLHF with General Function Approximation An Introduction to Variational Autoencoders
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42031251-04a4-4d60-98f6-1d92937d10ec · outbound
Thompson Sampling in Online RLHF with General Function Approximation Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0776b47-d595-408d-8199-0e20cbb831ad · outbound
Thompson Sampling in Online RLHF with General Function Approximation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e9a7c7f-6d8b-48bb-a13c-f3dd10c93cb9 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Statistical Rejection Sampling Improves Preference Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d07115f-7ad9-4934-859e-b3cfd268c94f · outbound
Thompson Sampling in Online RLHF with General Function Approximation Roberts, Matthew E
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99e536a6-804c-4e8e-9d7d-6a8133fa2060 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Dueling Posterior Sampling for Preference-Based Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2dfc404-e7b3-4272-a6ce-8e481e070dfe · outbound
Thompson Sampling in Online RLHF with General Function Approximation GPT-4 Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d119f7ce-7e96-44a5-8ae4-32de28a7a17b · outbound
Thompson Sampling in Online RLHF with General Function Approximation Randomized prior functions for deep reinforcement learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 49fef49c-a677-43d7-aabc-9d0303095e41 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Approximate thompson sampling via epistemic neural networks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4149ebe-e20f-44e2-98f4-d8e5ea133972 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f9f557f-4443-4c58-8032-dc79c85f99d1 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Training language models to follow instructions with human feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 861b40f8-ed0f-40de-9e17-2b78ec60c871 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 211392fb-78f6-416f-88c2-2596ae83e015 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Worst-Case Regret Bounds for Exploration via Randomized Value Functions
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 55c69687-57ec-4809-8c77-83e041a8c371 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Learning to optimize via posterior sampling.Math
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc4131b1-1e3a-4dc0-9db7-ff14eb6c686a · outbound
Thompson Sampling in Online RLHF with General Function Approximation Optimal algorithms for stochastic contextual preference bandits
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d14f95a0-6435-44a0-ad5c-650065971df0 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Efficient and optimal algorithms for contextual dueling bandits under realizability
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 96c7ef00-ab1d-4d6b-afc8-59b1df7108b3 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Dueling rl: Reinforcement learning with trajectory preferences
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 49191887-b0d3-43ff-95e7-cb122383c61b · outbound
Thompson Sampling in Online RLHF with General Function Approximation Proximal Policy Optimization Algorithms
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 279c74ef-69af-4a5a-80c5-1c6e4b99451a · outbound
Thompson Sampling in Online RLHF with General Function Approximation Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fa5d1bcb-9c9a-48b3-8d23-0f6bb19a2314 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Thompson
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0e6e6ee-2a2d-4ea5-9d0c-675fddd8b452 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ade00be-bf41-4888-ae26-23ff03ca7444 · outbound
Thompson Sampling in Online RLHF with General Function Approximation van de Geer
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 76980b50-eb90-47f5-9a42-dc7bc440db34 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f75d0cc-fa25-45bd-bef3-aa7b602e3ef7 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Thompson sampling for combinatorial semi-bandits
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66e2e15c-29c5-4fef-b039-035424f36808 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Is rlhf more difficult than standard rl? a theoretical perspective
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32e371bb-2530-41ee-91aa-d4f60570624f · outbound
Thompson Sampling in Online RLHF with General Function Approximation A survey of preference-based reinforcement learning methods.Journal of Machine Learning Research, 18(136):1–46, 2017
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aaaeabf1-1a17-4283-b3f8-afbfa8d4821b · outbound
Thompson Sampling in Online RLHF with General Function Approximation Making RL with Preference-based Feedback Efficient via Randomization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bff1bf3-7f85-42d6-9c5b-5fa3c5a8987b · outbound
Thompson Sampling in Online RLHF with General Function Approximation Borda regret minimization for generalized linear dueling bandits
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d633d1f9-2b01-4aa8-929c-e316d7ca4f71 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50e372fb-8206-41ca-be88-96913864589f · outbound
Thompson Sampling in Online RLHF with General Function Approximation Near-optimal randomized exploration for tabular markov decision processes
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 358a7f11-bc21-4f36-ae13-e03d2f64e846 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Yang, Aarti Singh, and Artur Dubrawski
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2224a913-2e3b-4ce0-89b8-bb0c530ca86a · outbound
Thompson Sampling in Online RLHF with General Function Approximation RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efdf6433-3dac-4f16-8619-95bc291194bb · outbound
Thompson Sampling in Online RLHF with General Function Approximation The k-armed dueling bandits problem.Journal of Computer and System Sciences, 78(5):1538–1556, 2012
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 74185e95-b077-4d1a-8ed4-77e557bb356e · outbound
Thompson Sampling in Online RLHF with General Function Approximation Frequentist Regret Bounds for Randomized Least-Squares Value Iteration
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be121500-6594-443b-9b65-15b395ac386f · outbound
Thompson Sampling in Online RLHF with General Function Approximation Provable Offline Preference-Based Reinforcement Learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52573980-9b4b-4c97-91d0-417248f344a7 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Lee, and Wen Sun
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b8a432e-e368-4c56-b05f-05153a48c8e2 · outbound
Thompson Sampling in Online RLHF with General Function Approximation SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c377c5-e3ab-436d-8210-6ccd562be404 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e137d6-8164-40ce-abb3-10a5008f955d · outbound
Thompson Sampling in Online RLHF with General Function Approximation Efficient active learning with abstention
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b14aa731-9812-46c9-a59b-4c5cdea52644 · outbound
Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.