Pith. sign in

Paper Citation Record · LEDGER

Thompson Sampling in Online RLHF with General Function Approximation

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.23927.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23927 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:50.639923Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4249897-9122-4b45-8413-d0da2bfbad51 · outbound

This paper cites Analysis of thompson sampling for the multi-armed bandit problem.

Thompson Sampling in Online RLHF with General Function Approximation Analysis of thompson sampling for the multi-armed bandit problem

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.353799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.472924Z digest=sha256:caf475b472dce0bde633a5a36a12b9e900bca1fbcf0afdc89efea2ed1e5d72f0

Observation 20d9d1ba-99c3-4248-8eb1-27b873063592 · outbound

This paper cites Thompson sampling for contextual bandits with linear payoffs.

Thompson Sampling in Online RLHF with General Function Approximation Thompson sampling for contextual bandits with linear payoffs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.169067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.503626Z digest=sha256:bd74fc17193a4d45de248171c5ba6e50f77c4f1e04940c894332d161b06fede3

Observation 123a3047-6d7c-4587-b00f-bc2f1b16ed18 · outbound

This paper cites Near-optimal regret bounds for thompson sampling.J.

Thompson Sampling in Online RLHF with General Function Approximation Near-optimal regret bounds for thompson sampling.J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.050843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.527002Z digest=sha256:b8cb51d51ae4cf3eaff141d9d5ca69d74af30037b57865bffe52a5fa94d79328

Observation ffa5cb5e-80ae-48f0-a949-f154c1aeaa4c · outbound

This paper cites Preference-based online learning with dueling bandits: a survey.J.

Thompson Sampling in Online RLHF with General Function Approximation Preference-based online learning with dueling bandits: a survey.J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.921778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.557153Z digest=sha256:884b9977427a7123e602eaf9f9cf3b8fa8713265769e9fafb4d38e109565a5ff

Observation 068d17e4-c7b3-451c-b836-a752a0dd49bd · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:56.839597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.588635Z digest=sha256:ec323ac058adea123358f80259e6ac5a18078bb40195a7c972f963f822987d49

Observation 1c0b73b2-aa5c-49b1-990e-9906ac6bb7ed · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Thompson Sampling in Online RLHF with General Function Approximation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.614974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.614974Z digest=sha256:a0ace425b2280597e9c21e442ee18d8bb874fdeb3d4469dff0cec744a6a365aa

Observation bf60b172-58a5-4b1c-a562-17d5db2b54ab · outbound

This paper cites Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation.

Thompson Sampling in Online RLHF with General Function Approximation Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.660403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.660403Z digest=sha256:fdc720bbf5527e07ada586db85ce82d6340b00d8c5cb71e100a11e41e054e676

Observation df94777c-9a6c-42ee-a834-a528c133d4c5 · outbound

This paper cites On the Weaknesses of Reinforcement Learning for Neural Machine Translation.

Thompson Sampling in Online RLHF with General Function Approximation On the Weaknesses of Reinforcement Learning for Neural Machine Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.697268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.697268Z digest=sha256:bf07e71b2af8b47f13882d870093c83e40569f26956f99b517455349aa7a13a1

Observation ec97d09e-1422-4c8e-beb9-24e7242a2e90 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Thompson Sampling in Online RLHF with General Function Approximation Christiano, Jan Leike, Tom B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.773058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.725692Z digest=sha256:ee634542b05ba2f1dbf6e3b489fcae082a43492a9f757d98ed81997f126fab9b

Observation c3b6e898-2535-472e-bb8e-1f0b4efb17ce · outbound

This paper cites RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023.

Thompson Sampling in Online RLHF with General Function Approximation RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.691296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.755339Z digest=sha256:c578cc89cdf103a6e14e1e16cec8f107b7308a2645e4939c9cc3ef3bc765586c

Observation 9b05ccd0-1758-4e94-898e-5e4cf8905e2f · outbound

This paper cites Schapire, Aleksandrs Slivkins, and Masrour Zoghi.

Thompson Sampling in Online RLHF with General Function Approximation Schapire, Aleksandrs Slivkins, and Masrour Zoghi

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.598480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.799854Z digest=sha256:dc0bf3c63fbbb86714de70b133c3a511866ad9875424993ee2dcf7c414801ae0

Observation 19324a54-6076-42a7-9974-d420050cd73f · outbound

This paper cites Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO.

Thompson Sampling in Online RLHF with General Function Approximation Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.833667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.833667Z digest=sha256:16b59f7517e0957dd8d678fcb8b36ab6dfc4371f134d05c0c0816cc6f536a0fa

Observation 3b0a6cb4-d055-45c1-963e-d178d68d1d07 · outbound

This paper cites Foster and Alexander Rakhlin.

Thompson Sampling in Online RLHF with General Function Approximation Foster and Alexander Rakhlin

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.515634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.872516Z digest=sha256:d177b4b6b984412c160cb636c1ef311dc04757064296b4dd47c288640a6bacf5

Observation 7e17df10-3712-4376-ba25-69acb9567738 · outbound

This paper cites Scaling laws for reward model overoptimization.

Thompson Sampling in Online RLHF with General Function Approximation Scaling laws for reward model overoptimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.445967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.899602Z digest=sha256:9c1ff131351bed25c047735e080a3d28743fa92c111412ff780216865a20d31e

Observation f93d914f-0c4c-473a-82e8-2c013645e3ae · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

Thompson Sampling in Online RLHF with General Function Approximation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.923648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.923648Z digest=sha256:db4617b484de218cc4319a9372284a62391f5f7fd55ff2180ca38bc81a52fb3f

Observation e83e8437-ccdd-42d5-b84d-f82d4a42da7e · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Thompson Sampling in Online RLHF with General Function Approximation Reinforced Self-Training (ReST) for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.949387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.949387Z digest=sha256:349ab6349c926fef5c01d2abd52f5d3e20486e2961a10098fc4e7b93a5b01505

Observation 577cacf1-c0d0-4a3d-81be-5207f551a71d · outbound

This paper cites Randomized Exploration for Reinforcement Learning with General Value Function Approximation.

Thompson Sampling in Online RLHF with General Function Approximation Randomized Exploration for Reinforcement Learning with General Value Function Approximation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:51.698037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:47.982915Z digest=sha256:d98c1f209f3bc210f19e9f1a814adf2565b90bd75c6045ec2e3a43e45331f09f

Observation f297a784-efaf-4ffe-89e3-7dcd0598597c · outbound

This paper cites Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo.

Thompson Sampling in Online RLHF with General Function Approximation Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.010070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.010070Z digest=sha256:d1d91635decba9880849edeffddc38596f0440ee0fe89655646028369999c6a1

Observation c42b36f4-36c4-48c0-91b7-2a61f39d6dd5 · outbound

This paper cites Learning trajectory preferences for manipulators via iterative improvement.

Thompson Sampling in Online RLHF with General Function Approximation Learning trajectory preferences for manipulators via iterative improvement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.372101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.044855Z digest=sha256:f9fa16fdc4635a9e514860dea0e8c23321c55d4f70055b14cb572cef797570b2

Observation 8fcfabfa-f0f1-4928-b963-d64cda7e98fd · outbound

This paper cites Schapire.

Thompson Sampling in Online RLHF with General Function Approximation Schapire

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.305869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.086054Z digest=sha256:d114fd181aa8de138e366c161709a2991161468d3e31776a305def6d5eedd2be

Observation 8d7c2828-6114-47ee-9b1c-75a721be65d2 · outbound

This paper cites Bellman eluder dimension: new rich classes of rl problems, and sample-efficient algorithms.

Thompson Sampling in Online RLHF with General Function Approximation Bellman eluder dimension: new rich classes of rl problems, and sample-efficient algorithms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.160284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.220473Z digest=sha256:92122fac0d547c7cb55a547fa88dcb297eecb61018308753bcc1984056e744f4

Observation 0c2c094a-4035-4108-932e-73172d5f9e74 · outbound

This paper cites An Introduction to Variational Autoencoders.

Thompson Sampling in Online RLHF with General Function Approximation An Introduction to Variational Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.246384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.246384Z digest=sha256:90dff8ce0f3c1948c37c11276500e626768db8da5228551eae119173c15e28dc

Observation 42031251-04a4-4d60-98f6-1d92937d10ec · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Thompson Sampling in Online RLHF with General Function Approximation Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.277156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.277156Z digest=sha256:c71424f369c57d6d4fd40e43bc8c295d388da14231b508623d4ed638855b5463

Observation f0776b47-d595-408d-8199-0e20cbb831ad · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Thompson Sampling in Online RLHF with General Function Approximation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.349648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.349648Z digest=sha256:720b809aa4bcc7e6843f5994efb1a44169bc4d490a62c8e46c6bb01cff9e6738

Observation 5e9a7c7f-6d8b-48bb-a13c-f3dd10c93cb9 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Thompson Sampling in Online RLHF with General Function Approximation Statistical Rejection Sampling Improves Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.431444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.431444Z digest=sha256:f87db01697fba5a46da1c10219220f66d6f09f9475bf368763a07f91e72cd1cb

Observation 2d07115f-7ad9-4934-859e-b3cfd268c94f · outbound

This paper cites Roberts, Matthew E.

Thompson Sampling in Online RLHF with General Function Approximation Roberts, Matthew E

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.983442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.478015Z digest=sha256:b8b264c9e03a0a5193c5dc673d3e1d132400f7f216e3a611cba5767e6a610746

Observation 99e536a6-804c-4e8e-9d7d-6a8133fa2060 · outbound

This paper cites Dueling Posterior Sampling for Preference-Based Reinforcement Learning.

Thompson Sampling in Online RLHF with General Function Approximation Dueling Posterior Sampling for Preference-Based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.515394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.515394Z digest=sha256:5658831ef0beedc34b57ecf0244cad2ca8c19f22922d606610e6673dc44f9ccc

Observation d2dfc404-e7b3-4272-a6ce-8e481e070dfe · outbound

This paper cites GPT-4 Technical Report.

Thompson Sampling in Online RLHF with General Function Approximation GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.557964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.557964Z digest=sha256:25754fe543b44e7e6b76fd0216f68b90baf190102ac1a8628011350ad987e5e6

Observation d119f7ce-7e96-44a5-8ae4-32de28a7a17b · outbound

This paper cites Randomized prior functions for deep reinforcement learning.

Thompson Sampling in Online RLHF with General Function Approximation Randomized prior functions for deep reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.724220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.595572Z digest=sha256:692e6bedfb51ff2df26da9cd00d292b5a4ae38904d1944b6cbc0bd9bb26fb69d

Observation 49fef49c-a677-43d7-aabc-9d0303095e41 · outbound

This paper cites Approximate thompson sampling via epistemic neural networks.

Thompson Sampling in Online RLHF with General Function Approximation Approximate thompson sampling via epistemic neural networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.458395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.687662Z digest=sha256:a902ab8b48ca8826484281f040e8084663f0378f9048bbeab20a0f1a75ccdc38

Observation a4149ebe-e20f-44e2-98f4-d8e5ea133972 · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:55.265065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.737226Z digest=sha256:ae3b8d1569c82a82309662edc52f73b561488d90a7da3eea6d142eae3accdb87

Observation 6f9f557f-4443-4c58-8032-dc79c85f99d1 · outbound

This paper cites Training language models to follow instructions with human feedback.

Thompson Sampling in Online RLHF with General Function Approximation Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.769424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.769424Z digest=sha256:d3a704afa981512517c9dafb97bc97f4ca968609284d721167f4a39b3eed32be

Observation 861b40f8-ed0f-40de-9e17-2b78ec60c871 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Thompson Sampling in Online RLHF with General Function Approximation Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.796716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.796716Z digest=sha256:bc229851441434975ec37d418e2349a63696e977011ed23eff1a1e033085ab54

Observation 211392fb-78f6-416f-88c2-2596ae83e015 · outbound

This paper cites Worst-Case Regret Bounds for Exploration via Randomized Value Functions.

Thompson Sampling in Online RLHF with General Function Approximation Worst-Case Regret Bounds for Exploration via Randomized Value Functions

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:51.321069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.814525Z digest=sha256:5b7aed63ab421fae450bd27e30597503ee8598a6fb1be6e712895ea67d05a537

Observation 55c69687-57ec-4809-8c77-83e041a8c371 · outbound

This paper cites Learning to optimize via posterior sampling.Math.

Thompson Sampling in Online RLHF with General Function Approximation Learning to optimize via posterior sampling.Math

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.197918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.838269Z digest=sha256:96a23a8ca3e0ec1247985b7ad0309286ce3e669b8c133e252d6c8f8ce8ae276e

Observation bc4131b1-1e3a-4dc0-9db7-ff14eb6c686a · outbound

This paper cites Optimal algorithms for stochastic contextual preference bandits.

Thompson Sampling in Online RLHF with General Function Approximation Optimal algorithms for stochastic contextual preference bandits

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.096880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.861004Z digest=sha256:88afe2bc0681f892a2b7482c39745c6b08b9b202eee8224f31ad94a0f20ddfaa

Observation d14f95a0-6435-44a0-ad5c-650065971df0 · outbound

This paper cites Efficient and optimal algorithms for contextual dueling bandits under realizability.

Thompson Sampling in Online RLHF with General Function Approximation Efficient and optimal algorithms for contextual dueling bandits under realizability

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.840829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.890800Z digest=sha256:355e0b13c84436447ede5d8b1a7e16bf981a207cdd38582fd2dd7f386dccf904

Observation 96c7ef00-ab1d-4d6b-afc8-59b1df7108b3 · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Thompson Sampling in Online RLHF with General Function Approximation Dueling rl: Reinforcement learning with trajectory preferences

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.465692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.921797Z digest=sha256:fb244d185918bd8e5dd866d613258492256d4f943c04d02c27cc0b2593a00dd4

Observation 49191887-b0d3-43ff-95e7-cb122383c61b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Thompson Sampling in Online RLHF with General Function Approximation Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.940005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.940005Z digest=sha256:2788c7a9292ae9e010ae9fefe6bbc61987ead711c7bdf266d459bf206a23c399

Observation 279c74ef-69af-4a5a-80c5-1c6e4b99451a · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano.

Thompson Sampling in Online RLHF with General Function Approximation Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.117813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.969661Z digest=sha256:45bc00a07d587353a0a3816fdff68aaa8cf78d0e0c1002b0bea3690a72567f32

Observation fa5d1bcb-9c9a-48b3-8d23-0f6bb19a2314 · outbound

This paper cites Thompson.

Thompson Sampling in Online RLHF with General Function Approximation Thompson

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.933023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:48.992418Z digest=sha256:8eb6f02e6da351e514bf34896c4da7a7d17686b018dba89bd972b28ef87e1609

Observation d0e6e6ee-2a2d-4ea5-9d0c-675fddd8b452 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Thompson Sampling in Online RLHF with General Function Approximation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.018024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.018024Z digest=sha256:0282e79290dd3cd7958dd865b2f739a8a267d3dee957900d35cb492342e61a25

Observation 5ade00be-bf41-4888-ae26-23ff03ca7444 · outbound

This paper cites van de Geer.

Thompson Sampling in Online RLHF with General Function Approximation van de Geer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.773838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:49.052190Z digest=sha256:d432d317978a2d1a2f61f287a7d4421103dedf8ab6cf1c6f67ac2260735416d5

Observation 76980b50-eb90-47f5-9a42-dc7bc440db34 · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Thompson Sampling in Online RLHF with General Function Approximation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.083529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.083529Z digest=sha256:fa0c15ea9b17a034ecbcbd85bb87353bc34b52a25f602462e46a8f448cbcf8e2

Observation 7f75d0cc-fa25-45bd-bef3-aa7b602e3ef7 · outbound

This paper cites Thompson sampling for combinatorial semi-bandits.

Thompson Sampling in Online RLHF with General Function Approximation Thompson sampling for combinatorial semi-bandits

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.720079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:49.108031Z digest=sha256:b1b16dd2375bcd3dc971041cb9089b9b1dda631648511be593cdb630b6f1cf9d

Observation 66e2e15c-29c5-4fef-b039-035424f36808 · outbound

This paper cites Is rlhf more difficult than standard rl? a theoretical perspective.

Thompson Sampling in Online RLHF with General Function Approximation Is rlhf more difficult than standard rl? a theoretical perspective

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.586179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:49.117791Z digest=sha256:bbe5d28664823ded034dfbb733abe62a2839e7131e7ac7120ecdec9a6e2a99e1

Observation 32e371bb-2530-41ee-91aa-d4f60570624f · outbound

This paper cites A survey of preference-based reinforcement learning methods.Journal of Machine Learning Research, 18(136):1–46, 2017.

Thompson Sampling in Online RLHF with General Function Approximation A survey of preference-based reinforcement learning methods.Journal of Machine Learning Research, 18(136):1–46, 2017

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.466104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:49.197952Z digest=sha256:d9254c038d64fa898a2d3f77e20a764c783328f420af3dd569c79c682db39807

Observation aaaeabf1-1a17-4283-b3f8-afbfa8d4821b · outbound

This paper cites Making RL with Preference-based Feedback Efficient via Randomization.

Thompson Sampling in Online RLHF with General Function Approximation Making RL with Preference-based Feedback Efficient via Randomization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.305402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.305402Z digest=sha256:f1b614840ac046395d105771440b08cdecdea53ff730bb35c6c15bf3ddb5c42f

Observation 3bff1bf3-7f85-42d6-9c5b-5fa3c5a8987b · outbound

This paper cites Borda regret minimization for generalized linear dueling bandits.

Thompson Sampling in Online RLHF with General Function Approximation Borda regret minimization for generalized linear dueling bandits

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.309730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:49.376621Z digest=sha256:599e17eb3da62621ffba9cf91a9a4fb7807a710e92bfef0a3010711546b8394c

Observation d633d1f9-2b01-4aa8-929c-e316d7ca4f71 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Thompson Sampling in Online RLHF with General Function Approximation Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.435213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.435213Z digest=sha256:e1053ef039de9bde3922c6a72888c44f96ec29a60f10e421fa0969ed3c5f4c3d

Observation 50e372fb-8206-41ca-be88-96913864589f · outbound

This paper cites Near-optimal randomized exploration for tabular markov decision processes.

Thompson Sampling in Online RLHF with General Function Approximation Near-optimal randomized exploration for tabular markov decision processes

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.149868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:49.519282Z digest=sha256:ac14e93f5a9d2838c7a1a34c7931e6510489cce443835b54cdca4539cdd80c22

Observation 358a7f11-bc21-4f36-ae13-e03d2f64e846 · outbound

This paper cites Yang, Aarti Singh, and Artur Dubrawski.

Thompson Sampling in Online RLHF with General Function Approximation Yang, Aarti Singh, and Artur Dubrawski

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.943796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:49.630519Z digest=sha256:e68d0dbde1d217a8c2b62d2b2541f9c4f29e25dce8855d96a4f112161070900e

Observation 2224a913-2e3b-4ce0-89b8-bb0c530ca86a · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Thompson Sampling in Online RLHF with General Function Approximation RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.730231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.730231Z digest=sha256:ce595d3af54af1549be69d962266b6ec6297ce3653d1609f3ddf748c0a65f7ba

Observation efdf6433-3dac-4f16-8619-95bc291194bb · outbound

This paper cites The k-armed dueling bandits problem.Journal of Computer and System Sciences, 78(5):1538–1556, 2012.

Thompson Sampling in Online RLHF with General Function Approximation The k-armed dueling bandits problem.Journal of Computer and System Sciences, 78(5):1538–1556, 2012

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.725910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:49.832227Z digest=sha256:5fe3778076a1fbae35d1afa41de1243ddef08147e091615e866b3002cf4dff2c

Observation 74185e95-b077-4d1a-8ed4-77e557bb356e · outbound

This paper cites Frequentist Regret Bounds for Randomized Least-Squares Value Iteration.

Thompson Sampling in Online RLHF with General Function Approximation Frequentist Regret Bounds for Randomized Least-Squares Value Iteration

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:50.963293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:49.926426Z digest=sha256:38effcc055c141d76e8f389f9811ba30c13ec53ac9806b6b647410d8426a55af

Observation be121500-6594-443b-9b65-15b395ac386f · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Thompson Sampling in Online RLHF with General Function Approximation Provable Offline Preference-Based Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.030374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.030374Z digest=sha256:a0d3a38833e48ddf130e6d3253622c990e6b7201f48c39c80f22d1152aac8652

Observation 52573980-9b4b-4c97-91d0-417248f344a7 · outbound

This paper cites Lee, and Wen Sun.

Thompson Sampling in Online RLHF with General Function Approximation Lee, and Wen Sun

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.502747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:50.138027Z digest=sha256:20c4f1b7e7c6e611f22d512f5ca9723ea186d06dd17554f2012b0076668964d3

Observation 0b8a432e-e368-4c56-b05f-05153a48c8e2 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Thompson Sampling in Online RLHF with General Function Approximation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.251391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.251391Z digest=sha256:d28230d1711dd90dd4b32a5463074014e35d7c3161915084f99e0215a81e13af

Observation 03c377c5-e3ab-436d-8210-6ccd562be404 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

Thompson Sampling in Online RLHF with General Function Approximation Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.418201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.418201Z digest=sha256:88d960222ae82cbd828997e18c7b3c754df06c63d0ed524ee34fa9ad76647611

Observation 77e137d6-8164-40ce-abb3-10a5008f955d · outbound

This paper cites Efficient active learning with abstention.

Thompson Sampling in Online RLHF with General Function Approximation Efficient active learning with abstention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.301764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:50.546951Z digest=sha256:4926d72e72dad9aa40e714aee34ecae8ae3248f5147af5690577f15efc05a4cd

Observation b14aa731-9812-46c9-a59b-4c5cdea52644 · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:52.029084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:50.639923Z digest=sha256:becfb0655880d03d93f2448e28aa35517a18e5acf87aed8101a294bff288ecfd

Pith citing papers

No inbound Pith citation observations are available.