Pith. sign in

Paper Citation Record · LEDGER

Thompson Sampling in Online RLHF with General Function Approximation

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.23927.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23927 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:50.639923Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4249897-9122-4b45-8413-d0da2bfbad51 · outbound

This paper cites Analysis of thompson sampling for the multi-armed bandit problem.

Thompson Sampling in Online RLHF with General Function Approximation Analysis of thompson sampling for the multi-armed bandit problem

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.353799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.472924Z digest=sha256:dde4385085856ebb0e26f3789750d39fdee5a06ea225cdde77e99043eb70c8e7

Observation 20d9d1ba-99c3-4248-8eb1-27b873063592 · outbound

This paper cites Thompson sampling for contextual bandits with linear payoffs.

Thompson Sampling in Online RLHF with General Function Approximation Thompson sampling for contextual bandits with linear payoffs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.169067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.503626Z digest=sha256:9153dedc6a364b214a33d9cf4662b8889f452e407680600f03cda015b760832e

Observation 123a3047-6d7c-4587-b00f-bc2f1b16ed18 · outbound

This paper cites Near-optimal regret bounds for thompson sampling.J.

Thompson Sampling in Online RLHF with General Function Approximation Near-optimal regret bounds for thompson sampling.J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:57.050843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.527002Z digest=sha256:cee50bc6f2917431ba22cb0c7cb73967c0c421038641f3c641284bdc15813dd7

Observation ffa5cb5e-80ae-48f0-a949-f154c1aeaa4c · outbound

This paper cites Preference-based online learning with dueling bandits: a survey.J.

Thompson Sampling in Online RLHF with General Function Approximation Preference-based online learning with dueling bandits: a survey.J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.921778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.557153Z digest=sha256:2eb3fabafbf6fab571f1b36c0f5123117a7dab2cd925d9c37f1c7c3789ebb942

Observation 068d17e4-c7b3-451c-b836-a752a0dd49bd · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:56.839597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.588635Z digest=sha256:eac4f9ee4b976ffef30e592093f34e5a4e82eff45fbe1a112da0a54db948ba58

Observation 1c0b73b2-aa5c-49b1-990e-9906ac6bb7ed · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Thompson Sampling in Online RLHF with General Function Approximation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.614974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.614974Z digest=sha256:a0ace425b2280597e9c21e442ee18d8bb874fdeb3d4469dff0cec744a6a365aa

Observation bf60b172-58a5-4b1c-a562-17d5db2b54ab · outbound

This paper cites Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation.

Thompson Sampling in Online RLHF with General Function Approximation Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.660403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.660403Z digest=sha256:fdc720bbf5527e07ada586db85ce82d6340b00d8c5cb71e100a11e41e054e676

Observation df94777c-9a6c-42ee-a834-a528c133d4c5 · outbound

This paper cites On the Weaknesses of Reinforcement Learning for Neural Machine Translation.

Thompson Sampling in Online RLHF with General Function Approximation On the Weaknesses of Reinforcement Learning for Neural Machine Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.697268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.697268Z digest=sha256:bf07e71b2af8b47f13882d870093c83e40569f26956f99b517455349aa7a13a1

Observation ec97d09e-1422-4c8e-beb9-24e7242a2e90 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Thompson Sampling in Online RLHF with General Function Approximation Christiano, Jan Leike, Tom B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.773058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.725692Z digest=sha256:eebd14c9fe7ffaf2854394743ac7a82377553d170eb947cc6da6902cabb9d8a5

Observation c3b6e898-2535-472e-bb8e-1f0b4efb17ce · outbound

This paper cites RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023.

Thompson Sampling in Online RLHF with General Function Approximation RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.691296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.755339Z digest=sha256:02171d708138121a0ab8eb3a16fb97823134c1d91bc15a4c534ac6080d418f56

Observation 9b05ccd0-1758-4e94-898e-5e4cf8905e2f · outbound

This paper cites Schapire, Aleksandrs Slivkins, and Masrour Zoghi.

Thompson Sampling in Online RLHF with General Function Approximation Schapire, Aleksandrs Slivkins, and Masrour Zoghi

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.598480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.799854Z digest=sha256:548025ebc6eda0889ee06622bba86612a3a14090b3481e537e8be1ab35669faa

Observation 19324a54-6076-42a7-9974-d420050cd73f · outbound

This paper cites Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO.

Thompson Sampling in Online RLHF with General Function Approximation Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.833667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.833667Z digest=sha256:16b59f7517e0957dd8d678fcb8b36ab6dfc4371f134d05c0c0816cc6f536a0fa

Observation 3b0a6cb4-d055-45c1-963e-d178d68d1d07 · outbound

This paper cites Foster and Alexander Rakhlin.

Thompson Sampling in Online RLHF with General Function Approximation Foster and Alexander Rakhlin

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.515634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.872516Z digest=sha256:a0301ef053244de5f1fc2e47566c2ab0926bb8b06a310f145bb2dae781d0a8e9

Observation 7e17df10-3712-4376-ba25-69acb9567738 · outbound

This paper cites Scaling laws for reward model overoptimization.

Thompson Sampling in Online RLHF with General Function Approximation Scaling laws for reward model overoptimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.445967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.899602Z digest=sha256:d6dd716337b70aa1d0a1aab7ad864f5e6e6d70994461416066919c6a210a3579

Observation f93d914f-0c4c-473a-82e8-2c013645e3ae · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

Thompson Sampling in Online RLHF with General Function Approximation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.923648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.923648Z digest=sha256:db4617b484de218cc4319a9372284a62391f5f7fd55ff2180ca38bc81a52fb3f

Observation e83e8437-ccdd-42d5-b84d-f82d4a42da7e · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Thompson Sampling in Online RLHF with General Function Approximation Reinforced Self-Training (ReST) for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.949387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.949387Z digest=sha256:349ab6349c926fef5c01d2abd52f5d3e20486e2961a10098fc4e7b93a5b01505

Observation 577cacf1-c0d0-4a3d-81be-5207f551a71d · outbound

This paper cites Randomized Exploration for Reinforcement Learning with General Value Function Approximation.

Thompson Sampling in Online RLHF with General Function Approximation Randomized Exploration for Reinforcement Learning with General Value Function Approximation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:51.698037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:47.982915Z digest=sha256:609c7963ab8d21366a6813b84e741051e5377e891cfb4dd026dcbc15a79fd835

Observation f297a784-efaf-4ffe-89e3-7dcd0598597c · outbound

This paper cites Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo.

Thompson Sampling in Online RLHF with General Function Approximation Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.010070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.010070Z digest=sha256:d1d91635decba9880849edeffddc38596f0440ee0fe89655646028369999c6a1

Observation c42b36f4-36c4-48c0-91b7-2a61f39d6dd5 · outbound

This paper cites Learning trajectory preferences for manipulators via iterative improvement.

Thompson Sampling in Online RLHF with General Function Approximation Learning trajectory preferences for manipulators via iterative improvement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.372101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.044855Z digest=sha256:6cda33fcc59c9d33d55d9d2c2ca11d676fba066121ceb03f760bd072daf08b86

Observation 8fcfabfa-f0f1-4928-b963-d64cda7e98fd · outbound

This paper cites Schapire.

Thompson Sampling in Online RLHF with General Function Approximation Schapire

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.305869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.086054Z digest=sha256:693c06ccc2875751e77125ea9878020379475a0cb82084ba5e767c9e3c3208a0

Observation 8d7c2828-6114-47ee-9b1c-75a721be65d2 · outbound

This paper cites Bellman eluder dimension: new rich classes of rl problems, and sample-efficient algorithms.

Thompson Sampling in Online RLHF with General Function Approximation Bellman eluder dimension: new rich classes of rl problems, and sample-efficient algorithms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:56.160284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.220473Z digest=sha256:376f30cc8dcc2f3f05f6084cbdc91b155a16a989204e0b1a07d642faafbb01ef

Observation 0c2c094a-4035-4108-932e-73172d5f9e74 · outbound

This paper cites An Introduction to Variational Autoencoders.

Thompson Sampling in Online RLHF with General Function Approximation An Introduction to Variational Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.246384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.246384Z digest=sha256:90dff8ce0f3c1948c37c11276500e626768db8da5228551eae119173c15e28dc

Observation 42031251-04a4-4d60-98f6-1d92937d10ec · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Thompson Sampling in Online RLHF with General Function Approximation Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.277156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.277156Z digest=sha256:c71424f369c57d6d4fd40e43bc8c295d388da14231b508623d4ed638855b5463

Observation f0776b47-d595-408d-8199-0e20cbb831ad · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Thompson Sampling in Online RLHF with General Function Approximation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.349648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.349648Z digest=sha256:720b809aa4bcc7e6843f5994efb1a44169bc4d490a62c8e46c6bb01cff9e6738

Observation 5e9a7c7f-6d8b-48bb-a13c-f3dd10c93cb9 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Thompson Sampling in Online RLHF with General Function Approximation Statistical Rejection Sampling Improves Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.431444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.431444Z digest=sha256:f87db01697fba5a46da1c10219220f66d6f09f9475bf368763a07f91e72cd1cb

Observation 2d07115f-7ad9-4934-859e-b3cfd268c94f · outbound

This paper cites Roberts, Matthew E.

Thompson Sampling in Online RLHF with General Function Approximation Roberts, Matthew E

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.983442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.478015Z digest=sha256:7130ae7d17a135170d93e0cff002a8704d221778115c73e9e33e4734a860bce5

Observation 99e536a6-804c-4e8e-9d7d-6a8133fa2060 · outbound

This paper cites Dueling Posterior Sampling for Preference-Based Reinforcement Learning.

Thompson Sampling in Online RLHF with General Function Approximation Dueling Posterior Sampling for Preference-Based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.515394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.515394Z digest=sha256:5658831ef0beedc34b57ecf0244cad2ca8c19f22922d606610e6673dc44f9ccc

Observation d2dfc404-e7b3-4272-a6ce-8e481e070dfe · outbound

This paper cites GPT-4 Technical Report.

Thompson Sampling in Online RLHF with General Function Approximation GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.557964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.557964Z digest=sha256:25754fe543b44e7e6b76fd0216f68b90baf190102ac1a8628011350ad987e5e6

Observation d119f7ce-7e96-44a5-8ae4-32de28a7a17b · outbound

This paper cites Randomized prior functions for deep reinforcement learning.

Thompson Sampling in Online RLHF with General Function Approximation Randomized prior functions for deep reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.724220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.595572Z digest=sha256:00762d4a445dc79c65eac81031e7d877d3a34c78f93cc060fcde3f7d7bec3208

Observation 49fef49c-a677-43d7-aabc-9d0303095e41 · outbound

This paper cites Approximate thompson sampling via epistemic neural networks.

Thompson Sampling in Online RLHF with General Function Approximation Approximate thompson sampling via epistemic neural networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.458395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.687662Z digest=sha256:edbea01cc893a28b4bc2e24ceb279878e6ec2e395ef0de7b1310a818edb2654b

Observation a4149ebe-e20f-44e2-98f4-d8e5ea133972 · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:55.265065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.737226Z digest=sha256:3204bd49b425373776f18e7866f0a444f649db4c766f91195609e9c629d4c925

Observation 6f9f557f-4443-4c58-8032-dc79c85f99d1 · outbound

This paper cites Training language models to follow instructions with human feedback.

Thompson Sampling in Online RLHF with General Function Approximation Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.769424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.769424Z digest=sha256:d3a704afa981512517c9dafb97bc97f4ca968609284d721167f4a39b3eed32be

Observation 861b40f8-ed0f-40de-9e17-2b78ec60c871 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Thompson Sampling in Online RLHF with General Function Approximation Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.796716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.796716Z digest=sha256:bc229851441434975ec37d418e2349a63696e977011ed23eff1a1e033085ab54

Observation 211392fb-78f6-416f-88c2-2596ae83e015 · outbound

This paper cites Worst-Case Regret Bounds for Exploration via Randomized Value Functions.

Thompson Sampling in Online RLHF with General Function Approximation Worst-Case Regret Bounds for Exploration via Randomized Value Functions

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:51.321069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.814525Z digest=sha256:773954c54a08b94ae1bb692b7ff11682aeb07d83de733278c24ad4ab1b7ce988

Observation 55c69687-57ec-4809-8c77-83e041a8c371 · outbound

This paper cites Learning to optimize via posterior sampling.Math.

Thompson Sampling in Online RLHF with General Function Approximation Learning to optimize via posterior sampling.Math

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.197918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.838269Z digest=sha256:e918f55cdfc43b7c99dfd1f58841e4cf65c7d48974058c7cf56fcf3df7af775a

Observation bc4131b1-1e3a-4dc0-9db7-ff14eb6c686a · outbound

This paper cites Optimal algorithms for stochastic contextual preference bandits.

Thompson Sampling in Online RLHF with General Function Approximation Optimal algorithms for stochastic contextual preference bandits

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:55.096880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.861004Z digest=sha256:aaa0f4832fc669c34b81c22fb00f08943863734699fed99b4f4eeae0d71e3eae

Observation d14f95a0-6435-44a0-ad5c-650065971df0 · outbound

This paper cites Efficient and optimal algorithms for contextual dueling bandits under realizability.

Thompson Sampling in Online RLHF with General Function Approximation Efficient and optimal algorithms for contextual dueling bandits under realizability

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.840829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.890800Z digest=sha256:15fecce93587819cc9fd96cd624b1df6d0846d72737b7dc2011835d97549ea91

Observation 96c7ef00-ab1d-4d6b-afc8-59b1df7108b3 · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Thompson Sampling in Online RLHF with General Function Approximation Dueling rl: Reinforcement learning with trajectory preferences

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.465692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.921797Z digest=sha256:3a8de33a4b032c81c63476f5f5f12ba4d123bf03e373527a190df6cb537a677e

Observation 49191887-b0d3-43ff-95e7-cb122383c61b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Thompson Sampling in Online RLHF with General Function Approximation Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.940005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.940005Z digest=sha256:2788c7a9292ae9e010ae9fefe6bbc61987ead711c7bdf266d459bf206a23c399

Observation 279c74ef-69af-4a5a-80c5-1c6e4b99451a · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano.

Thompson Sampling in Online RLHF with General Function Approximation Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:54.117813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.969661Z digest=sha256:a0dca49f69d20ca874ddc2d0b93e2eb3d5b31fa971cb2411d211f98849d4bf2a

Observation fa5d1bcb-9c9a-48b3-8d23-0f6bb19a2314 · outbound

This paper cites Thompson.

Thompson Sampling in Online RLHF with General Function Approximation Thompson

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.933023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:48.992418Z digest=sha256:bb8c71d6f9a67796ce8b6c4f140f0af141fe5bf5eac7ed20c30796c4342289ca

Observation d0e6e6ee-2a2d-4ea5-9d0c-675fddd8b452 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Thompson Sampling in Online RLHF with General Function Approximation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.018024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.018024Z digest=sha256:0282e79290dd3cd7958dd865b2f739a8a267d3dee957900d35cb492342e61a25

Observation 5ade00be-bf41-4888-ae26-23ff03ca7444 · outbound

This paper cites van de Geer.

Thompson Sampling in Online RLHF with General Function Approximation van de Geer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.773838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:49.052190Z digest=sha256:8200c0d29b6e1743d9c7c55576af798494437dd9ce743f70d872f5dc657564c3

Observation 76980b50-eb90-47f5-9a42-dc7bc440db34 · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Thompson Sampling in Online RLHF with General Function Approximation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.083529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.083529Z digest=sha256:fa0c15ea9b17a034ecbcbd85bb87353bc34b52a25f602462e46a8f448cbcf8e2

Observation 7f75d0cc-fa25-45bd-bef3-aa7b602e3ef7 · outbound

This paper cites Thompson sampling for combinatorial semi-bandits.

Thompson Sampling in Online RLHF with General Function Approximation Thompson sampling for combinatorial semi-bandits

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.720079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:49.108031Z digest=sha256:ce6d2b4c0b6118cd022a0ff8e51d4618db7928bccc5f9770f4053c392e98c942

Observation 66e2e15c-29c5-4fef-b039-035424f36808 · outbound

This paper cites Is rlhf more difficult than standard rl? a theoretical perspective.

Thompson Sampling in Online RLHF with General Function Approximation Is rlhf more difficult than standard rl? a theoretical perspective

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.586179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:49.117791Z digest=sha256:ee9f1b143aa22db2582d1b83461a0e626d7a027a9b03fae804b8b742aa94f4b2

Observation 32e371bb-2530-41ee-91aa-d4f60570624f · outbound

This paper cites A survey of preference-based reinforcement learning methods.Journal of Machine Learning Research, 18(136):1–46, 2017.

Thompson Sampling in Online RLHF with General Function Approximation A survey of preference-based reinforcement learning methods.Journal of Machine Learning Research, 18(136):1–46, 2017

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.466104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:49.197952Z digest=sha256:435a15236437e68a9b3ea79eee9461663abbed726efb1b597c6bc044fc26b619

Observation aaaeabf1-1a17-4283-b3f8-afbfa8d4821b · outbound

This paper cites Making RL with Preference-based Feedback Efficient via Randomization.

Thompson Sampling in Online RLHF with General Function Approximation Making RL with Preference-based Feedback Efficient via Randomization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.305402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.305402Z digest=sha256:f1b614840ac046395d105771440b08cdecdea53ff730bb35c6c15bf3ddb5c42f

Observation 3bff1bf3-7f85-42d6-9c5b-5fa3c5a8987b · outbound

This paper cites Borda regret minimization for generalized linear dueling bandits.

Thompson Sampling in Online RLHF with General Function Approximation Borda regret minimization for generalized linear dueling bandits

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.309730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:49.376621Z digest=sha256:74a3e6404ad760840975090289c544ae0cbd8aaeab5f0abf0bb05b6760b471d1

Observation d633d1f9-2b01-4aa8-929c-e316d7ca4f71 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Thompson Sampling in Online RLHF with General Function Approximation Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.435213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.435213Z digest=sha256:e1053ef039de9bde3922c6a72888c44f96ec29a60f10e421fa0969ed3c5f4c3d

Observation 50e372fb-8206-41ca-be88-96913864589f · outbound

This paper cites Near-optimal randomized exploration for tabular markov decision processes.

Thompson Sampling in Online RLHF with General Function Approximation Near-optimal randomized exploration for tabular markov decision processes

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:53.149868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:49.519282Z digest=sha256:597fd4e124a26754cba3dea62a87181bd3ce38703966c104610ce3c76120a5f8

Observation 358a7f11-bc21-4f36-ae13-e03d2f64e846 · outbound

This paper cites Yang, Aarti Singh, and Artur Dubrawski.

Thompson Sampling in Online RLHF with General Function Approximation Yang, Aarti Singh, and Artur Dubrawski

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.943796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:49.630519Z digest=sha256:3c3021166a3ce042068e79acf3c401c8efaf16fd726c1f3dedbc02f78b820903

Observation 2224a913-2e3b-4ce0-89b8-bb0c530ca86a · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Thompson Sampling in Online RLHF with General Function Approximation RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.730231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.730231Z digest=sha256:ce595d3af54af1549be69d962266b6ec6297ce3653d1609f3ddf748c0a65f7ba

Observation efdf6433-3dac-4f16-8619-95bc291194bb · outbound

This paper cites The k-armed dueling bandits problem.Journal of Computer and System Sciences, 78(5):1538–1556, 2012.

Thompson Sampling in Online RLHF with General Function Approximation The k-armed dueling bandits problem.Journal of Computer and System Sciences, 78(5):1538–1556, 2012

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.725910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:49.832227Z digest=sha256:bacc85fb33c6a2d8d7abeb4b953c03836298d98f81b30694ee1063cf4fd4b7e6

Observation 74185e95-b077-4d1a-8ed4-77e557bb356e · outbound

This paper cites Frequentist Regret Bounds for Randomized Least-Squares Value Iteration.

Thompson Sampling in Online RLHF with General Function Approximation Frequentist Regret Bounds for Randomized Least-Squares Value Iteration

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:43:50.963293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:49.926426Z digest=sha256:b23718366d87d480077059fc7bea5a083be64e10ddba3b171380617333f526a6

Observation be121500-6594-443b-9b65-15b395ac386f · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Thompson Sampling in Online RLHF with General Function Approximation Provable Offline Preference-Based Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.030374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.030374Z digest=sha256:a0d3a38833e48ddf130e6d3253622c990e6b7201f48c39c80f22d1152aac8652

Observation 52573980-9b4b-4c97-91d0-417248f344a7 · outbound

This paper cites Lee, and Wen Sun.

Thompson Sampling in Online RLHF with General Function Approximation Lee, and Wen Sun

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.502747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:50.138027Z digest=sha256:77e86adb2aa3fd4251cebffcaf702a34fb0cf245da96670e6c889d7199ddf300

Observation 0b8a432e-e368-4c56-b05f-05153a48c8e2 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Thompson Sampling in Online RLHF with General Function Approximation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.251391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.251391Z digest=sha256:d28230d1711dd90dd4b32a5463074014e35d7c3161915084f99e0215a81e13af

Observation 03c377c5-e3ab-436d-8210-6ccd562be404 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

Thompson Sampling in Online RLHF with General Function Approximation Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.418201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.418201Z digest=sha256:88d960222ae82cbd828997e18c7b3c754df06c63d0ed524ee34fa9ad76647611

Observation 77e137d6-8164-40ce-abb3-10a5008f955d · outbound

This paper cites Efficient active learning with abstention.

Thompson Sampling in Online RLHF with General Function Approximation Efficient active learning with abstention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:52.301764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:50.546951Z digest=sha256:0b33e259efa5cca141c3743d6335e3fcf0c1f23d94bb63a4b15f4c6501f50209

Observation b14aa731-9812-46c9-a59b-4c5cdea52644 · outbound

This paper cites an unresolved cited work.

Thompson Sampling in Online RLHF with General Function Approximation Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:52.029084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:43:50.639923Z digest=sha256:894ae4d9afd4f645ab723770cc5fb8250351c990f27590b5f106131932475aac

Pith citing papers

No inbound Pith citation observations are available.