Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-01T00:27:39.321158Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2605.04477.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-01T00:27:39.321158Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ac52137f-0dfd-4b70-ba0a-5242fd9060fa · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Abbasi-Yadkori, D
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d2f0328-4118-4919-88fb-c3732ed2f9f9 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Online Least Squares Estimation with Self-Normalized Processes: An Application to Bandit Problems
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 66a85e3f-1c65-4521-8204-01fc191a02e6 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7b97dbb8-36ec-4ea5-be3b-368108431a23 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 59201b32-1257-40de-88fe-897dc28932cf · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 66ea7f5d-b4a0-4328-a510-91626b9c075a · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7a9fdac5-4d95-473d-bd4a-9286626ed371 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90272d87-abbb-4d79-8a22-333ef821c809 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9efe4dd6-a8ed-4ecf-9bd2-3fe5f8687293 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b79f5054-d602-481f-b0c5-e88b6c475db9 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1cfbeb39-5772-4cf1-87ff-b161addf6b30 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Training Verifiers to Solve Math Word Problems
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f482ad3-ca24-4ceb-b281-713a95ec8b78 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 46b304d5-ee22-4ca3-954b-a6aa53e05467 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7c9527d3-0653-49f7-a9f0-54ffaa087035 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Dwaracherla, S
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a5303a5-7923-4d03-9a3c-e3cf9e802322 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b14ac81c-580e-4e83-a32e-a52c9ba14453 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Hendrycks, C
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e7db404c-5e41-4dd2-b6dc-774931c8b50d · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 724e391a-7667-4e6b-9198-de01403f8f90 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe4aadb7-74b9-4a75-92f4-54b29626623d · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5444e995-32b5-4ee4-915a-1eec75cef248 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d431dc00-ee7e-4d9b-86dc-7367b19acf20 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 68dddbd1-720d-40d7-af77-7963577cb0b1 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3b3419c9-0a24-4988-93c5-61ab5f6ce848 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 528f463a-61e5-4873-83e7-e6c332bcc039 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Ouyang, J
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6124bd09-3912-4a28-8d51-701fbf7b9fed · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Rafailov, A
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5157bf7e-6766-43e0-9d5f-8b6426faa818 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a8e835f5-b00a-45b1-9e56-e379ddcbcfe4 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Proximal Policy Optimization Algorithms
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a70acb00-a1dc-45e2-9587-e5226f3a5188 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Sherman and W
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9bd38557-1297-4da1-8826-c588fb31939f · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 147d0e39-30c0-4dd4-aa95-3b3d87ae0173 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dfdd086b-4452-41ad-96de-66c7666fd2de · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00237546-cee1-417e-8e8f-c4d121cbf96d · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 507e69ee-2fcf-496e-9aa6-a591109b1581 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cd512669-f46d-4609-8894-6fad6d930528 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 011847c1-0ac6-4c7a-aea4-405c0aca7de3 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Xiong, C
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 982bf544-798f-40b6-b81c-1003b1591837 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf74eceb-12de-4c8e-bcff-d0a127d82a6d · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Zhang, H
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a34848c7-db17-4966-bfdf-1ae8fd5908ea · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Zhang, S.-A
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04cdc8d3-3859-4aaa-a0d2-d389120ca9a6 · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 754f1ef1-1628-4e0b-b4f3-c6b36a820e0c · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94b7afdb-a103-4396-80ad-e6776af0e48f · outbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.