Pith. sign in

Paper Citation Record · LEDGER

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback

As of 6 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2605.04477.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.04477 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T00:27:39.321158Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact10
  • verified fuzzy9
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ac52137f-0dfd-4b70-ba0a-5242fd9060fa · outbound

This paper cites Abbasi-Yadkori, D.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Abbasi-Yadkori, D

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T08:43:28.785978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:2ccb86c51e5e9b6c9e3cdb5f4c0bb3dd393ac7c338a1704885ddf837a59adcf1

Observation 8d2f0328-4118-4919-88fb-c3732ed2f9f9 · outbound

This paper cites Online Least Squares Estimation with Self-Normalized Processes: An Application to Bandit Problems.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Online Least Squares Estimation with Self-Normalized Processes: An Application to Bandit Problems

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:35:10.418229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:92fcf5dc5f3be62eacb58303de56f073e5f793c20b5ad865d1b09239177a89b0

Observation 66a85e3f-1c65-4521-8204-01fc191a02e6 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.796424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:9def772efa29245707573aea737899cf65a97ef4c9c63929d9b514b78834238f

Observation 7b97dbb8-36ec-4ea5-be3b-368108431a23 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.796828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:fb46d1574879ca556be40b6d6fa71e35b12a5cb28e7aa3f1bb6f2c42e57ee3f3

Observation 59201b32-1257-40de-88fe-897dc28932cf · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.778932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:805dc2ee12b3dcc08e72c86bfc2e288d8fd2dd9b47987fb7c9363d424fef4a78

Observation 66ea7f5d-b4a0-4328-a510-91626b9c075a · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.421147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:6f45bdde5cbc2a42507ca6f661742342130597543347e8d3e06ab1a791d031f1

Observation 7a9fdac5-4d95-473d-bd4a-9286626ed371 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.792798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:be5a7a592740a9e01bda195a1f95ff80f1c5229ec59faa0773741bec97cd5299

Observation 90272d87-abbb-4d79-8a22-333ef821c809 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.791510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:6c56ee3edf25cfef2329958608cb83c6e201ae484b6b187ad97f52b56820759b

Observation 9efe4dd6-a8ed-4ecf-9bd2-3fe5f8687293 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.780933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:5c9f04b8b37722289a33171e5d0a76deaf773e416c43e632efc39814eb9600e5

Observation b79f5054-d602-481f-b0c5-e88b6c475db9 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.784289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:a809463cc22b38579ed5932bfe6c6d068fed1b7c17db3d297ce0ba0573494b78

Observation 1cfbeb39-5772-4cf1-87ff-b161addf6b30 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:35:10.415630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:c3326e19c46a2dba547390d01fa0870177ab200f1ad1a6492b752042890370c5

Observation 1f482ad3-ca24-4ceb-b281-713a95ec8b78 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.789130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:87d232c83a8f3dc833da81feb409e7acb17345d1320b408b5bce77978d2463ef

Observation 46b304d5-ee22-4ca3-954b-a6aa53e05467 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:35:10.421157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:69f8bb95308d5f14f68a740d0cabcae9724d88f57ef3c8875959fc5e975c27ff

Observation 7c9527d3-0653-49f7-a9f0-54ffaa087035 · outbound

This paper cites Dwaracherla, S.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Dwaracherla, S

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T08:43:28.798501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:b20a3653107a47a987a404537adc0b3a8f56e837f5328797bbd11cc08479f2e6

Observation 1a5303a5-7923-4d03-9a3c-e3cf9e802322 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.418679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:d75a66fa49cb38b87afeb1b0088ae9c695358364750c23b5f8fd5509b47b715a

Observation b14ac81c-580e-4e83-a32e-a52c9ba14453 · outbound

This paper cites Hendrycks, C.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Hendrycks, C

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T08:43:28.774075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:56a4f8c659964089b56726bb3ee1fe1b200d7b45d004784ae253c3f282d758a2

Observation e7db404c-5e41-4dd2-b6dc-774931c8b50d · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.795110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:cf6670276254a2535103cd9a4ddc2ba7095d4f5930432008a3f5a83644fda4fe

Observation 724e391a-7667-4e6b-9198-de01403f8f90 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.782442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:bac69bef80d0bdb143777a721b5d02bee036ab7a04eb99811d3af221df934373

Observation fe4aadb7-74b9-4a75-92f4-54b29626623d · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.784100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:56faf46aacbdbbe8a3a08d4ed028e21fd46667b2fcb2f30102a4ee6e1b93f164

Observation 5444e995-32b5-4ee4-915a-1eec75cef248 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.779135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:7ef69733492a4a11941b36eb04b314effe3c6f433e247d79f88fffd5b2e49b53

Observation d431dc00-ee7e-4d9b-86dc-7367b19acf20 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.752413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:c7ea7422a9996da8bf763cecd3117b83fd6b8dcf36ebb63a5e5722e1f191fd86

Observation 68dddbd1-720d-40d7-af77-7963577cb0b1 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.746772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:53dca596ebebc5a20db79bf079456399a1288152e448ad8041e93fee4cd6bce3

Observation 3b3419c9-0a24-4988-93c5-61ab5f6ce848 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.782635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:22669f323625a1879c4daae3118314d5c109e15ab31c46d66fdd781ca137d4da

Observation 528f463a-61e5-4873-83e7-e6c332bcc039 · outbound

This paper cites Ouyang, J.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Ouyang, J

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T08:43:28.787447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:ce873e3a811ab87d40d3aa7951eb9f5b494fc20ebd95ad918d83177ddb7e569c

Observation 6124bd09-3912-4a28-8d51-701fbf7b9fed · outbound

This paper cites Rafailov, A.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Rafailov, A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T08:43:28.777526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:9c30e214053015d5f0b3b478ce4c729353ea28f69fb571aa959afff13f808d26

Observation 5157bf7e-6766-43e0-9d5f-8b6426faa818 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.785810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:2afb16767ca067289b3cfdcbb8a7e603931fc0b8cd02ccd9c9df76c081879803

Observation a8e835f5-b00a-45b1-9e56-e379ddcbcfe4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Proximal Policy Optimization Algorithms

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:35:10.423898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:add6de7e055f010bf4ba3f9a7ef0b5c393f5ff82060f87394a07584f0d1565b2

Observation a70acb00-a1dc-45e2-9587-e5226f3a5188 · outbound

This paper cites Sherman and W.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Sherman and W

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T08:43:28.758420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:f7dec018949b1d8d1fd6ee380085cf0b873a13198f88d6483b2c3a662f18f4ce

Observation 9bd38557-1297-4da1-8826-c588fb31939f · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.787643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:c912e028c9bcf5ebee90b3e293fcd0db07413831eba408476896a71f8e37286a

Observation 147d0e39-30c0-4dd4-aa95-3b3d87ae0173 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:35:10.413107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:06d8c686f148e2d3402a6bfac47d78afc0720fd3c9f583267cb3e8f3b1dcf821

Observation dfdd086b-4452-41ad-96de-66c7666fd2de · outbound

This paper cites Representation-Based Exploration for Language Models: From Test-Time to Post-Training.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:dd8b8dfbf928f353054d80af9e4e6492c9309e060902037f0d00ae4ae00e8627

Observation 00237546-cee1-417e-8e8f-c4d121cbf96d · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.761434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:a5c47196c35d1f4e0051a23ff2595bf4d6794d509d6e2ce7c0b1504521987e0b

Observation 507e69ee-2fcf-496e-9aa6-a591109b1581 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.770518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:129ac04119838d1487976ff8f69e418539291d031f6aee8027b63f5c75b169de

Observation cd512669-f46d-4609-8894-6fad6d930528 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.789832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:04a8d30b6bd863da7892872960c138af7de1182e9a951f499df139b659eb0426

Observation 011847c1-0ac6-4c7a-aea4-405c0aca7de3 · outbound

This paper cites Xiong, C.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Xiong, C

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T08:43:28.772066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:c33b9d62f2fcbf427e58aa530759dd9ae79ea35d242a137a6cf04beff5d49a1a

Observation 982bf544-798f-40b6-b81c-1003b1591837 · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.410391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:193307f30e935a44409e8b09ba15bbee6da6361732fba726ccae14d003cbed8e

Observation bf74eceb-12de-4c8e-bcff-d0a127d82a6d · outbound

This paper cites Zhang, H.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Zhang, H

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T08:43:28.777130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:5684629c01889afc5d9f2793ac3b83a06ec40f8072d312d628d1996cbbca6cd2

Observation a34848c7-db17-4966-bfdf-1ae8fd5908ea · outbound

This paper cites Zhang, S.-A.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Zhang, S.-A

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T08:43:28.793228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:ddc78c47fe03df9c8ad0b135fc141945cc7cee1f7d4a83c061d77ce0e9b0441c

Observation 04cdc8d3-3859-4aaa-a0d2-d389120ca9a6 · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.754696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:882214038ce49e42ecc516e079180fe50cb4403b2ea34de60f84314ee5ef8728

Observation 754f1ef1-1628-4e0b-b4f3-c6b36a820e0c · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.407105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:b831b10f5cb1030f2ca3b772a355da44a75000cc2035c6e4ee64d127b99ba1a4

Observation 94b7afdb-a103-4396-80ad-e6776af0e48f · outbound

This paper cites an unresolved cited work.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-07-07T08:43:28.773665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:c78d28e2e69da0763f11efe3f5f2b7729fa4238939d63b703ef51498788b3ee1

Pith citing papers

No inbound Pith citation observations are available.