Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:47:06.636420Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2504.14732.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:47:06.636420Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
12 of 12 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7dc0ae4c-2657-497d-bf44-2b8eb12c8231 · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Proof of Theorem 3 The optimistic reward function is defined as follows: R(bwn,τ ) = min R(bwn,τ ) + 4Kexp(4B) ηλmin(ΣDn) r C2 2n log 4 δ, K− 1
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 83dd88f6-48f6-4f8b-9526-e7840fc9eb7d · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Playing Atari with Deep Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9422b27b-8d00-417b-b05b-bc5ffe9b67d1 · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd920b3a-51b6-46b1-ba3c-441ec1baea27 · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Benchmarks and Algorithms for Offline Preference-Based Reward Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 073d7591-ac3d-45bb-bbc8-532936bd7d35 · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Is RLHF More Difficult than Standard RL?
Reference 2005
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b15f9e9f-7c11-4329-8018-8987f5ea50ea · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Tamer: Training an agent manually via evaluative reinforcement
Reference 2013
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0d9849dd-3b73-4e28-9b55-105f62577960 · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Provable Offline Preference-Based Reinforcement Learning
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34bfff6d-8e60-41a9-b0fc-719573c10ef4 · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Inverse Reinforcement Learning by Estimating Expertise of Demonstrators
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03cfa0b1-3e01-4f43-97c0-6fbb73a6b765 · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Inverse reinforcement learning with learning and leveraging demonstrators’ varying expertise levels
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8b0bccfa-c0d6-4eb7-8844-7b91015e5161 · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Diversity is All You Need: Learning Skills without a Reward Function
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 045ae104-de2c-4a14-a542-e9a172f9d9da · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback Regret Bounds for Discounted MDPs
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4d25f93-6eda-42bf-9b74-568c7980d848 · outbound
Reinforcement Learning from Multi-level and Episodic Human Feedback On Lower Bounds for Regret in Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.