Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:59:51.021706Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2506.09183.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:59:51.021706Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6f6bb597-f374-4dc0-9f5c-80702e8d0973 · outbound
Multi-Task Reward Learning from Human Ratings F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32e67037-1214-456a-be17-292daea69630 · outbound
Multi-Task Reward Learning from Human Ratings Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1ddc9f2-0cb7-47ac-9dbe-9e248d529878 · outbound
Multi-Task Reward Learning from Human Ratings Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a3f0b867-8090-403b-99ff-95b8e16ce518 · outbound
Multi-Task Reward Learning from Human Ratings Playing Atari with Deep Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 198b248d-6de3-42a4-93d3-e0d195550c7f · outbound
Multi-Task Reward Learning from Human Ratings Performance Optimization of Ratings-Based Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 193da9e2-238b-4da1-b164-5096c803f4a9 · outbound
Multi-Task Reward Learning from Human Ratings Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d213a956-cdbe-4e94-ad33-4dadd1eb6b94 · outbound
Multi-Task Reward Learning from Human Ratings Proximal Policy Optimization Algorithms
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f6a1e6-9583-4904-9999-8c30e3110ad4 · outbound
Multi-Task Reward Learning from Human Ratings Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5dc880e9-1868-4bfc-82c9-3e2c3b0c3851 · outbound
Multi-Task Reward Learning from Human Ratings Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a9c1ad1-640c-4584-826a-93eef9aba65b · outbound
Multi-Task Reward Learning from Human Ratings DeepMind Control Suite
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 860d5170-a297-4dbd-a06f-a95d0eb4ea47 · outbound
Multi-Task Reward Learning from Human Ratings Mujoco: A physics engine for model-based control
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92368bed-3358-4f79-be95-3e301767955c · outbound
Multi-Task Reward Learning from Human Ratings J., Waytowich, N., and Cao, Y
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 48c9184b-f062-4a8c-afa4-d0394dbd217d · outbound
Multi-Task Reward Learning from Human Ratings Value of potential field in reward specification for robotic control via deep reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2bcbd690-7010-4138-9b8d-6227b1256353 · outbound
Multi-Task Reward Learning from Human Ratings Offline reinforcement learning with failure under sparse reward environments
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5ba111c-b936-4da8-af73-df95b2eae750 · outbound
Multi-Task Reward Learning from Human Ratings R., and Cao, Y
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93cd9e43-9014-42cc-9656-347a30e9c078 · outbound
Multi-Task Reward Learning from Human Ratings write newline
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.