Pith. sign in

Paper Citation Record · LEDGER

Multi-Task Reward Learning from Human Ratings

As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2506.09183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09183 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:59:51.021706Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f6bb597-f374-4dc0-9f5c-80702e8d0973 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Multi-Task Reward Learning from Human Ratings F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:49.760680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:49.760680Z digest=sha256:6d739d84fbe6beda08e1080bda9555aed881e043bdb38def3f8876ff0aee20ee

Observation 32e67037-1214-456a-be17-292daea69630 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Multi-Task Reward Learning from Human Ratings Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:49.832394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:49.832394Z digest=sha256:078ec045e8d06e454bd4fa6abaad3761c854f5af18545af29fde5c6f8d7a5a1b

Observation e1ddc9f2-0cb7-47ac-9dbe-9e248d529878 · outbound

This paper cites Multi-task learning using uncertainty to weigh losses for scene geometry and semantics.

Multi-Task Reward Learning from Human Ratings Multi-task learning using uncertainty to weigh losses for scene geometry and semantics

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:52.806826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:59:49.916874Z digest=sha256:bdb761315406849eadf342b7c2dc2c047e486e6fda42ca71e5ff6b791c59e1d2

Observation a3f0b867-8090-403b-99ff-95b8e16ce518 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Multi-Task Reward Learning from Human Ratings Playing Atari with Deep Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:49.996007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:49.996007Z digest=sha256:f4899878cdd5de5adbadcbaf1c7ccb40f787ae4eb4e918d101b16476dc175aff

Observation 198b248d-6de3-42a4-93d3-e0d195550c7f · outbound

This paper cites Performance Optimization of Ratings-Based Reinforcement Learning.

Multi-Task Reward Learning from Human Ratings Performance Optimization of Ratings-Based Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:59:51.596559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.092283Z digest=sha256:2af7040db06d21c29563e442af0f68d438f2856c0dd8bfe22f64d617c763458c

Observation 193da9e2-238b-4da1-b164-5096c803f4a9 · outbound

This paper cites an unresolved cited work.

Multi-Task Reward Learning from Human Ratings Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:59:52.709502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.156703Z digest=sha256:f40fec1e9c82996450eae9887494d54fb72941467b486d61612a958dccebcb15

Observation d213a956-cdbe-4e94-ad33-4dadd1eb6b94 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Multi-Task Reward Learning from Human Ratings Proximal Policy Optimization Algorithms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:50.248231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:50.248231Z digest=sha256:00813cd9a0f0878adbd7cf05cdac080dcaf902d701c38416a5f2a5bbdf6d07fb

Observation 79f6a1e6-9583-4904-9999-8c30e3110ad4 · outbound

This paper cites an unresolved cited work.

Multi-Task Reward Learning from Human Ratings Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:59:52.556661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.328564Z digest=sha256:f07651b8c6e7e838e8de569db3326d6c80f167169d6566799a6df081b7fd094b

Observation 5dc880e9-1868-4bfc-82c9-3e2c3b0c3851 · outbound

This paper cites an unresolved cited work.

Multi-Task Reward Learning from Human Ratings Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:59:52.281107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.424254Z digest=sha256:c9b0f976c5d7fbd9ed12168e39229552570d109fed29cce1037c1c1c5badf7f6

Observation 8a9c1ad1-640c-4584-826a-93eef9aba65b · outbound

This paper cites DeepMind Control Suite.

Multi-Task Reward Learning from Human Ratings DeepMind Control Suite

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:50.587696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:50.587696Z digest=sha256:fd12dccd48e9ccb1e57283eef3b3d48d68430322045525ae5265bce01211340c

Observation 860d5170-a297-4dbd-a06f-a95d0eb4ea47 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Multi-Task Reward Learning from Human Ratings Mujoco: A physics engine for model-based control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:50.678056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:50.678056Z digest=sha256:0295fc75ed92497c9a25b01aaa4bd00145bf3031214e87dfcbbae45f00fda138

Observation 92368bed-3358-4f79-be95-3e301767955c · outbound

This paper cites J., Waytowich, N., and Cao, Y.

Multi-Task Reward Learning from Human Ratings J., Waytowich, N., and Cao, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:52.122919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.777524Z digest=sha256:d1fd99ed9eb81c0f02f640166f5921089325930dd7419b1a9dd5cb7514cedb87

Observation 48c9184b-f062-4a8c-afa4-d0394dbd217d · outbound

This paper cites Value of potential field in reward specification for robotic control via deep reinforcement learning.

Multi-Task Reward Learning from Human Ratings Value of potential field in reward specification for robotic control via deep reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:52.004313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.847757Z digest=sha256:6f0b1ae924feb478c7e1393a8992e07953cc75752749dddd6f57083c24a5cb60

Observation 2bcbd690-7010-4138-9b8d-6227b1256353 · outbound

This paper cites Offline reinforcement learning with failure under sparse reward environments.

Multi-Task Reward Learning from Human Ratings Offline reinforcement learning with failure under sparse reward environments

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:51.876339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.891547Z digest=sha256:944949f27127870c1778dfe281a8a32696a92f9e1423279a3d44a413cb7caf65

Observation b5ba111c-b936-4da8-af73-df95b2eae750 · outbound

This paper cites R., and Cao, Y.

Multi-Task Reward Learning from Human Ratings R., and Cao, Y

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:59:51.343300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.949341Z digest=sha256:dc7bbe2563be08a9bb57abd1beb24c5386b125e05eac6dffeb795ba4449d7d2a

Observation 93cd9e43-9014-42cc-9656-347a30e9c078 · outbound

This paper cites write newline.

Multi-Task Reward Learning from Human Ratings write newline

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:51.021706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:51.021706Z digest=sha256:c37f1613536543a1fb115d52b6f077d5e9aa963473ed25a381a18958b5a3afbf

Pith citing papers

No inbound Pith citation observations are available.