Pith. sign in

Paper Citation Record · LEDGER

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning

As of 6 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 1 inbound Pith citation observation for arXiv:2606.03962.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.03962 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T10:59:44.092482Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:41:12.498785Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

84 of 84 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved75
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0334e11-9b62-407c-b3cc-2954d0fccbbb · outbound

This paper cites Mastering the game of.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Mastering the game of

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:b05566af7fdddf59c9b62a531e09ca12b7d7413c6a6c42579ecac5fdc06a3f71

Observation e45bb1cd-831d-425a-b45b-f220c994391c · outbound

This paper cites Olympiad-level formal mathematical reasoning with reinforcement learning , journal =.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Olympiad-level formal mathematical reasoning with reinforcement learning , journal =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:0abfed0f7f17aadb1d7e65dc75876acbb921f199053b679465cd9f5e1e2629bd

Observation 3852df05-fc32-4efe-a773-0cec3f69c399 · outbound

This paper cites Pawan and Dupont, Emilien and Ruiz, Francisco J.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Pawan and Dupont, Emilien and Ruiz, Francisco J

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:a5b81e9b2c8e789ae87b12b388d7a0e345c86252b7da6ff30748f83f92a8afa0

Observation 6489a272-e33b-45e0-9453-591bcc1db965 · outbound

This paper cites Proceedings of the International Conference on Learning Representations , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Learning Representations , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:751b53f84b29763b66e483a576bb1a66ccd6eb5df653546ab599651f1aa5828a

Observation 7a376905-93aa-405b-9ddc-d4c3f25ecbf5 · outbound

This paper cites Proceedings of the International Conference on Learning Representations , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Learning Representations , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:32cacc5e9fd7092b1b5cba38692391a7ae7176fba6364dfd55a8c77bde4ade68

Observation 7b0d20fa-b12c-4f5e-b204-bf548fb277f6 · outbound

This paper cites McLean and Peter Norgaard and Zahra Shamsi and David Smalling and James Thompson and Subhashini Venugopalan and Brian P.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning McLean and Peter Norgaard and Zahra Shamsi and David Smalling and James Thompson and Subhashini Venugopalan and Brian P

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:cad7abfca427cd594086a786414ac261c69bf9495fcb54bbdf82d7cf58748dc7

Observation ce6e45ad-4ac9-435d-a34f-78fe5e155754 · outbound

This paper cites Proceedings of the International Conference on Learning Representations , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Learning Representations , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:c6f71f4d799da96a4c809a494877a21487216ff02525c5e542fec3c2b106af8b

Observation 7dfd984b-fe57-47f4-b7c2-b521c0ef9031 · outbound

This paper cites Proceedings of the International Conference on Machine Learning , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Machine Learning , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:4977e22feb95575a71ddbc3a9974218c90d447cf935be89bc1fd1a96ccb78a63

Observation 67e969d6-5bd3-466c-a9de-6b15e93ef6be · outbound

This paper cites Proceedings of the International Conference on Machine Learning , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Machine Learning , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:f4273edf54cdaeb1955f224f519a3b95fddbf261c04dd379a0c04836b2af1c01

Observation 0c7b5a01-b716-44b9-8fcd-a2d95a688f2f · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Advances in Neural Information Processing Systems , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:c3025e2ae3e167d2fdbd538d7ba101f0b000d4d6d65010986de22b5b6a102328

Observation adb501ea-c47d-4479-9437-d7ec066e45e0 · outbound

This paper cites Proceedings of the International Conference on Learning Representations , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Learning Representations , year=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:aca5cad736fa1094631b13e26c866a880bf806e48cff1311b7ed3497f92c2f01

Observation eed184b3-4317-4ac5-a57b-07e48e8baf58 · outbound

This paper cites Jointly Reinforcing Diversity and Quality in Language Model Generations.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:26.270305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:0e9eaa3d916d0b29d06f894c7fb2b4d981174c042ded9d070d6a65f1aab6207c

Observation 15d71f0b-372a-442e-85e9-ac7a4cd61ee3 · outbound

This paper cites Proceedings of the International Conference on Machine Learning , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Machine Learning , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:ae2c7c7796293e829ae7c37604964315f222ac9330b14d403ca1d0eac8e0ece4

Observation 27db969e-c4c6-4309-b5a9-a67756535c29 · outbound

This paper cites 2025 , eprint=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning 2025 , eprint=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:5db9e610997619a67dd3e1501b7701f00fe7c3ac2ffd692733c27852514cb4cd

Observation dfd8f3b6-8b46-45f0-9a40-85f04ca4f540 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Advances in Neural Information Processing Systems , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:f08f8a13b01ac06ef2975023a8c06d045494f7a175d63e3f4c52b7b23ed62d03

Observation 7855d5a1-25be-4149-8ee8-2971df240daa · outbound

This paper cites and Barto, Andrew G.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning and Barto, Andrew G

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:1c35f2f288d51bd5c76a0dea7ec07d91b01ef9a337b75596d0f38e2da6c1d738

Observation e65eeebe-823e-45b5-9657-9d69c31413e0 · outbound

This paper cites Machine Learning , volume=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Machine Learning , volume=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:089d3e594d7d6f263d7b10b281cf6d789e5af655afff6a9e2ed6d30439c98881

Observation b5a0d56f-637c-443a-9683-c59102396a16 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Advances in Neural Information Processing Systems , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:f904d871fdcd0ff0e3dd6dfdb9365bf4e24052ccb4c1e1589348ffd3c535a30b

Observation 61c54d98-d33b-45b6-9284-e24b1263093e · outbound

This paper cites Proceedings of the International Conference on Learning Representations , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Learning Representations , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:46a6c183aa26c6ccc62337cad4be41d5f2e45c68ec6d33a80664bc31079b1229

Observation dedab193-fbce-4fea-8ec7-c995c64eda08 · outbound

This paper cites 2016 , publisher=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning 2016 , publisher=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:28a889b2bfeb9f33f8deb0f39b724f87f289e100a727c33930fd0be91e6fff4b

Observation db1085b1-fb21-481a-8012-8bea94ad6acd · outbound

This paper cites Machine learning , volume=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Machine learning , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:5e6f2d353311bbacc8bc906db989014fb5132ea369497ca11224a64e84df2e67

Observation d3d330d4-2cf5-413a-8bcb-ffada8b3737b · outbound

This paper cites an unresolved cited work.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:a29b023f98bee20a8c4838a85ac8fb9175a3f4e2a4ebc9e93b4cc897442d4d4a

Observation 1338b182-987f-4a36-af92-1ce7dab2f49d · outbound

This paper cites Autonomous Agents and Multi-Agent Systems , volume=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Autonomous Agents and Multi-Agent Systems , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:d84e8dc9709b2a157bfc3cb95b16ad58647548b7b9858bb10c5e76da7d2adb09

Observation 573d8d2f-e6df-47e2-b109-84b54bda503f · outbound

This paper cites an unresolved cited work.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:4158075fd40ff80a368a936b94ff3fe602b82d6fbd5f59972efd765e617e30ed

Observation ffe32597-7c30-4bcc-ad63-5f96221ca80e · outbound

This paper cites Proceedings of the International Conference on Learning Representations , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Learning Representations , year=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:17b690e01a0299960836fc1577545ac0ecd441d08980b2510e364fd99961b363

Observation 1f384100-3b5f-406d-a9ea-d6331cb995e3 · outbound

This paper cites Linearly-solvable.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Linearly-solvable

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:a7a9d28887cce3bc7621900c5318848949eacb1db06c7b746a6edf5c9138e717

Observation 53760e6c-6f37-49e2-8b2d-45e69df3cc0a · outbound

This paper cites , author=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning , author=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:a7cfe2cc6b1280e32101636da25c0c2a708ecf9f51487b086d569e46e2482f2a

Observation 1e7972be-9860-42a0-bbce-959b60eb6f4a · outbound

This paper cites Proceedings of the International Conference on Machine Learning , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Machine Learning , year=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:ff53a09ed527e8fd89847573439561479fb91c796496070853a22f7f00353e82

Observation 30eb67e9-ea52-4e5a-9714-48c48c925c33 · outbound

This paper cites Journal of Artificial Intelligence Research , volume=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Journal of Artificial Intelligence Research , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:9ef3dc25ba0be08f00e07270b462fdc78d63ea259f964a2bb796ae36d618eded

Observation 7bca290c-4c68-4815-8e66-ad8da3569fcd · outbound

This paper cites Constrained.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Constrained

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:edb0848c55fe19177d2fe2988a68d27f4c25b239ac75d724a65eade998916f9f

Observation 67c51fed-1526-40bc-8877-8aba06aafeb8 · outbound

This paper cites , author=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning , author=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:82f887474d3c5ba86db11ee02c4a9324185b91c6d08f197f62cb610df0668925

Observation 86e49675-bf1a-41d3-88a2-b52ea730fd48 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Advances in Neural Information Processing Systems , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:0d5bb560d7a588c11a4228386eacd985edab769f4c7977de42e899e307c2600f

Observation b517656e-46fa-4cd3-a0b7-6e915d970682 · outbound

This paper cites Deep reinforcement learning with attention for slate.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Deep reinforcement learning with attention for slate

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:31d5c353491d22c5bce68be593b620daacff9dbb0e512282bc50849a610daeb5

Observation 81faf75b-b6b4-4bcf-929e-00d8c7fe7c1c · outbound

This paper cites Non-deterministic policies in.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Non-deterministic policies in

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:198e4fcc894678687187a924b5462d9868f75fd1232cf4c5b61dfc7b638842cf

Observation 92735ab3-27bb-4e44-8519-a8f4cacdfaa8 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Advances in Neural Information Processing Systems , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:a2ffb1527efa8edb759396319bc70012c7949edb0764711b163cc7d814cd8f7d

Observation 976fc8c1-8753-4c1a-a043-f1644897b69f · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Advances in Neural Information Processing Systems , year=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:a0a62a4e2ca32f6433eeff4e5623b61fe1ed38c01b7943497354eebdfe15ef60

Observation c97bffd0-4573-4712-99bf-fad58e629c7b · outbound

This paper cites Proceedings of the International Conference on Machine Learning , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Machine Learning , year=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:3e1dfb365e892775644802f70aab25eed4f3619e6f34a3672253282f8cb8fd5d

Observation 236f4585-8999-4734-a020-1bfd9b421313 · outbound

This paper cites , author=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning , author=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:34e2c5c15f2234652a621e17bc7a07d96789cecf11eb900cb6539558765940e4

Observation f36f5ef8-d9ed-4bb8-a715-ffb991e69e4f · outbound

This paper cites Journal of Artificial Intelligence Research , volume=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Journal of Artificial Intelligence Research , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:85505b518228a26f8877a8d870be43d214159e0e266af490960e52c408ec7db2

Observation c34d1c6c-0005-47c5-b351-39242b233fb6 · outbound

This paper cites Breaking the bias barrier in concave multi-objective reinforcement learning.arXiv preprint arXiv:2603.08518, 2026.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Breaking the bias barrier in concave multi-objective reinforcement learning.arXiv preprint arXiv:2603.08518, 2026

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.264707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:9304a55a2b914286a0188fda12723ecef55b2be0407db0f4f707f625d7a14915

Observation 811acb8f-5fcf-4baf-a79b-076a3b99727d · outbound

This paper cites Puterman , title =.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Puterman , title =

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:5257b13812f32c6dce5a3ae06c6a7939b67b5de6a7aaf8ffa7ac7b255f4c4b20

Observation d9610089-cf61-4c86-bd8d-8d3f6f0d92db · outbound

This paper cites Understanding the effects of.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Understanding the effects of

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:ad6198c4e4378d2b42f733a1fa3145f15ac7de837aa717cf59852a262254cf71

Observation a2f3ba32-a4b8-45d9-991a-fc0364310cbb · outbound

This paper cites Proceedings of the Conference on Language Modeling , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the Conference on Language Modeling , year=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:a657581a46cd8aced8dd793a4ab5681697a2b7e524698195cf208e1311d3fe5b

Observation 4cf23039-e341-4bd6-a438-1e0c2e23ad2e · outbound

This paper cites Findings of the Association for Computational Linguistics , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Findings of the Association for Computational Linguistics , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:476aa6733a13c6f41f8b42b1db785e88f6b005e68f2b3815f8da7386d49913de

Observation 705db951-4606-4573-bf38-ddfe0b6182c0 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:26:26.263139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:086a452fa6142f4ad5a6d26ccf4d2fab36b12dbc66de47fc4c693225980316d5

Observation 70c40a63-4681-4129-b016-7a5f5524e943 · outbound

This paper cites Self-Improvement in Language Models: The Sharpening Mechanism.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:26.274337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:39906f56383100c53364c39d63d380d7de92d82b40f7e103bd89326804f910c6

Observation ea32348d-caf9-4294-ab34-c6e84837953d · outbound

This paper cites an unresolved cited work.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:56dd3b6a5fa827d924caf951f699ed48665f6213d93f5ca13bf87d4d8f7ada43

Observation eae08f34-3010-46fa-ad0c-1cd933c7f1af · outbound

This paper cites The Price of Format: Diversity Collapse in.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning The Price of Format: Diversity Collapse in

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:aaab1350943445406bc48d1e5d68ec91899fc3cc67baf6859b9852bf30e8164c

Observation d05c2bf4-0263-483e-b282-2603445e5f7b · outbound

This paper cites GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.267568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:bbb34d280e1a9d95a9c3db5bd00980ca54e51b6bde4289f4e77796faf55b30f2

Observation 6199598d-34af-48db-b9cc-5f268c67d7da · outbound

This paper cites Echo chamber:.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Echo chamber:

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:4ac5d45c80901dbfaf6d3b1eefa4c44a7ab07beaa0cb33c12b8c21ab5933e278

Observation fea41adb-636c-42c9-bc97-8d1d2d431f5e · outbound

This paper cites Evaluating the diversity and quality of.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Evaluating the diversity and quality of

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:2ffceb57ef193f59123c6e5f3a70d52f2b6609e635d417f23b32366bc92ab7c2

Observation f596246d-22af-416f-895e-dd629f01857b · outbound

This paper cites Proceedings of the Conference on Language Modeling , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the Conference on Language Modeling , year=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:667883a18e9d2c82e8a3ba7bdebcf931f6379161453559887a997bffd684c760

Observation cfbde808-e086-48c3-a262-22ea20a23115 · outbound

This paper cites Outcome-based Exploration for LLM Reasoning.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Outcome-based Exploration for LLM Reasoning

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:26.262089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:539da7bf0fe6799ac7e5a7599c4f0c1023fc3c3820c49f9891e064419694df18

Observation 667d9f4a-7019-4c70-9708-3f3afa185015 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Scaling Laws for Reward Model Overoptimization

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:26:26.254040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:d50ca5e85127c735681e6b1c5f253978386b55e378bdc5d6f6271b0d7d7e51f3

Observation b99aa16b-fa0b-4d8f-b728-17abc074ddde · outbound

This paper cites Confronting reward model overoptimization with constrained.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Confronting reward model overoptimization with constrained

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:1b4d7189ea6e138660776155238c67e85174458b4bf24412d13049a80356b42e

Observation c020c849-0d17-4732-bfa9-f0fef9ee6f5a · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Advances in Neural Information Processing Systems , year=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:ed7ac6fa97a2aff025cf992479cad8e0e24bdf7d42dc71be457479184e2f2e93

Observation 26cdedd0-ebde-4041-a6aa-bd650efed26f · outbound

This paper cites Reinforcement Learning from Human Feedback , year =.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Reinforcement Learning from Human Feedback , year =

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:66a8fa2ab217b65d4e734e17ae69a9b5507d5013cded098c194f62c357a98d37

Observation c0f7609d-a8bc-41dd-95d2-f5f4a645a7a8 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Advances in Neural Information Processing Systems , year=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:3bab00f90d12e50c79a2df022889094ceb81b02d227f86ce740abf439dddca03

Observation 72b3dd61-e719-4e44-a141-63f72523ef95 · outbound

This paper cites Connection Science , volume=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Connection Science , volume=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:e287023a2fe971a25e8117aa88a4e80429c67f6e38fc01d5165947325ec5f4e5

Observation 418ffb64-0b84-4508-b25e-3dd79e5f5eaf · outbound

This paper cites an unresolved cited work.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:0a1ca28c30a9b344eaa8bcc5da6ad26526cae06cd718832948930322af27a980

Observation eff136d7-53a4-4553-bdae-1e23a966afa6 · outbound

This paper cites Proceedings of the International Conference on Learning Representations , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Learning Representations , year=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:42aaccbe1cd9bac2f16f76854f407a5abb82b81203be625806c95e2e7e172881

Observation 8d9a7028-8681-4ad7-8ff2-d04c7770397e · outbound

This paper cites Proceedings of the International Conference on Machine Learning , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Machine Learning , year=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:cf255794503c7e69fbfee682594cde9cd46cc254c8d49da49e874431beb4e569

Observation 2810ed06-b5de-4fca-aff9-45e36ddfa698 · outbound

This paper cites Proceedings of the International Conference on Autonomous Agents and Multiagent Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Autonomous Agents and Multiagent Systems , year=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:18be13ba12050482bf77d6a2ace67f49d9641195f657c34172173e3b4a584c46

Observation 9c15bfe4-f859-45d7-8123-8d6053a75c8e · outbound

This paper cites Econometrica: Journal of the Econometric Society , pages=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Econometrica: Journal of the Econometric Society , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:37cadd920b02691b46f4cc4e9742f398251aa1872c9b887a7cabcee1d4b062d6

Observation d528fea6-74a0-4412-86b2-f21ad98c71a3 · outbound

This paper cites ASTIN Bulletin: The Journal of the IAA , volume=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning ASTIN Bulletin: The Journal of the IAA , volume=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:52cd8569359e149edbc938f7bd554cf4ebe81872a8a5f0d736a63dd0b7e59fc8

Observation 4fc18bd8-eed4-4d15-a8ba-6b586d9b9c0a · outbound

This paper cites Proceedings of the International Conference on Machine Learning , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Proceedings of the International Conference on Machine Learning , year=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:dc1b7b6e90b43c838ade110fc29609f850f0daac180c36f046125b4ca9b0134d

Observation 9724549f-f14c-4e93-ae09-517b4296fa44 · outbound

This paper cites Journal of Risk and Uncertainty , volume=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Journal of Risk and Uncertainty , volume=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:02baed0407cfacf7cf37f3112616f3fc7fa0969f2a20c1ef86d22157c9509fdd

Observation 802f5a39-6813-4310-9354-ff1eaa4c83c4 · outbound

This paper cites Algorithms for.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Algorithms for

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:883468087b495064916aa1cf4332ceade2c46a2e342d04c209ca5dcf2376ba13

Observation a2b11b09-263f-459d-8442-ddc6818f722a · outbound

This paper cites 2023 , publisher=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning 2023 , publisher=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:77089f521fb26c079b3e259593aa80abb3401bf64b52d7a45e3cb9556639397a

Observation 5d121a0b-282e-4202-8ea4-1789258df642 · outbound

This paper cites Vector Policy Optimization: Training for Diversity Improves Test-Time Search.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Vector Policy Optimization: Training for Diversity Improves Test-Time Search

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:26:26.268801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:ba148c836600375ed67c305b516c0c12ef43209031b4778548347398a6498922

Observation 2d05cf72-4630-4938-8084-8db171440a36 · outbound

This paper cites Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:26:26.271488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:3574d060da9841985e74039c69b1548e5a072241a9508bdcaed4ca1dd17928f1

Observation bc6feaca-5886-4535-b64d-a54fb8296fdd · outbound

This paper cites 1999 , publisher=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning 1999 , publisher=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:504f5c18aae541ea2005769d09a0223fa8012f32c1d5fc81e78ff4a77790e309

Observation 57127b27-6417-47b6-9f2d-50e374c11071 · outbound

This paper cites 2005 , publisher=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning 2005 , publisher=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:72d1aa345742c36cc23379496465a6ab705ab32bfa1fb649b1fcdb154a02e1e8

Observation ad36887f-4794-485e-82b8-7d84536772be · outbound

This paper cites On the relationship of the.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning On the relationship of the

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:488310121758c5eaa246b4aafa6f1737c687ffa2dba8f63f819ebfe8f484ae1a

Observation 785ee7b0-1d97-4914-9cc6-d8a87c38b02e · outbound

This paper cites An interactive weighted.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning An interactive weighted

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:f043516dee6693c053aa6a4fc870f5f81caf0e6a4a3719a4c1f329a0a461676f

Observation af0544fb-9c70-4c10-b64e-c73f02244130 · outbound

This paper cites an unresolved cited work.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:48da1dedf45fffeba8b329ea5f904627f68dfd601e9c6d747dfcc189c5e8e373

Observation 42378077-b10f-4161-bc3a-660a8cf29e06 · outbound

This paper cites Operations Research , volume=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Operations Research , volume=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:6d285e03b3ed29f6f4a00371d7c2073ab91403a74225248fa0f74f9cd1e74144

Observation 26849ec4-ff07-4541-809c-b200b5ebcc8b · outbound

This paper cites A closer look at drawbacks of minimizing weighted sums of objectives for.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning A closer look at drawbacks of minimizing weighted sums of objectives for

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:24e9ad201b2d89bc7782ccf7763d0bbd1df94f006b639243633dcd7529795055

Observation ac7670c0-4525-4cc1-83ae-cb29ae6eae9a · outbound

This paper cites an unresolved cited work.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:4457ba34f1e263e20a635ad4ff9ae31504724421cf1d8bbedadce569bdeb98e5

Observation 1ed99b1c-8414-4362-ac42-858bdf6bee93 · outbound

This paper cites an unresolved cited work.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:b62f0e8ce1a55edd1522da9fd4083e661bc53cb74aee4e76d649d60166b59c1c

Observation bdf27403-d3c4-4a86-a1f9-011577fed60a · outbound

This paper cites SIAM Journal on Mathematics of Data Science , volume=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning SIAM Journal on Mathematics of Data Science , volume=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:1a9ad665b7d212ae692f7d4d9ff2e6b63e92e65e93ef81e365b5528ef64f8c4b

Observation a38ed160-5879-4eda-ab21-471431aaceff · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Advances in Neural Information Processing Systems , year=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:ab253cc75d0559953a4d705b11ae54f84610414b52ef4ce4421927e6c898352c

Observation a30a4ca8-69d1-40fa-a405-9bf7058e7367 · outbound

This paper cites IEEE International Symposium on Information Theory (ISIT) , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning IEEE International Symposium on Information Theory (ISIT) , year=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:8d386c1a310275e47733203b53c2c04341deac660c911b1a1fdb26a52c482d92

Observation ab7d5905-79dd-4331-9aee-44d33c9043c2 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Advances in Neural Information Processing Systems , year=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-28T10:59:44.092482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:dbdeb45cef0c959dddfbb2c7f968156e3625d0eda9aac826ac582cd62005fb20

Pith citing papers

Observation f00ee3cd-ecd5-47ae-b844-71a294de6562 · inbound

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation cites this paper.

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:12.498785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:12.498785Z digest=sha256:33156002f2d25dabbd45ba751c4e74c48b8a655cbbd190c05b752a724118b776