Pith. sign in

Paper Citation Record · LEDGER

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2506.13741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13741 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:04.401624Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy25
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d3ca3b8-39bc-4df9-9bda-4faae0bb9f96 · outbound

This paper cites Forecasting with Multiple Seasonality.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Forecasting with Multiple Seasonality

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:33:04.542813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.096688Z digest=sha256:8fffecab8dba80619893cd8366d8397ea4e3985d87aedba7f6200dfad082aa82

Observation 0ff5fe61-f701-4d0e-98cf-f78aa95601dc · outbound

This paper cites Batch active preference-based learning of reward functions.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Batch active preference-based learning of reward functions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.769359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.165524Z digest=sha256:dedb22bb5c74c4912372ff1659d387e16147274ea558eba395d80c543b8b87a7

Observation 3ca188d0-d1aa-4191-bee9-ae299a861ddf · outbound

This paper cites Landolfi, Dylan P.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Landolfi, Dylan P

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.761093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.194615Z digest=sha256:b0cc013d4b4a0a791409dc20f56b9bbb9dd0a11d0fdc2f2879586f7d39c190ad

Observation 16f14e3a-6d11-45d0-a1a3-9dbcc885d1ff · outbound

This paper cites Rank analysis of incomplete block designs: I.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Rank analysis of incomplete block designs: I

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.271354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.271354Z digest=sha256:94c9f3e390a83b4e26443c1be4fff60671840a200d247b243199a44a182a4835

Observation bb8ca152-c1ee-466f-8552-22153d40bb14 · outbound

This paper cites Christiano, Jan Leike, Tom B.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Christiano, Jan Leike, Tom B

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.747117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.342628Z digest=sha256:4c47e7d419b421b826c2256494772925bbb85cc625cf5e7e7a9527b574186837

Observation 75ee63ad-1625-4f68-87ca-0306261e0258 · outbound

This paper cites Autonomous skill discovery with quality-diversity and unsupervised descriptors.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Autonomous skill discovery with quality-diversity and unsupervised descriptors

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.738394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.370760Z digest=sha256:5452515f4890ead6e8cf58b99b76d1de38aa22838477862df5ebe925d8a65827

Observation ee5644ab-4798-4701-ab72-26e93eb93764 · outbound

This paper cites Faster Improvement Rate Population Based Training.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Faster Improvement Rate Population Based Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:33:04.530615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.441449Z digest=sha256:b62ec0c328283077c99d622a7bf52c0a7d9fb8ef246ee480bc31534ad30cd58b

Observation 98854a81-2221-4aef-b8d9-d1f072050231 · outbound

This paper cites Diversity is all you need: Learning skills without a reward function.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Diversity is all you need: Learning skills without a reward function

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.730037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.502883Z digest=sha256:71af59657143a16a830b22032551ebf216c533752e0d43a395a31b20245e2faa

Observation 8b8647fa-d609-449b-9991-3b6ff6253b6e · outbound

This paper cites Jax-lob: A gpu-accelerated limit order book simulator to unlock large scale reinforcement learning for trading.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Jax-lob: A gpu-accelerated limit order book simulator to unlock large scale reinforcement learning for trading

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.721729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.566117Z digest=sha256:f479034651c10ae514b332c14b743dd5ea8d8548aa94eafa4a928cd0301a8439

Observation 7018460c-2e08-436c-9e75-6936fccb3a54 · outbound

This paper cites u rnkranz, Eyke H \.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning u rnkranz, Eyke H \

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.622234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.622234Z digest=sha256:f40a11059c10fd91820ffa355012644513770bfdd439ff4fdfa51eb380eee4ec

Observation 845db6e5-ee25-43ac-98cb-4ebbacc3c51d · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.685817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.685817Z digest=sha256:ed156775cb2a7aacc4004276a6dacc4fb63eb1612e3591678c6b57ef4d657ebc

Observation 20f8516b-548f-49ce-b3b0-246b310d433a · outbound

This paper cites Russell, and Anca D.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Russell, and Anca D

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.707679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.753777Z digest=sha256:4ccbd1b760ef26d4e65d6f2cc0863312ae1fb0d9b0b60ce77b69fb63eb56436a

Observation 2dc743d1-2a3b-46ca-9edf-4d6bdf6d34c9 · outbound

This paper cites Query-policy misalignment in preference-based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Query-policy misalignment in preference-based reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.698293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.789081Z digest=sha256:961ca1e7444d88938acc2c6d2dcd524e4ff1169e249a2231a9de54a58fb91a95

Observation b34a91bc-ede1-4ca6-9fce-25bd053e69b7 · outbound

This paper cites TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.820968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.820968Z digest=sha256:db304ae4fe5ff3efabf15efbfb19cb484844ddc1293c14434f5a8f1e1b71af4a

Observation fc53ef44-60ad-4df3-ad40-b688b6a136cd · outbound

This paper cites Population Based Training of Neural Networks.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Population Based Training of Neural Networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.870174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.870174Z digest=sha256:847724d5aafe7d13ad5bcbcb7b29b3e9807a65c81da18f8014828dd174e2227d

Observation 4a483ab6-d007-40f8-8f9b-c6b3a0bf95af · outbound

This paper cites One solution is not all you need: Few-shot extrapolation via structured maxent rl.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning One solution is not all you need: Few-shot extrapolation via structured maxent rl

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.689812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.922990Z digest=sha256:0b21ddccf811ee75705bfd126c433d4eb53421fc5e7d59fa8c2e5603d91c790d

Observation b0a333ab-5437-4079-b7de-e88058b5fd29 · outbound

This paper cites B-pref: Benchmarking preference-based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning B-pref: Benchmarking preference-based reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.682059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.965011Z digest=sha256:8bdd11c11f8f8bf70ba8672e88a85571fcba44fc459bc354dd14f26ea7f0aadc

Observation dea18995-5bd0-4844-b2e9-15ef7185ed2b · outbound

This paper cites Smith, and Pieter Abbeel.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Smith, and Pieter Abbeel

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.674287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.009880Z digest=sha256:a7b4b566329c606cfde5a8707513bfcf812035233c28da5b11c58f4cd25cab82

Observation 22081c72-68ed-4432-befd-c43cfb4912b9 · outbound

This paper cites Evolution through the search for novelty.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Evolution through the search for novelty

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.666897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.076339Z digest=sha256:0c5ccfb4935f45bd42a06a0d473f1800ca32e55308ebf9533031a3668a7aed7b

Observation 7394ec4d-2201-40c3-adc0-e427cf205823 · outbound

This paper cites Evolving a diversity of virtual creatures through novelty search and local competition.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Evolving a diversity of virtual creatures through novelty search and local competition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.659279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.161033Z digest=sha256:a1280f9f441a8149b3cd3a0cfbc2563600d91a5472f1220227895e11ba1d1014

Observation 8f765897-1cd7-4aa4-b3e0-3abdf2ead841 · outbound

This paper cites Reward uncertainty for exploration in preference-based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Reward uncertainty for exploration in preference-based reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.651000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.215678Z digest=sha256:9904a36af7325619c62c51989e7a7af9610232b9daa48f4c4557d83c9b8c331a

Observation d7e0be92-746a-4f3c-b441-eb77449568ab · outbound

This paper cites Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.643003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.280078Z digest=sha256:7ffed65ef51cdb33c903384a913e225317ca77966ee5ceb7c52db5868ad30cf1

Observation 8b9e4ef6-97bd-43a6-9945-99b25d8274ca · outbound

This paper cites Efficient preference-based reinforcement learning using learned dynamics models.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Efficient preference-based reinforcement learning using learned dynamics models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.634836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.339377Z digest=sha256:9d90b41c69a76990f773190fdf65dfd352f0b3ef094ee518316e0b97e42401d3

Observation 6856a5f5-707d-44d1-863e-9b9c6e4b3f86 · outbound

This paper cites Variquery: Vae segment-based active learning for query selection in preference-based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Variquery: Vae segment-based active learning for query selection in preference-based reinforcement learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.626755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.389931Z digest=sha256:c05d417e10f41a3a32e712a7e4700dd5bb23fe96f1aca83f24fffd12028a9b1d

Observation 5e0af0b1-3115-40a0-80bb-045cb1e8d4b5 · outbound

This paper cites Rewards Encoding Environment Dynamics Improves Preference-based Reinforcement Learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Rewards Encoding Environment Dynamics Improves Preference-based Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:03.458839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:03.458839Z digest=sha256:fd1699c8a88ebb7a381d7b559cdc497ed56218aa047ab19b9d2a5f107763ef64

Observation 10ecdd6d-fdf7-4749-a094-8816e6e2010b · outbound

This paper cites Illuminating search spaces by mapping elites.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Illuminating search spaces by mapping elites

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:03.520942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:03.520942Z digest=sha256:10151958d7eeed4ed42e2b6887f10452589457c023a0513751f92d541fa7dda1

Observation cf25001d-17db-4534-a18f-7a5581b6f219 · outbound

This paper cites Xland-minigrid: Scalable meta-reinforcement learning environments in jax.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Xland-minigrid: Scalable meta-reinforcement learning environments in jax

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.618180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.600291Z digest=sha256:192d6ab0de50ddace0c863eef20a74099a4f47df13006deb470b75cc247260bc

Observation 6f73c65c-9e1d-49c8-be74-979b7d1236bb · outbound

This paper cites Policy gradient assisted MAP-Elites.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Policy gradient assisted MAP-Elites

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.608896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.650089Z digest=sha256:7eb44ce852b48aabd887fe23a5ffd8af73998dbf8dacc2a9e41d3ac610d357c1

Observation a166f64c-c237-412e-972a-f1dbb3340dce · outbound

This paper cites Discovering diverse solutions in deep reinforcement learning by maximizing state--action-based mutual information.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Discovering diverse solutions in deep reinforcement learning by maximizing state--action-based mutual information

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.600764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.732003Z digest=sha256:c12793158f93f84c3e7cbdded790bcbd369c68d2be4ebf7d2e901fbc253093b0

Observation c0ed7339-82d8-44be-a7b7-a506a191fd40 · outbound

This paper cites SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:03.800435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:03.800435Z digest=sha256:a5c0354bd103cc1d31e81c7cdb1fbef9086b0166642b048f1761e7113d507f05

Observation 47273825-0496-4a5f-87f8-5a271a79c99c · outbound

This paper cites Effective diversity in population based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Effective diversity in population based reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.592571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.883178Z digest=sha256:0da884785df71ea8a1e4004c6d375cd10857748b72600a2a480fd9e8662b410b

Observation 4d3691d9-59db-49ee-868b-4f9f7203e226 · outbound

This paper cites Jaxmarl: Multi-agent rl environments in jax.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Jaxmarl: Multi-agent rl environments in jax

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.583969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.948814Z digest=sha256:8fbce8241db44a00ced110628beb605f9dd3a9e39b6adafbfafd10a30e264842

Observation cba07e3f-6afc-47cc-8173-ad5d6961a15d · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:04.015747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:04.015747Z digest=sha256:a23f6adeadcbca561e8eb78c75652033a2f6b4fc1efcb6684ef602deb73bd7e9

Observation e602c915-ec49-4297-8476-52b38f3e57eb · outbound

This paper cites Dynamics-aware unsupervised discovery of skills.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Dynamics-aware unsupervised discovery of skills

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.575052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:04.098928Z digest=sha256:9460e5e14b2d6bc607d8c3dc418e91260be04d9845718c47a5eef974a25c3e48

Observation 100cb00e-0586-444b-b90d-34a2793461b6 · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Reinforcement learning: An introduction, volume 1

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:04.163238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:04.163238Z digest=sha256:72d361b6955dd968d5455c307578ca9804abea5e93777c4ce0aeb5b142cbab77

Observation 8ea88ba8-32ca-403b-9456-f0b60699cb2d · outbound

This paper cites DeepMind Control Suite.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning DeepMind Control Suite

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:04.252878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:04.252878Z digest=sha256:615d3482c2c96a586a91348efcdd689e91a89e54d11b92dfe0954d18ce747080

Observation c6be78c0-e1dc-45bd-92cf-e897ea2f2946 · outbound

This paper cites Ball, Vu Nguyen, Binxin Ru, and Michael A.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Ball, Vu Nguyen, Binxin Ru, and Michael A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.560627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:04.332746Z digest=sha256:1caaa75323503c969c0a043157a2959930dad137328c17bbbb839c7fa19975cf

Observation 1591528d-92f0-4345-a103-0f26b8bbe050 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning A survey of preference-based reinforcement learning methods

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.552166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:04.401624Z digest=sha256:bb30f28b04f3b6fdaa79daaf60e8e973eba2d37a8387d849bfe0de9133bce3b8

Pith citing papers

No inbound Pith citation observations are available.