Pith. sign in

Paper Citation Record · LEDGER

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations

As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2507.08707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08707 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:20:25.838910Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c3e263e-9bb5-4523-80cf-5f2812f2552e · outbound

This paper cites Human-level control through deep reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Human-level control through deep reinforcement learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.710016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.710016Z digest=sha256:cc8017548621de9b13b96971a197961a4f58091eec23255736b3204100079fde

Observation d76da8ca-b402-47e8-a52d-4225da23a541 · outbound

This paper cites an unresolved cited work.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.714530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.714530Z digest=sha256:ff72cae2a488bcb3a7c9437ff919033018ada440a1fe7320cbd4f437d0eac65f

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.717886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.717886Z digest=sha256:6123f2800939691f35c4f7f46398ce75832d1d49eebc2fff5b7446940ebcac40

Observation 232e92b0-299c-48a1-b84e-f9bd9d3527ec · outbound

This paper cites Deep reinforcement learning with double q-learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Deep reinforcement learning with double q-learning,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.721431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.721431Z digest=sha256:53e659d7d79c77d1046a32552eba088ba0b326e23459c5fa9c8d6c55f12003ea

Observation a8029730-e6e3-4ef3-9f85-0dc3f74ce68a · outbound

This paper cites Dueling network architectures for deep reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Dueling network architectures for deep reinforcement learning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.724923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.724923Z digest=sha256:f5535114489371f4888b7523ab8d22f205adc9d950fe2f0db051b877773dd5f2

Observation d4627158-c1ed-486a-a136-c297b7d3b994 · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Rainbow: Combining improvements in deep reinforcement learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.171698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.728465Z digest=sha256:9a5e9df3664223d2108c90969b1038e5be764ebd58e945755495599d9c0e13f8

Observation f7fffa32-e1b7-4a6c-b266-2e628bf3c1fc · outbound

This paper cites Agent57: Outperforming the atari human benchmark,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Agent57: Outperforming the atari human benchmark,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.161510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.731984Z digest=sha256:f86ae613f5c62e2d10a989a00c321c494e1853d78d37a4ce9fb31477797e69cd

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.735117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.735117Z digest=sha256:8658c97ce060b3c05cb09b8ccd3975ac8984289bc24f26c5b5655ff7ca994e76

Observation d660e7e5-ed9c-48b2-9f6a-de95d503967d · outbound

This paper cites Hind- sight experience replay,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Hind- sight experience replay,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.738520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.738520Z digest=sha256:ce9a77cdcf3eea4f831756362ffdc55b2a1f940e069cdabaa588ba57b377a2f4

Observation ca24fd49-4464-4192-9cba-c1e09a061b69 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Policy invariance under reward transformations: Theory and application to reward shaping,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.741654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.741654Z digest=sha256:fd3f97504f255be2eacc5a808c09c61ff83936ed97dad0db69ae08042a79d3fb

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.744740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.744740Z digest=sha256:8ef303970536c969162065b302dbcde7b254b3ea1e600bdfa54f1d0edc3d68a6

Observation fb9405a3-f9f2-4628-b6b6-fc5fd8a4bb8f · outbound

This paper cites Algorithms for inverse reinforcement learning.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Algorithms for inverse reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.139975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.748757Z digest=sha256:99d5cf1e251874161c73f223add3893915bbe82f07cab75b2e45109b4bb0b3de

Observation 99d5a334-74b2-43f4-9d6f-0c0398f2c446 · outbound

This paper cites A survey of inverse reinforcement learning: Challenges, methods and progress,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations A survey of inverse reinforcement learning: Challenges, methods and progress,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.751966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.751966Z digest=sha256:ddc8028fcf4c946dca51aa88ac7dd0581899ecdf5d89d21746e817d90834759e

Observation bf9d4eba-acb4-40ea-91d0-954e62a9e65d · outbound

This paper cites Maximum entropy inverse reinforcement learning.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Maximum entropy inverse reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.124150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.754940Z digest=sha256:fa2f26474c384cfd81e60700ef4fd6468d4efb2cf72ef3715097b2e81df5be13

Observation 1ec7ecbd-e008-40f9-bb35-1315b5ff8b45 · outbound

This paper cites Relative entropy inverse reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Relative entropy inverse reinforcement learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.114120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.757931Z digest=sha256:0a4442fa0316a59c38e435dd710235384be7ea1be8ac9ca9a3d3fb486245660b

Observation f91dc819-13e8-4b2b-b8e7-8bda180ffd72 · outbound

This paper cites Guided cost learning: Deep inverse optimal control via policy optimization,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Guided cost learning: Deep inverse optimal control via policy optimization,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.104633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.760986Z digest=sha256:939a6ef8350fd78a0d7ea079d9d4a40d4030921776bb3267d37a758f9074475a

Observation 7c849bc3-582a-45e8-87a3-a85ab7d30ad1 · outbound

This paper cites Preference-learning based inverse reinforcement learning for dialog control,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Preference-learning based inverse reinforcement learning for dialog control,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.092792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.764010Z digest=sha256:57282b1be1c6e039d6b0a7a029fee055218aa17de34d1879162a222ba8ec6894

Observation d7964256-0650-499d-bdeb-f811b26c75f1 · outbound

This paper cites Model-free preference- based reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Model-free preference- based reinforcement learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.081885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.766940Z digest=sha256:cb43e5f21ef389897ebe169533421da5b0cda11ebbcb56b5a86b4e47569e20f9

Observation cd855081-3c73-4431-8099-6ef536c9c79d · outbound

This paper cites Reward learning from human preferences and demonstrations in atari,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Reward learning from human preferences and demonstrations in atari,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.770793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.770793Z digest=sha256:cb47afdaa2515e2e3ab984b26e7050e00316d98a060e5112fde07bafa964029d

Observation bee54f88-2d1a-4a0e-8b97-7260168fc507 · outbound

This paper cites Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.066841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.774286Z digest=sha256:21ad6d285f99eb1696eb9bcb84d2fb863ccd35032dd9b9f5db14a5ef0720fc33

Observation 8ee9bfe2-6d15-484f-bfd3-47cde9b0ec2c · outbound

This paper cites Learning Reward Functions by Integrating Human Demonstrations and Preferences.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Learning Reward Functions by Integrating Human Demonstrations and Preferences

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.778489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.778489Z digest=sha256:15bf815a3ffa8dc20f46ba9ca51c07b97d7b69317c6421abad166454736ee08c

Observation f4e3a1e0-6a47-45dd-93b3-b46eed409d7d · outbound

This paper cites Better-than-demonstrator imitation learning via automatically-ranked demonstrations,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Better-than-demonstrator imitation learning via automatically-ranked demonstrations,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.056843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.782420Z digest=sha256:1d6db119c67c4fccb632ab11cfb1e027a906e39672f7f41a4cd9cdbab8be1dc5

Observation 326b094d-4a6c-4a6c-8014-ead0214a095c · outbound

This paper cites Learning from suboptimal demonstration via self-supervised reward regression,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Learning from suboptimal demonstration via self-supervised reward regression,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.046640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.786139Z digest=sha256:56f71233d059db88ba3c19cdfe8b381ecb254cd44c51ea4e1d773b32e58868ea

Observation d43b979c-628b-4b0f-9e37-ea1c51da1bf6 · outbound

This paper cites Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.789375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.789375Z digest=sha256:5e1e0fa851abddca692d3ebb42b39eb73b44c3e11fea6e6cb7b06597776db48e

Observation ed52cbac-4d3b-4e44-81ce-a7deef39d4f6 · outbound

This paper cites Pyquaticus capture the flag gymnasium,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Pyquaticus capture the flag gymnasium,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.029779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.792772Z digest=sha256:5b16af972b593ac2782813e81397b4984ec33c25d580fae3b55daadce06c2ccf

Observation e0446422-77f7-45bf-a6be-caaccd8f4c4e · outbound

This paper cites Nested autonomy for unmanned marine vehicles with moos-ivp,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Nested autonomy for unmanned marine vehicles with moos-ivp,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.019423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.796031Z digest=sha256:08b442803c064618a386cd68cd7c60f4e76f09e0569d159073d154efabcc6615

Observation 7429e3e1-cd43-4450-bc6e-024f70d379bc · outbound

This paper cites Efficient training of artificial neural networks for autonomous navigation,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Efficient training of artificial neural networks for autonomous navigation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.008756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.799195Z digest=sha256:227857789628f3579520b7740ff8197d9eb8cfa742e65100aa400d6c98c742cb

Observation a11709ba-bd7c-4f6a-b7a1-b50978201d51 · outbound

This paper cites Behavioral cloning from observation,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Behavioral cloning from observation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.997906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.802431Z digest=sha256:d010e447a98752a8f99d6e95acd01e480cee59984fe1ea649a078c9a68511244

Observation a55694fd-142c-4519-80b2-6d469ca7ed00 · outbound

This paper cites Generative adversarial nets,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Generative adversarial nets,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.809841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.809841Z digest=sha256:041317bb750763c398f7d6208db47ec0d5859b55d7c65d861f31866155b79b91

Observation cd703e8b-e143-4e15-97ea-ab212b85b628 · outbound

This paper cites Generative adversarial imitation learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Generative adversarial imitation learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.980808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.813120Z digest=sha256:f9217cdc155780985b312380101c1cccaf1a7b17f5f7868ad277892cabbff1a7

Observation ea2561d4-701f-4003-ac5a-b8d56cddd217 · outbound

This paper cites Learning Robust Rewards with Adversarial Inverse Reinforcement Learning.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Learning Robust Rewards with Adversarial Inverse Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.816962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.816962Z digest=sha256:8080cffa959b652f0c98947179dda68cda6a517cafb866b6d0c12eb09a1fe6f7

Observation fcdadd5c-ed15-4e7a-8265-1406a9c589bf · outbound

This paper cites Inverse reinforcement learning for video games.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Inverse reinforcement learning for video games

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:20:25.879792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.820875Z digest=sha256:f4b85f95d10d42f0fdecde5f66b8558d27c04ed607a72567138f5605ed65d6a2

Observation 7ffeea1a-ab22-4eb1-879e-b2a9728b90f5 · outbound

This paper cites A survey of preference-based reinforcement learning methods,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations A survey of preference-based reinforcement learning methods,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.970626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.824428Z digest=sha256:b30cf0456950b88e4a25013210ed1df8cf6cec4141206ab9097d49683208bd2a

Observation a14d0fad-69fa-42c5-89df-5d23f111d458 · outbound

This paper cites Optiongan: Learning joint reward-policy options using generative adversarial inverse reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Optiongan: Learning joint reward-policy options using generative adversarial inverse reinforcement learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.959978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.827400Z digest=sha256:ba9d2a876c33a5c7d7312771493ea910a27ae537dc1c371def9bc4904d3b1afb

Observation 441a28b5-767e-48b6-ab07-f752e206dd26 · outbound

This paper cites Hierarchical relative entropy policy search,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Hierarchical relative entropy policy search,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.949493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.831645Z digest=sha256:8a4c20a4ce5f50cd87f99a723ca1ea2a0d9f0036c18be051642c3a7f703eea0e

Observation 325fe346-047f-40c1-b2bd-063739624956 · outbound

This paper cites Deep reinforcement learning from human preferences,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Deep reinforcement learning from human preferences,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.835801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.835801Z digest=sha256:d05a2e238222561745f80f272daa4cfb98fbb2d32ca1a55187b27cbdfa5025d4

Observation 5e122e81-04f9-4e42-9508-a37889ed5663 · outbound

This paper cites Inverse reinforcement learning from failure,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Inverse reinforcement learning from failure,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.934336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:20:25.838910Z digest=sha256:4b8cfd8a9a55d697f3cf6370ff5d95ad058a0e2b4d24a78a4b34fed331d67de6

Observation 395cd817-782a-463d-85a0-500b086c9701 · outbound

This paper cites Available: https://doi.org/10.24963/ijcai.2018/687.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Available: https://doi.org/10.24963/ijcai.2018/687

Reference 4957

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.805843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.805843Z digest=sha256:08b92e3378e754238918e763283cd2d90f2cb0dd4dc42624b2c437a32f3be7ee

Pith citing papers

No inbound Pith citation observations are available.