Pith. sign in

Paper Citation Record · LEDGER

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2506.12529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12529 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:53:13.730009Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a838ea7f-1367-4cab-ade4-ff13cad9272b · outbound

This paper cites Sutton and Andrew G.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Sutton and Andrew G

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.636568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.574686Z digest=sha256:05523bc93ae9758f2b54fa10ba4955e2b5f7966c2774373e8a294b0679623382

Observation e41adb7b-fc05-47b2-8087-650acdee650b · outbound

This paper cites Reinforcement learning can be more efficient with multiple rewards.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Reinforcement learning can be more efficient with multiple rewards

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.629294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.577558Z digest=sha256:066e1fbc08e74c477ddebdec7563ed0bfcd931264dfc42a34dd6c175d036ade4

Observation 31528ee8-6aa9-4fcc-830b-25e50396a312 · outbound

This paper cites Learning Agile Robotic Locomotion Skills by Imitating Animals.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning Agile Robotic Locomotion Skills by Imitating Animals

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.580172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.580172Z digest=sha256:02c3b8bd1abb41c14f70aae4ef0f288567f3ca09ab66e3ef71ec742f73948516

Observation 668de56f-1921-46d4-b2ec-7ffbe1be87e2 · outbound

This paper cites Deep object-centric represen- tations for generalizable robot learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Deep object-centric represen- tations for generalizable robot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.583522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.583522Z digest=sha256:43de0facec4a041dbe4d37f17d8af3eef41a1321c00f12ca17f42e4785819e2f

Observation b89ad36b-5273-462d-9fa7-3068b8804939 · outbound

This paper cites The ingredients of real world robotic reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning The ingredients of real world robotic reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.622631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.586670Z digest=sha256:7e2c4f74a109dcb6984a954f2fac34e3953db8ee528e76744c7c8a21325e5f40

Observation 1b95cad8-1d9e-4eea-b430-9a45b1b06e2e · outbound

This paper cites Preference Transformer: Modeling Human Preferences using Transformers for RL.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.589096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.589096Z digest=sha256:0eb51cb9fa925b3fa7b04a017d12574d466bc921c290ba9b23495a2dc018e9e7

Observation 1211a9fc-d637-440a-8732-934998ae1c9a · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Fine-Tuning Language Models from Human Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.591993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.591993Z digest=sha256:99f48cbd2100284a3209ceb75014c82c780b14a205353800698abd336edf4dcf

Observation bcdaafa4-37b0-4089-8eca-e979692aa58c · outbound

This paper cites GPT-4 Technical Report.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.594860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.594860Z digest=sha256:ddf16d9a5196aa8755a8eb778cc95786c5dde9da1a363d338b729cc03e8fe418

Observation 69641a7f-1abc-4504-9b55-3c54a6ceec2f · outbound

This paper cites Training language models to follow instructions with human feedback.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Training language models to follow instructions with human feedback

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.615042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.597295Z digest=sha256:5a4e3e520649b7e712611f5df8bb0544a52533f56fedfe4ceb30070fb0a72842

Observation 19b879c1-f247-4979-b6df-efeccc565ea6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.599621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.599621Z digest=sha256:28cbb5bd7204d44d465e7404ec182a883069a0668e8c04b9c62fde1ead0efa05

Observation a2f8fbe3-ce5d-4f95-a555-30270e9eab9b · outbound

This paper cites Active Preference-Based Learning of Reward Functions.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Active Preference-Based Learning of Reward Functions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.602307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.602307Z digest=sha256:605e5b07ed3cfb2369590bfb5713597a41c0150f2917808795f7db811bea6741

Observation ca135057-2688-40fe-a535-3bc88ad74e26 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Deep Reinforcement Learning from Human Preferences

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.608347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.604747Z digest=sha256:0b1f32e6b424cada95a08652180df76e61364594de604af01a0664c3b2735735

Observation f60f215e-05b2-4d97-b993-3dbe38972c4d · outbound

This paper cites A Survey of Preference-Based Reinforcement Learning Methods.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A Survey of Preference-Based Reinforcement Learning Methods

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.601692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.607066Z digest=sha256:4990083ed48b34356c61d383dca3e1a7fee37d0c9d7ec2239d8fe7ee4160cffa

Observation dcc01f40-2cf0-4b44-9e41-5b13da365c11 · outbound

This paper cites an unresolved cited work.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.609596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.609596Z digest=sha256:3e619e3378e6db59d9558beae0d17f99881c335295e969fe16078f244fc4165b

Observation bbed6be8-6b7e-4898-b8a2-4b2515e6542b · outbound

This paper cites Smith, and Pieter Abbeel.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Smith, and Pieter Abbeel

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.594628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.611610Z digest=sha256:66a8fbc1c9d000d6c8ac3107ed3f454f53e7e2e543e10ff0904623ed61ebd9cd

Observation 5a1f93fd-7264-4d23-8023-83386e640017 · outbound

This paper cites Few-shot preference learning for human-in-the-loop RL.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Few-shot preference learning for human-in-the-loop RL

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.579789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.616579Z digest=sha256:3b668ed144e7c2d9b31b433d2fe991230cb819adf0b641eca231c3c685129e3e

Observation 886092de-c884-4349-aee7-dec74b9d0023 · outbound

This paper cites Bradley Knox, and Dorsa Sadigh.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Bradley Knox, and Dorsa Sadigh

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.572904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.618745Z digest=sha256:6037e89f4b720f0db8d117c609863382589b31fb32cde0168562cec6c6c247d7

Observation aa775772-5afe-4056-bb72-43b3df819895 · outbound

This paper cites Inverse Preference Learning: Preference-based RL without a Reward Function.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Inverse Preference Learning: Preference-based RL without a Reward Function

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.559555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.623886Z digest=sha256:bcda148a6d9dc8c5cbfac93bef31d0773b5acf7ef2bed9ccefd0cfb661e7bf5c

Observation 327d9590-55cc-4c53-9ddb-7df41ee30b9b · outbound

This paper cites Direct Preference-based Policy Optimization without Reward Modeling.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Direct Preference-based Policy Optimization without Reward Modeling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.552744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.626123Z digest=sha256:e520982e69c1f88643120c507ac3858a5a1ffb6e78d11ccf1a0213d1ddb00e10

Observation 364218c0-d7c7-44f6-a6e0-26c1e2dd405c · outbound

This paper cites Beyond Reward: Offline Preference-guided Policy Optimization.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Beyond Reward: Offline Preference-guided Policy Optimization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.545970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.628391Z digest=sha256:db136802fc65b8f0106284285c9635e54004478216dfc22d3b98d90cd36b5738

Observation cba5540a-8275-4414-afb8-118e0aa90b58 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.538656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.631200Z digest=sha256:7fbfc136001615cf639cb5a826e96f62a1ab85ffc87c3bb26b43d620f49f0410

Observation 1cc52139-f6aa-46bc-8f50-89faafdb063a · outbound

This paper cites Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.531895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.633432Z digest=sha256:2b9923dae8de5fa99a3b3e567e565a8f7b2a2dc9b38eeb1b3d4247ff27de2089

Observation fb61f357-0b37-4a7c-bee7-13dff66eaae2 · outbound

This paper cites Rethinking reward modeling in preference-based large language model alignment.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Rethinking reward modeling in preference-based large language model alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.524528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.635700Z digest=sha256:acb16c3d254a00a39a3569b8edd41123678302b00a5f485f2d5f07eb663eb0f1

Observation 4533caba-3b52-4501-8e24-c7b60990f44a · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Generalized preference optimization: A unified approach to offline alignment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.517863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.638150Z digest=sha256:3973399e8c8f8bbe0c4afbc15c316f14f761a9d935218f3b6690045a69345895

Observation 8bad5c92-683e-4ad4-a002-cc55117e83cb · outbound

This paper cites Nash learning from human feedback.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Nash learning from human feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.511138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.640321Z digest=sha256:83ddd9ef74c088aefc17bd48170d7d7d0d6ab290ac485dafbaf8a738d8cc300d

Observation 13bdd45a-96b2-43aa-865f-65b560504c57 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A general theoretical paradigm to understand learning from human preferences

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.504228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.643844Z digest=sha256:9b4f547a8e84527d9f7cc8b9bdde833d02268dae50331686b6eae40acc7c50b6

Observation 70cd97ea-e593-49fa-9ef1-b44d30a25af8 · outbound

This paper cites Online Iterative Reinforcement Learning from Human Feedback with General Preference Model.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.497602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.646342Z digest=sha256:2ccdf9f7357e56a81029ae7c498a4883ff1c81a23fdc0f1ebe9d55d91870855a

Observation 681f80ec-a3bf-4e3e-96a8-0c1c4f32eda0 · outbound

This paper cites Intransitivity of preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Intransitivity of preferences

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.648723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.648723Z digest=sha256:e9ef75400dab47d56ee8b336aa7ccd6c45f6436353036de2e6fe44af8c6536d2

Observation a5f4cb82-180f-41ad-bc80-26a2ee1c3fe5 · outbound

This paper cites an unresolved cited work.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:53:14.491103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.651354Z digest=sha256:dfe1a3f335b81cadd53bd0177a011aa75f480967dd369b67462f7a8229ef6061

Observation fa32ebb1-b6de-484f-865b-45bde68c1df8 · outbound

This paper cites RIME: robust preference-based reinforcement learning with noisy preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning RIME: robust preference-based reinforcement learning with noisy preferences

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.484703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.653713Z digest=sha256:e22fb2d7ad65fe349164025df54e97caf747c5a701dbf8d9c220b92ac5c142c8

Observation e1b4705a-362a-4749-a5dc-5b177b88445a · outbound

This paper cites AlpacaFarm: A Sim- ulation Framework for Methods that Learn from Human Feedback.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning AlpacaFarm: A Sim- ulation Framework for Methods that Learn from Human Feedback

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.477470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.656095Z digest=sha256:08b4c3efaf2d5ba6e000072105129bc335c7760ae7951656e7a82e4df1ddd857

Observation 08f7bed2-0bca-414a-bf50-9ebb75e284de · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Reward model ensembles help mitigate overoptimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.470636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.658610Z digest=sha256:60f1a68d70ab7f565c9d4433581c6e1e725f2a25c8e34a2a7122df5341601a7d

Observation 49ca57c0-ffb5-4bfc-bdde-debf0a6536f2 · outbound

This paper cites B-pref: Benchmarking preference- based reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning B-pref: Benchmarking preference- based reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.462842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.662379Z digest=sha256:b227e8ec2d78bd2b43584f870a43c912db642b6f02250d5972b5d5e5848e258b

Observation f35ae298-7b63-4b7c-b111-39ab708b06b1 · outbound

This paper cites Rethinking decision transformer via hierarchical reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Rethinking decision transformer via hierarchical reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.456224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.665757Z digest=sha256:87a58fb1f92c9ec6b4c8122a18056be6c8822df2f5409e156eaa28f225315d9b

Observation 93960999-efcb-4eb0-8a44-5b54d0d724f8 · outbound

This paper cites WHEN SHOULD WE PREFER DECISION TRANSFORMERS FOR OFFLINE REINFORCE- MENT LEARNING? 2024.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning WHEN SHOULD WE PREFER DECISION TRANSFORMERS FOR OFFLINE REINFORCE- MENT LEARNING? 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.449568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.668108Z digest=sha256:d4e8cf6a432f3ef3ddba674b0d33098ed1d178cedb96700606ad61aa05abfbc6

Observation f7b64e6a-3356-4f98-b0e9-4dda53306843 · outbound

This paper cites OFFLINE REINFORCEMENT LEARNING WITH IMPLICIT Q-LEARNING.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning OFFLINE REINFORCEMENT LEARNING WITH IMPLICIT Q-LEARNING

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.442807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.670493Z digest=sha256:74469589f283da05fe1b02028f4fdb5cc7413387a5ac7e4a0dde2ed92fb2f99c

Observation fd154f9e-c94f-4d55-adf5-b1d0e5a085e7 · outbound

This paper cites Preference Alignment with Flow Matching.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Preference Alignment with Flow Matching

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.435365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.672687Z digest=sha256:e58e05993a21b3496bdf2b511e95ea3c8fcc6e52da671f7f637b23ee0779b671

Observation 275949b1-55c3-4527-a651-7e1c23fb4b66 · outbound

This paper cites Flow to better: Offline preference-based reinforcement learning via preferred trajectory generation.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Flow to better: Offline preference-based reinforcement learning via preferred trajectory generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.428600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.675423Z digest=sha256:17ff32587cc0f5560dd3738b1bb0f90f7dbefdaf2198df7315e7255415535697

Observation cd767918-e406-4296-b1ad-44c438625715 · outbound

This paper cites Denoising Implicit Feedback for Recommendation.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Denoising Implicit Feedback for Recommendation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.677623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.677623Z digest=sha256:80ea1287982131a12381930e6dfc449ab18c247d467aefa8097874538bba2a08

Observation 356483ba-d9a8-4434-8c00-6fd4d6b272cb · outbound

This paper cites mixup: BEYOND EMPIRICAL RISK MINIMIZATION.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning mixup: BEYOND EMPIRICAL RISK MINIMIZATION

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.420313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.679660Z digest=sha256:4c6a9514ace99dcef06fb7c5503e94a6d6cffc836fb502747d13fd9c795ba46f

Observation be9b2e01-c247-47f0-8baa-332bbc521971 · outbound

This paper cites Co-teaching: Robust training of deep neural networks with extremely noisy labels.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Co-teaching: Robust training of deep neural networks with extremely noisy labels

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.413068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.681877Z digest=sha256:dbb2da52eebee59e363a48015fc8d33016b97f5d983d149e153a5866a2ce9cad

Observation 29d4218a-fc9c-49e8-aa23-01dd9802a5bf · outbound

This paper cites Does label smoothing mitigate label noise? In Proceedings of the 37th International Conference on Machine Learning, volume 119 of ICML’20, pages 6448–6458.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Does label smoothing mitigate label noise? In Proceedings of the 37th International Conference on Machine Learning, volume 119 of ICML’20, pages 6448–6458

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.405932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.684352Z digest=sha256:33b4d6a1aa84fe0c90170148fe43f0e7de1127e174f2c39bda1750ce6c606355

Observation 5bb06fb2-59c1-4f04-81d9-f9e6de1d2d0b · outbound

This paper cites Learning from noisy labels with deep neural networks: A survey.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning from noisy labels with deep neural networks: A survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.686484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.686484Z digest=sha256:abc4a494cf76c515f82ccf1fa0eefc6df6a2ca1a03521e5a54770d6425487371

Observation 340f57f2-d046-4e68-93fb-8f4a829786b0 · outbound

This paper cites Attention is All you Need.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Attention is All you Need

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.399011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.688693Z digest=sha256:98b22d21582f1be57dbf6f1fd6f2cbb449886e8755fcdb30c5da6fa7cc9eb972

Observation 07b25ea8-2a8e-4604-810c-a4532a8e2b03 · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A Simple Framework for Contrastive Learning of Visual Representations

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.392260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.691043Z digest=sha256:6cc3b5ce103cfac86e313e95fad36a4bc762622fb7c2c29ff53fc6f44ef933c5

Observation d4dcb6c3-5bc3-4569-a471-ca88d75bbcca · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.693302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.693302Z digest=sha256:95aa60f90c8e73a8ca401333beb88837c8620eed0c999580b9e8a22bbe962584

Observation ebefa7eb-9474-40c3-84b9-61dc1a853f1b · outbound

This paper cites D4RL: Building Better Benchmarks for Offline Reinforcement Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning D4RL: Building Better Benchmarks for Offline Reinforcement Learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.385421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.695735Z digest=sha256:aae09a5d5d42f49c588334712271b4f0958d7231c40b2854f61e2c033e029bfd

Observation 0c06b20e-a5dc-42b0-86c5-cf226630a8c7 · outbound

This paper cites URL https://www.gymlibrary.dev/environments/ mujoco/index.html.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/index.html

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.378249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.698093Z digest=sha256:8bc9804aeb67459b928451d9feeedcdc494a42741493e13940a9312b5b5c5836

Observation 5883a3ed-b7f8-4bac-8410-542612030a70 · outbound

This paper cites Offlinerl-kit: An elegant pytorch offline reinforcement learning library.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Offlinerl-kit: An elegant pytorch offline reinforcement learning library

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.370970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.700406Z digest=sha256:8627393e132e593a64cd87a6a781468344c96e3d2e427b186a4bde44a54f8d70

Observation acc20467-b73a-4e61-aafb-0e89c97011d2 · outbound

This paper cites Hierarchically decoupled imitation for morpho- logical transfer.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Hierarchically decoupled imitation for morpho- logical transfer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.363018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.703454Z digest=sha256:01b96dc7fddfefc87136a2a0f2a58d80262d862fa38df34bd534fcb4e213c4ab

Observation b8b6faf0-90c2-4efb-a873-f36946864a93 · outbound

This paper cites PEARL: Zero-shot cross-task preference alignment and robust reward learning for robotic manipulation.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning PEARL: Zero-shot cross-task preference alignment and robust reward learning for robotic manipulation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.354877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.705831Z digest=sha256:285e57a764fdbb77f4f5c6e7dd88c8a83f15dfd141b1cad564a39586542f40aa

Observation faa778ae-004a-4dde-8ba4-0f0d7a677f57 · outbound

This paper cites Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.708152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.708152Z digest=sha256:508cca99e430a055f5f83897a6e5d1c2b472c4e679119975f7c92b1b0f1e893b

Observation 69bc5927-4172-4956-8c45-9083497c7694 · outbound

This paper cites Learning robust perceptive locomotion for quadrupedal robots in the wild.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning robust perceptive locomotion for quadrupedal robots in the wild

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.710490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.710490Z digest=sha256:32bfc6cf51cd7d42c29be0e8b4e58bd80ba7a745976f603bf5df4394fb237a6f

Observation ab1552f6-f733-4b52-9d09-1b90227dd761 · outbound

This paper cites dm_control: Software and tasks for continuous control.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning dm_control: Software and tasks for continuous control

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.713036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.713036Z digest=sha256:feaaf7c949384f1c38a2d32ec884da6556ea4f0dd126416853e0809baa5fec92

Observation 8e0ccee2-5793-45ba-b7b4-5e6d7602df08 · outbound

This paper cites URLB: Unsupervised reinforcement learning benchmark.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URLB: Unsupervised reinforcement learning benchmark

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.346513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.715496Z digest=sha256:d59febbecdb73893ac1b9145625099687fc9e326083985714b9b23d220baf2a7

Observation 103bccc0-2245-448a-b121-76970064a4ac · outbound

This paper cites Continuous control with deep reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Continuous control with deep reinforcement learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.717937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.717937Z digest=sha256:4ce934fa80a676312ef8336cfb65aae5f7b508afa846e2157ef47b60bb9a60d7

Observation ce58efbb-6332-4e03-81a4-6bd802df36c9 · outbound

This paper cites URL https://docs.pytorch.org/ docs/stable/generated/torch.nn.TransformerEncoder.html.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://docs.pytorch.org/ docs/stable/generated/torch.nn.TransformerEncoder.html

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.338969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.721084Z digest=sha256:4641845188585e55158c0020117727665defd4c359d74ec45fafda03bd839bf0

Observation 9c3c1e64-16a1-4f96-92c9-bfcd397821ca · outbound

This paper cites URL https://www.gymlibrary.dev/environments/ mujoco/hopper/.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/hopper/

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.331172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.724721Z digest=sha256:9bf64ffde0fa096113aa04c8b838dbc01a343254b2133765f9ef465c3c70ac7f

Observation 26b71533-3a9e-45e8-96ae-1e3ad80676d8 · outbound

This paper cites URL https://www.gymlibrary.dev/environments/ mujoco/walker2d/.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/walker2d/

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.323213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.727226Z digest=sha256:a10219408509f020497054f40d03a8990080fae51d29479fe34908423b30c737

Observation ab7ec7ee-36e9-419b-9380-ac83076703d1 · outbound

This paper cites For the Franka Kitchen tasks, we use the preference datasets from An et al.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning For the Franka Kitchen tasks, we use the preference datasets from An et al

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.314920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.730009Z digest=sha256:cb6b0e6d1ba44051925396fba8247d02e09ddd202ac3267086a748ac2fddc55c

Observation 41a941b2-d2f4-404f-8b60-2b89f0b89425 · outbound

This paper cites ISSN: 2640-3498.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning ISSN: 2640-3498

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.588090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.613923Z digest=sha256:3775909f37c31475cf364d51830fde0acafe519d7ba6ffc639b283a5a485c152

Observation 0347b73b-0a39-4a05-b278-31a057847404 · outbound

This paper cites an unresolved cited work.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:53:14.566095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:53:13.621171Z digest=sha256:49e204c59edcc14e97b0d9458365d287e37c39be3eeca22c6ca76d42e7eb0460

Pith citing papers

No inbound Pith citation observations are available.