Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:53:13.730009Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2506.12529.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:53:13.730009Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a838ea7f-1367-4cab-ade4-ff13cad9272b · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Sutton and Andrew G
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e41adb7b-fc05-47b2-8087-650acdee650b · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Reinforcement learning can be more efficient with multiple rewards
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 31528ee8-6aa9-4fcc-830b-25e50396a312 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning Agile Robotic Locomotion Skills by Imitating Animals
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 668de56f-1921-46d4-b2ec-7ffbe1be87e2 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Deep object-centric represen- tations for generalizable robot learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b89ad36b-5273-462d-9fa7-3068b8804939 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning The ingredients of real world robotic reinforcement learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1b95cad8-1d9e-4eea-b430-9a45b1b06e2e · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1211a9fc-d637-440a-8732-934998ae1c9a · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Fine-Tuning Language Models from Human Preferences
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcdaafa4-37b0-4089-8eca-e979692aa58c · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning GPT-4 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69641a7f-1abc-4504-9b55-3c54a6ceec2f · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Training language models to follow instructions with human feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 19b879c1-f247-4979-b6df-efeccc565ea6 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2f8fbe3-ce5d-4f95-a555-30270e9eab9b · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Active Preference-Based Learning of Reward Functions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca135057-2688-40fe-a535-3bc88ad74e26 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Deep Reinforcement Learning from Human Preferences
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f60f215e-05b2-4d97-b993-3dbe38972c4d · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A Survey of Preference-Based Reinforcement Learning Methods
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dcc01f40-2cf0-4b44-9e41-5b13da365c11 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbed6be8-6b7e-4898-b8a2-4b2515e6542b · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Smith, and Pieter Abbeel
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a1f93fd-7264-4d23-8023-83386e640017 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Few-shot preference learning for human-in-the-loop RL
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 886092de-c884-4349-aee7-dec74b9d0023 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Bradley Knox, and Dorsa Sadigh
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aa775772-5afe-4056-bb72-43b3df819895 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Inverse Preference Learning: Preference-based RL without a Reward Function
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 327d9590-55cc-4c53-9ddb-7df41ee30b9b · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Direct Preference-based Policy Optimization without Reward Modeling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 364218c0-d7c7-44f6-a6e0-26c1e2dd405c · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Beyond Reward: Offline Preference-guided Policy Optimization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cba5540a-8275-4414-afb8-118e0aa90b58 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1cc52139-f6aa-46bc-8f50-89faafdb063a · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb61f357-0b37-4a7c-bee7-13dff66eaae2 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Rethinking reward modeling in preference-based large language model alignment
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4533caba-3b52-4501-8e24-c7b60990f44a · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Generalized preference optimization: A unified approach to offline alignment
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8bad5c92-683e-4ad4-a002-cc55117e83cb · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Nash learning from human feedback
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 13bdd45a-96b2-43aa-865f-65b560504c57 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A general theoretical paradigm to understand learning from human preferences
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 70cd97ea-e593-49fa-9ef1-b44d30a25af8 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Online Iterative Reinforcement Learning from Human Feedback with General Preference Model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 681f80ec-a3bf-4e3e-96a8-0c1c4f32eda0 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Intransitivity of preferences
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5f4cb82-180f-41ad-bc80-26a2ee1c3fe5 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa32ebb1-b6de-484f-865b-45bde68c1df8 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning RIME: robust preference-based reinforcement learning with noisy preferences
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e1b4705a-362a-4749-a5dc-5b177b88445a · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning AlpacaFarm: A Sim- ulation Framework for Methods that Learn from Human Feedback
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 08f7bed2-0bca-414a-bf50-9ebb75e284de · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Reward model ensembles help mitigate overoptimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 49ca57c0-ffb5-4bfc-bdde-debf0a6536f2 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning B-pref: Benchmarking preference- based reinforcement learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f35ae298-7b63-4b7c-b111-39ab708b06b1 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Rethinking decision transformer via hierarchical reinforcement learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93960999-efcb-4eb0-8a44-5b54d0d724f8 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning WHEN SHOULD WE PREFER DECISION TRANSFORMERS FOR OFFLINE REINFORCE- MENT LEARNING? 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f7b64e6a-3356-4f98-b0e9-4dda53306843 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning OFFLINE REINFORCEMENT LEARNING WITH IMPLICIT Q-LEARNING
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fd154f9e-c94f-4d55-adf5-b1d0e5a085e7 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Preference Alignment with Flow Matching
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 275949b1-55c3-4527-a651-7e1c23fb4b66 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Flow to better: Offline preference-based reinforcement learning via preferred trajectory generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd767918-e406-4296-b1ad-44c438625715 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Denoising Implicit Feedback for Recommendation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 356483ba-d9a8-4434-8c00-6fd4d6b272cb · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning mixup: BEYOND EMPIRICAL RISK MINIMIZATION
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation be9b2e01-c247-47f0-8baa-332bbc521971 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Co-teaching: Robust training of deep neural networks with extremely noisy labels
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 29d4218a-fc9c-49e8-aa23-01dd9802a5bf · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Does label smoothing mitigate label noise? In Proceedings of the 37th International Conference on Machine Learning, volume 119 of ICML’20, pages 6448–6458
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5bb06fb2-59c1-4f04-81d9-f9e6de1d2d0b · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning from noisy labels with deep neural networks: A survey
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 340f57f2-d046-4e68-93fb-8f4a829786b0 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Attention is All you Need
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 07b25ea8-2a8e-4604-810c-a4532a8e2b03 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A Simple Framework for Contrastive Learning of Visual Representations
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4dcb6c3-5bc3-4569-a471-ca88d75bbcca · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebefa7eb-9474-40c3-84b9-61dc1a853f1b · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning D4RL: Building Better Benchmarks for Offline Reinforcement Learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c06b20e-a5dc-42b0-86c5-cf226630a8c7 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/index.html
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5883a3ed-b7f8-4bac-8410-542612030a70 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Offlinerl-kit: An elegant pytorch offline reinforcement learning library
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation acc20467-b73a-4e61-aafb-0e89c97011d2 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Hierarchically decoupled imitation for morpho- logical transfer
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8b6faf0-90c2-4efb-a873-f36946864a93 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning PEARL: Zero-shot cross-task preference alignment and robust reward learning for robotic manipulation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation faa778ae-004a-4dde-8ba4-0f0d7a677f57 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69bc5927-4172-4956-8c45-9083497c7694 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning robust perceptive locomotion for quadrupedal robots in the wild
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1552f6-f733-4b52-9d09-1b90227dd761 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning dm_control: Software and tasks for continuous control
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e0ccee2-5793-45ba-b7b4-5e6d7602df08 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URLB: Unsupervised reinforcement learning benchmark
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 103bccc0-2245-448a-b121-76970064a4ac · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Continuous control with deep reinforcement learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce58efbb-6332-4e03-81a4-6bd802df36c9 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://docs.pytorch.org/ docs/stable/generated/torch.nn.TransformerEncoder.html
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c3c1e64-16a1-4f96-92c9-bfcd397821ca · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/hopper/
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26b71533-3a9e-45e8-96ae-1e3ad80676d8 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/walker2d/
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ab7ec7ee-36e9-419b-9380-ac83076703d1 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning For the Franka Kitchen tasks, we use the preference datasets from An et al
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 41a941b2-d2f4-404f-8b60-2b89f0b89425 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning ISSN: 2640-3498
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0347b73b-0a39-4a05-b278-31a057847404 · outbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.