Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:52:24.226578Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2505.06079.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:52:24.226578Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:02.820968Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T02:25:55.886294Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a2e85a03-dd5b-4903-99a0-9c65b6b6053f · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e0148500-b221-443c-9ac9-83b1ad8eaf3f · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Surf: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 347a37e1-a656-476b-bf35-ca5ddc5569fd · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Active preference-based learning of reward functions,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d5f7e1f5-b275-4787-ae92-5636312e01f9 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Wayex: Waypoint exploration using a single demonstration,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bf61eb4a-90a0-4d9f-98ce-2673d3b3e562 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Online human training of a myoelec- tric prosthesis controller via actor-critic reinforcement learn- ing,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d8f80d31-5c8f-4286-9e69-0b591ec015d7 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations A bayesian approach for policy learning from trajectory preference queries,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e4f19d52-326b-4655-b7c2-8bb5130b9fac · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Model-free preference-based reinforcement learning,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cb0a0399-3f18-42a0-bdf4-76f5bfcbbff7 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 425e39b7-0f1b-437d-87a1-b2bf58bb4687 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations B-pref: Benchmarking preference-based reinforcement learning,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 57d4e03e-e1e6-43bc-ba8e-64325dcbdd8c · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Rime: Robust preference-based reinforcement learning with noisy preferences,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9df6eba4-ca71-4096-a591-f4f0e531ab76 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd98498-658b-4815-8a3a-ce7582b22bd4 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Deep reinforcement learning from human pref- erences,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d06e3473-61aa-433c-b696-6440b961915a · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reward learning from human preferences and demonstrations in atari,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 454dfea3-18a0-47fb-b593-bdaaea830df0 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Batch active preference-based learn- ing of reward functions,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3369010c-f1da-4ca0-a102-b717adf3fbd0 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Active Preference-Based Gaussian Process Regression for Reward Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b516d6-919d-424e-8a61-dd977c1e1cb9 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reward uncertainty for exploration in preference-based reinforcement learning,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 17db3971-04d1-4ce7-8fea-79c9fe144022 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Learning from noisy labels with deep neural networks: A survey,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d7b30a1e-8da2-4bbb-9df4-941b35ff97c5 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Denoising implicit feedback for recommendation,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 94d01dd7-6ee4-4b61-adc5-73f494a19a1e · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Co-teaching: Robust training of deep neural networks with extremely noisy labels,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 659a0e5e-01df-4e28-8da3-f142664a56e7 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Training deep neural- networks using a noise adaptation layer,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bf61dbfc-642d-4fdc-8098-093dbc1a813b · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Does label smoothing mitigate label noise?
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 809c688c-896b-4e6a-8f51-dfc29ab12a7b · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reinforcement Learning from Diverse Human Preferences
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94d72d0d-ba29-4fb6-8d76-85939b25c775 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f35022f0-c43f-4fb3-8710-8f8f839e699e · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations What is point supervision worth in video instance segmentation?
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e445d69a-2941-40cc-b4e9-9e1f01268b61 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Uvis: Unsupervised video instance segmen- tation,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c825d4c1-175b-4a08-b48a-477946a698da · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations LMPriors: Pre-Trained Language Models as Task-Specific Priors
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d42d39ea-908a-4dc0-9bd3-e1475742e268 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Guiding pretraining in reinforce- ment learning with large language models,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6732c782-95b1-4d40-87b0-5f456182d21d · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f78414c-e0db-4f4e-9249-de8d935ecf70 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Structured attentions for visual question answering,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 16abd0f0-a1cc-45f7-9da4-832a4dc15240 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Learning semantic correspondence with sparse an- notations,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 605291f7-9742-4acf-ae0b-745152967fb8 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Confidence-aware adversarial learning for self-supervised semantic matching,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 755a868c-18a9-4cab-b531-912632d5778f · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Motif: Intrinsic motivation from artificial intelligence feedback,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d5c9583-b766-4635-88ae-8745f8e5e585 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Towards scalable neural representation for diverse videos,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a9f2a9e1-6d3c-4bd7-a1a9-099d2fa47a32 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Dynamic context correspondence network for semantic alignment,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e0c4e6b8-b6d7-433b-b37e-b3560a0073ee · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations ARDuP: Active Region Video Diffusion for Universal Policies
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b036ebd2-9411-435b-a204-df80976d3183 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations P3-po: Prescriptive point priors for visuo-spatial generalization of robot policies,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1f38efc1-8e5f-40a4-9542-e481cabe3278 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4347a6e-3b8c-48cc-99ab-08d9de61f246 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Is Imitation All You Need? Generalized Decision-Making with Dual-Phase Training
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 291a3c4f-2bab-4cd5-917b-73a32052f909 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations PRISE: LLM-style sequence compression for learning temporal action abstractions in control,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 30359195-88b4-414e-8c2a-b27de4879bc7 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Premier-taco is a few-shot policy learner: Pretraining mul- titask representation via temporal action-driven contrastive loss,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eaa4af82-dfed-4b83-a9c5-36fdb8271e13 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations TACO: Temporal latent action- driven contrastive loss for visual reinforcement learning,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f084a526-0144-4825-bd6c-5133a1132641 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58299a99-58f8-457c-b70b-e78e651891c2 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reward Design with Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf60b1ca-c877-4d7c-a23c-5396d3d49dbd · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reinforcement learning: An introduction,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7e3206d1-26b7-4905-804f-7aa75a600cb2 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 22b1ecfc-84e9-48c5-b1a7-520499d955fe · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d2fc51-f8c5-4372-a4fd-435e3661a722 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Rank analysis of incom- plete block designs: I. the method of paired comparisons,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2592a74d-851e-46be-83ba-5d3bd231ff49 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Three-teaching: A three-way decision framework to handle noisy labels,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8a917d95-51e0-4e03-b7e6-0bde1080d9d9 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e65c652-ed01-4230-93aa-2bc5d7e87d61 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Nearest neighbor estimates of entropy,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5b66dbde-c3fe-490b-9fa8-e377e24a9b87 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Soft actor- critic: Off-policy maximum entropy deep reinforcement learn- ing with a stochastic actor,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 05e69f5a-eb38-4101-b8fc-c2d38e73e0c8 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Gemini: A family of highly capable multimodal models,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2ca1742f-2fdb-480e-83e3-c40e37230947 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reward Uncertainty for Exploration in Preference-based Reinforcement Learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d09e3dd3-6de9-48af-818e-5448436fe3e6 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Motif: Intrinsic Motivation from Artificial Intelligence Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee37f6f-7f98-430f-8e2a-9b515e862e19 · outbound
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Available: https://openreview.net/forum?id= p225Od0aYt
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b34a91bc-ede1-4ca6-9fce-25bd053e69b7 · inbound
PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce6d9bd-f0e1-424a-8712-342fe507c107 · inbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.