Pith. sign in

Paper Citation Record · LEDGER

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2505.06079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06079 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:52:24.226578Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:02.820968Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T02:25:55.886294Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy39
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2e85a03-dd5b-4903-99a0-9c65b6b6053f · outbound

This paper cites Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:27.356075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.680173Z digest=sha256:ea8a8bee05fd7b2963e4585f9626a44845bc4ce1e4803cbb6b2ca0c02c0428a4

Observation e0148500-b221-443c-9ac9-83b1ad8eaf3f · outbound

This paper cites Surf: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Surf: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:27.331141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.688126Z digest=sha256:c662f8657dbf22004694b5bb7ab421fea38cf270040dddbc09979d53da219987

Observation 347a37e1-a656-476b-bf35-ca5ddc5569fd · outbound

This paper cites Active preference-based learning of reward functions,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Active preference-based learning of reward functions,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:27.306467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.704362Z digest=sha256:bc83704b6e883a6402ff12d477a436c7c5245666b43cd7bc5fed63cf2af7362a

Observation d5f7e1f5-b275-4787-ae92-5636312e01f9 · outbound

This paper cites Wayex: Waypoint exploration using a single demonstration,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Wayex: Waypoint exploration using a single demonstration,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:27.268190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.712170Z digest=sha256:e58827982b601c262bb6ac838c6fd108fac5d1dee843aecff7ea9adf57edaf36

Observation bf61eb4a-90a0-4d9f-98ce-2673d3b3e562 · outbound

This paper cites Online human training of a myoelec- tric prosthesis controller via actor-critic reinforcement learn- ing,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Online human training of a myoelec- tric prosthesis controller via actor-critic reinforcement learn- ing,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:27.212477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.722615Z digest=sha256:d999cb9d7b2dd53c2782c180b8397b04d79088d05fd0a54a0e0426dba415b92d

Observation d8f80d31-5c8f-4286-9e69-0b591ec015d7 · outbound

This paper cites A bayesian approach for policy learning from trajectory preference queries,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations A bayesian approach for policy learning from trajectory preference queries,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:27.178340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.727644Z digest=sha256:f3ee11ebe2578740b175f7317cfd22b855139887a1ea4dad0595216987825f4f

Observation e4f19d52-326b-4655-b7c2-8bb5130b9fac · outbound

This paper cites Model-free preference-based reinforcement learning,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Model-free preference-based reinforcement learning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:27.135395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.754379Z digest=sha256:0d97f02da5968af48926fdd59837f6d29510bc00f644d69394972a97fcd82b15

Observation cb0a0399-3f18-42a0-bdf4-76f5bfcbbff7 · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:23.759155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:23.759155Z digest=sha256:fa06a169a315a082441e8e3491a80c3b8930da645c3ef7d6c22a82ae16549edb

Observation 425e39b7-0f1b-437d-87a1-b2bf58bb4687 · outbound

This paper cites B-pref: Benchmarking preference-based reinforcement learning,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations B-pref: Benchmarking preference-based reinforcement learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:27.087427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.764185Z digest=sha256:0dd5867723798b3f675279b747cdc985c4aadc982901aa7e275f3828420661ea

Observation 57d4e03e-e1e6-43bc-ba8e-64325dcbdd8c · outbound

This paper cites Rime: Robust preference-based reinforcement learning with noisy preferences,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Rime: Robust preference-based reinforcement learning with noisy preferences,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:27.036722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.772793Z digest=sha256:ee981c2505d8841cd1297723cef6d0d94e05bdf5daaceb57434501706bb7479d

Observation 9df6eba4-ca71-4096-a591-f4f0e531ab76 · outbound

This paper cites Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:23.783357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:23.783357Z digest=sha256:c00d7a7415bb0e9f3d5f97af16eaa8cd6ed2f62daca76b3db0a688a045ce50dc

Observation bcd98498-658b-4815-8a3a-ce7582b22bd4 · outbound

This paper cites Deep reinforcement learning from human pref- erences,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Deep reinforcement learning from human pref- erences,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.987163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.796861Z digest=sha256:fcfd0b177fbb477149f8a80a2bb7fb83bcf2640b69c50841f8e34ca193648a7a

Observation d06e3473-61aa-433c-b696-6440b961915a · outbound

This paper cites Reward learning from human preferences and demonstrations in atari,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reward learning from human preferences and demonstrations in atari,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.956679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.810404Z digest=sha256:5f292083c35e2ad3f41c3ad04928b74d56a35c75d1f2c518b7f32832caadb4cd

Observation 454dfea3-18a0-47fb-b593-bdaaea830df0 · outbound

This paper cites Batch active preference-based learn- ing of reward functions,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Batch active preference-based learn- ing of reward functions,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.927089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.815216Z digest=sha256:e0c019f53fd58198ba4cee9222f91a3b5e29355cc9543fb877361859833cf2d1

Observation 3369010c-f1da-4ca0-a102-b717adf3fbd0 · outbound

This paper cites Active Preference-Based Gaussian Process Regression for Reward Learning.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Active Preference-Based Gaussian Process Regression for Reward Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:23.819700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:23.819700Z digest=sha256:51ce3724e09d99a22a0f209bff30990d2f8cc0d3f67615a0941baff6e7bf0afa

Observation a9b516d6-919d-424e-8a61-dd977c1e1cb9 · outbound

This paper cites Reward uncertainty for exploration in preference-based reinforcement learning,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reward uncertainty for exploration in preference-based reinforcement learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.846709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.824839Z digest=sha256:194e0f04a5a159f916dac118d771ef2a94f21fd426d5374f57c9f6f1fec06a54

Observation 17db3971-04d1-4ce7-8fea-79c9fe144022 · outbound

This paper cites Learning from noisy labels with deep neural networks: A survey,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Learning from noisy labels with deep neural networks: A survey,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.756271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.845159Z digest=sha256:a940ae1e65a1cd37c52e11a5836df22b4a86e2d38088dce71719e2adebe7f354

Observation d7b30a1e-8da2-4bbb-9df4-941b35ff97c5 · outbound

This paper cites Denoising implicit feedback for recommendation,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Denoising implicit feedback for recommendation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.620633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.852870Z digest=sha256:ac727298f183cb6391d6f7c3751e16679c56e7a4931a351d825e48a3fbb219a8

Observation 94d01dd7-6ee4-4b61-adc5-73f494a19a1e · outbound

This paper cites Co-teaching: Robust training of deep neural networks with extremely noisy labels,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Co-teaching: Robust training of deep neural networks with extremely noisy labels,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.548642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.868255Z digest=sha256:b2afab8df57a5da9d21cf55e9d7b3fd00b70a41e2907672d0abb011d72dba75b

Observation 659a0e5e-01df-4e28-8da3-f142664a56e7 · outbound

This paper cites Training deep neural- networks using a noise adaptation layer,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Training deep neural- networks using a noise adaptation layer,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.433851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.873657Z digest=sha256:83013166307033a77e25d57ae1c38ad5e320b9fd4a2875ef84e4662cab21728e

Observation bf61dbfc-642d-4fdc-8098-093dbc1a813b · outbound

This paper cites Does label smoothing mitigate label noise?.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Does label smoothing mitigate label noise?

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.398056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.878520Z digest=sha256:5172dcabde8f0d3e37e7041c0a3a1bb8b55cbfd1f53c71c1eefa912b230017d6

Observation 809c688c-896b-4e6a-8f51-dfc29ab12a7b · outbound

This paper cites Reinforcement Learning from Diverse Human Preferences.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reinforcement Learning from Diverse Human Preferences

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:23.896876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:23.896876Z digest=sha256:088470a78c6fcf4e01555f76049a6a1e98d3d179abe1e4d05821be8df154d280

Observation 94d72d0d-ba29-4fb6-8d76-85939b25c775 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:23.911852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:23.911852Z digest=sha256:3ff734b34866208901fb0277ea0783fa5c4378f5919cf1f6ab59ed792d8cb55b

Observation f35022f0-c43f-4fb3-8710-8f8f839e699e · outbound

This paper cites What is point supervision worth in video instance segmentation?.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations What is point supervision worth in video instance segmentation?

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.377251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.920429Z digest=sha256:2a2389a7f52ade563bf4398507cf4bb71ea059a3aa7f94218ee3b1c3c4cd4b05

Observation e445d69a-2941-40cc-b4e9-9e1f01268b61 · outbound

This paper cites Uvis: Unsupervised video instance segmen- tation,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Uvis: Unsupervised video instance segmen- tation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.350294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.931684Z digest=sha256:bfd189db902a35152b1826da1c2d785284493bd38392f4394c7a57fdb6bd405a

Observation c825d4c1-175b-4a08-b48a-477946a698da · outbound

This paper cites LMPriors: Pre-Trained Language Models as Task-Specific Priors.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations LMPriors: Pre-Trained Language Models as Task-Specific Priors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:23.945688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:23.945688Z digest=sha256:db55009b498098edbff354e98a99a934db87097bd42007a860d5e061bb601464

Observation d42d39ea-908a-4dc0-9bd3-e1475742e268 · outbound

This paper cites Guiding pretraining in reinforce- ment learning with large language models,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Guiding pretraining in reinforce- ment learning with large language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.308706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.950878Z digest=sha256:df7b908340ba4c2eae742609f696481e98167bcf355f18e63032dfa1f5d2dd83

Observation 6732c782-95b1-4d40-87b0-5f456182d21d · outbound

This paper cites Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:23.958237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:23.958237Z digest=sha256:d70eacf5a4a3ee851a4d5fda00c1d96c4026002a474ecfd725836d4122ff8c67

Observation 4f78414c-e0db-4f4e-9249-de8d935ecf70 · outbound

This paper cites Structured attentions for visual question answering,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Structured attentions for visual question answering,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.248067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.963460Z digest=sha256:0a96a926ba958775baab00129def632e4861561499da5faec06e1ea7a14e7852

Observation 16abd0f0-a1cc-45f7-9da4-832a4dc15240 · outbound

This paper cites Learning semantic correspondence with sparse an- notations,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Learning semantic correspondence with sparse an- notations,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.152235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.968321Z digest=sha256:7b6a157ed249b4f16aa0bdf1f8f3173bd4ab583d9502a5638493afa7cc674a51

Observation 605291f7-9742-4acf-ae0b-745152967fb8 · outbound

This paper cites Confidence-aware adversarial learning for self-supervised semantic matching,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Confidence-aware adversarial learning for self-supervised semantic matching,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.080627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.978519Z digest=sha256:bbe23837a1d131f17eb4478c738bd16ad1d1fc914b75bf0106f6d148ae7b1af9

Observation 755a868c-18a9-4cab-b531-912632d5778f · outbound

This paper cites Motif: Intrinsic motivation from artificial intelligence feedback,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Motif: Intrinsic motivation from artificial intelligence feedback,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:26.001012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:23.983533Z digest=sha256:c269861551c37ee4759654357a5dc08fc3d749afd8264749198a63c2d176739d

Observation 3d5c9583-b766-4635-88ae-8745f8e5e585 · outbound

This paper cites Towards scalable neural representation for diverse videos,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Towards scalable neural representation for diverse videos,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.892235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.005974Z digest=sha256:aa9d8e501f5831fc045ff906dc151b42af0369025b6541ec3a721a53e20ef748

Observation a9f2a9e1-6d3c-4bd7-a1a9-099d2fa47a32 · outbound

This paper cites Dynamic context correspondence network for semantic alignment,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Dynamic context correspondence network for semantic alignment,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.869633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.018464Z digest=sha256:d9dcc3c1d419ee985f7a78286e0f56b3f60a7b802a23bea1bd3d140bfea4d381

Observation e0c4e6b8-b6d7-433b-b37e-b3560a0073ee · outbound

This paper cites ARDuP: Active Region Video Diffusion for Universal Policies.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations ARDuP: Active Region Video Diffusion for Universal Policies

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:24.024394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:24.024394Z digest=sha256:7faf2d384b3fec8b47366c5daf617aebd129c0496f4be9290594d30488cd9698

Observation b036ebd2-9411-435b-a204-df80976d3183 · outbound

This paper cites P3-po: Prescriptive point priors for visuo-spatial generalization of robot policies,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations P3-po: Prescriptive point priors for visuo-spatial generalization of robot policies,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.804858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.029571Z digest=sha256:6ea553fcfda80d6425d2538cc5f84410a4db7dbf8f5af0573bef1eda794b4998

Observation 1f38efc1-8e5f-40a4-9542-e481cabe3278 · outbound

This paper cites AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:24.036110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:24.036110Z digest=sha256:d79b2f7c50de9121a091de4bba59bc56cd03cc1e20fd5782adbedd0a6886036e

Observation f4347a6e-3b8c-48cc-99ab-08d9de61f246 · outbound

This paper cites Is Imitation All You Need? Generalized Decision-Making with Dual-Phase Training.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Is Imitation All You Need? Generalized Decision-Making with Dual-Phase Training

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:52:24.540392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.040918Z digest=sha256:abccc6f94944cb1addcc87c4d80d903618165e2d8549f502081344e37be3b6c6

Observation 291a3c4f-2bab-4cd5-917b-73a32052f909 · outbound

This paper cites PRISE: LLM-style sequence compression for learning temporal action abstractions in control,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations PRISE: LLM-style sequence compression for learning temporal action abstractions in control,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.780442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.051639Z digest=sha256:c27758deeed63d1424c130e1a1524964cf0567ed727b4b2279a0abddcd4cceb8

Observation 30359195-88b4-414e-8c2a-b27de4879bc7 · outbound

This paper cites Premier-taco is a few-shot policy learner: Pretraining mul- titask representation via temporal action-driven contrastive loss,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Premier-taco is a few-shot policy learner: Pretraining mul- titask representation via temporal action-driven contrastive loss,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.692940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.073680Z digest=sha256:d8ebcb21fa4bafd57d63ce055e67d6e09aaff1ef43ebf11bc9b7e0a59993f47b

Observation eaa4af82-dfed-4b83-a9c5-36fdb8271e13 · outbound

This paper cites TACO: Temporal latent action- driven contrastive loss for visual reinforcement learning,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations TACO: Temporal latent action- driven contrastive loss for visual reinforcement learning,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.655812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.088667Z digest=sha256:293049032cb75d2ffa38d99717c86e9aa874cdf9eceb81dc186afdb83e2a9f9b

Observation f084a526-0144-4825-bd6c-5133a1132641 · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:24.105867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:24.105867Z digest=sha256:422fa8419cf7ad150dbfcc4d2b6ae65211b19d8f3191b45f86eeb2101e34c9e4

Observation 58299a99-58f8-457c-b70b-e78e651891c2 · outbound

This paper cites Reward Design with Language Models.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reward Design with Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:24.111055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:24.111055Z digest=sha256:907fc43dc8a2182acd535e65ee4f2a6cf55364060958e470956fc1462392f504

Observation bf60b1ca-c877-4d7c-a23c-5396d3d49dbd · outbound

This paper cites Reinforcement learning: An introduction,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reinforcement learning: An introduction,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.622118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.117001Z digest=sha256:2d68fc7adab987b6ae10ef53ceb3a2e4109fbe5259203a2f2e2695fa65741c71

Observation 7e3206d1-26b7-4905-804f-7aa75a600cb2 · outbound

This paper cites Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.586756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.134948Z digest=sha256:6bbbf566dda4674f706a470cc26210213d8e02f15bf01b3d74f582708724c5ad

Observation 22b1ecfc-84e9-48c5-b1a7-520499d955fe · outbound

This paper cites Preference Transformer: Modeling Human Preferences using Transformers for RL.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:24.143019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:24.143019Z digest=sha256:67b4ca5f6fff1072bfb0230ca7fc83e661ac569650a636706faa6984aac30eb8

Observation d3d2fc51-f8c5-4372-a4fd-435e3661a722 · outbound

This paper cites Rank analysis of incom- plete block designs: I. the method of paired comparisons,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Rank analysis of incom- plete block designs: I. the method of paired comparisons,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.543883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.155929Z digest=sha256:3f7690b7acaf08a6fd59f8d99b95d747b56ea9ced80e35e26bc393dd22abe378

Observation 2592a74d-851e-46be-83ba-5d3bd231ff49 · outbound

This paper cites Three-teaching: A three-way decision framework to handle noisy labels,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Three-teaching: A three-way decision framework to handle noisy labels,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.476004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.174521Z digest=sha256:61a1f18f211b254a0b98232e5adf2eeb2ca3031972c0c21b81a843185576f28b

Observation 8a917d95-51e0-4e03-b7e6-0bde1080d9d9 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:24.198999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:24.198999Z digest=sha256:21545362f67d76b692c62e55130b52c658969cf501718c36d623118b4ac928c0

Observation 1e65c652-ed01-4230-93aa-2bc5d7e87d61 · outbound

This paper cites Nearest neighbor estimates of entropy,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Nearest neighbor estimates of entropy,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.422313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.208749Z digest=sha256:ab4eb9384c5c7ba82d40f07950cd8b409281b8e68170209c080415d5b00d0ddf

Observation 5b66dbde-c3fe-490b-9fa8-e377e24a9b87 · outbound

This paper cites Soft actor- critic: Off-policy maximum entropy deep reinforcement learn- ing with a stochastic actor,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Soft actor- critic: Off-policy maximum entropy deep reinforcement learn- ing with a stochastic actor,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.376078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.218391Z digest=sha256:9925f08e3a2881cabe85fe8a448bc1a82958c1f62779a3e571cc632e1928b725

Observation 05e69f5a-eb38-4101-b8fc-c2d38e73e0c8 · outbound

This paper cites Gemini: A family of highly capable multimodal models,.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Gemini: A family of highly capable multimodal models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.338214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.226578Z digest=sha256:0ab72cbabfc1d3a2d6de6af90e6d9a2494d202a228980b103d413d603e0efc31

Observation 2ca1742f-2fdb-480e-83e3-c40e37230947 · outbound

This paper cites Reward Uncertainty for Exploration in Preference-based Reinforcement Learning.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:23.829934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:23.829934Z digest=sha256:fafcf03c6fb5aba697d4a3317b79c4d8e83db0dc91992361383d249cac496dac

Observation d09e3dd3-6de9-48af-818e-5448436fe3e6 · outbound

This paper cites Motif: Intrinsic Motivation from Artificial Intelligence Feedback.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Motif: Intrinsic Motivation from Artificial Intelligence Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:23.991919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:23.991919Z digest=sha256:7e0f17442e1051c6793a28c77c7291850b7623ec5dd793cdc1a68ddca2195347

Observation 5ee37f6f-7f98-430f-8e2a-9b515e862e19 · outbound

This paper cites Available: https://openreview.net/forum?id= p225Od0aYt.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Available: https://openreview.net/forum?id= p225Od0aYt

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:52:25.729269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:52:24.059555Z digest=sha256:82b6add76cf380523f2e58a39868d3ec1c263f5e621a357680d32793a854f18b

Pith citing papers

Observation b34a91bc-ede1-4ca6-9fce-25bd053e69b7 · inbound

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning cites this paper.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.820968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.820968Z digest=sha256:f79ad827093beb05afbee8a8f2ea1b3ec6b3297635c813e90c4d46a369bfe08a

Observation 1ce6d9bd-f0e1-424a-8712-342fe507c107 · inbound

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF cites this paper.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.887521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:09016af87fec4f1b822b9e2e07e251c3445c54fa15319e471e4a296ed93714ab