Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2106.05091.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:59.030007Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:40:06.257016Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 006cde4c-6c80-49fd-81c9-d12af47de548 · inbound
Active teacher selection for reward learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8190da58-22ab-43de-a3f7-162654f53e83 · inbound
CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a675f9e-6ba7-40da-b27d-cadf3d44b6e5 · inbound
SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f862af4-11fc-402b-ac64-e871402ee531 · inbound
Residual Reward Models for Preference-based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee2d89a7-e9a9-4d9b-ab0d-6fb0937dee98 · inbound
CueLearner: Bootstrapping and local policy adaptation from relative feedback PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf990ad4-6fe6-4e04-848f-c4f42f575879 · inbound
HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fb53c3-1d03-4e44-9b60-4f6e10056144 · inbound
Active Query Selection for Crowd-Based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10b2cd0d-6808-4b23-a199-26ae8d81cedd · inbound
Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294511cf-d0ef-4e06-a517-4c9dee980e38 · inbound
RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7127ee49-baaf-4be9-99f8-d999492dfd71 · inbound
Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8544e529-e9bf-48bc-baa1-3c64508c4798 · inbound
MAPL: Multi-Objective Preference Learning for Robot Locomotion PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.