Pith. sign in

Paper Citation Record · LEDGER

PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2106.05091.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.05091 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:59.030007Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:06.257016Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 006cde4c-6c80-49fd-81c9-d12af47de548 · inbound

Active teacher selection for reward learning cites this paper.

Active teacher selection for reward learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:56:01.750823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T05:54:43.873174Z digest=sha256:79319896e00943b74428545cbad9b47da707e748cd3a9413aee6d0d23f1db2d4

Observation 8190da58-22ab-43de-a3f7-162654f53e83 · inbound

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries cites this paper.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.030007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.030007Z digest=sha256:03abee6c2b43971b5690a319a3758abcb9c36bd8ee17bb8bddaa2642651c4d96

Observation 2a675f9e-6ba7-40da-b27d-cadf3d44b6e5 · inbound

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning cites this paper.

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:31:36.503042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T13:27:48.404574Z digest=sha256:174ef81939fda6bebca7190d89eb83b3a756eb020cdf7b677bb91e7188135a22

Observation 9f862af4-11fc-402b-ac64-e871402ee531 · inbound

Residual Reward Models for Preference-based Reinforcement Learning cites this paper.

Residual Reward Models for Preference-based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:36.384982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:36.384982Z digest=sha256:daf218810bc0ddfe259878d79f538f983bce2dc9df4aec5dcd8b76ded783ca74

Observation ee2d89a7-e9a9-4d9b-ab0d-6fb0937dee98 · inbound

CueLearner: Bootstrapping and local policy adaptation from relative feedback cites this paper.

CueLearner: Bootstrapping and local policy adaptation from relative feedback PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:57.368857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:57.368857Z digest=sha256:004c19ca4ff0226be7224e5eb09efb3c2d0ea20438131fe747ba5b8b60d0b1ab

Observation bf990ad4-6fe6-4e04-848f-c4f42f575879 · inbound

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation cites this paper.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.210585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.210585Z digest=sha256:ddfd1bc4ea8e75522cbece8f318b0abf4f5b5323720e462f28b6fb680eb17fd7

Observation 58fb53c3-1d03-4e44-9b60-4f6e10056144 · inbound

Active Query Selection for Crowd-Based Reinforcement Learning cites this paper.

Active Query Selection for Crowd-Based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:02:25.976213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:02:25.976213Z digest=sha256:8266aa9b2ad12daec4d38130da67e6f2e19749c768d2eee7ec6b1c232581bb9d

Observation 10b2cd0d-6808-4b23-a199-26ae8d81cedd · inbound

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning cites this paper.

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T16:06:56.860428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:06:56.860428Z digest=sha256:a1b0a8d4eae3b52b769b9455b86eb395eabbc964d6ed66193d71845f1a7c9da6

Observation 294511cf-d0ef-4e06-a517-4c9dee980e38 · inbound

RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview cites this paper.

RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T22:33:25.723393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:33:25.723393Z digest=sha256:e894eaadfea1ca0e9903d5e4321a9cdaadc6631ecd45933d8e623954b5ac1583

Observation 7127ee49-baaf-4be9-99f8-d999492dfd71 · inbound

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback cites this paper.

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.090762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T23:35:03.577967Z digest=sha256:e5407b5a5bbea6d474d713cbec0c05327d0d6e273ce8275df0866b7bf00ef68a

Observation 8544e529-e9bf-48bc-baa1-3c64508c4798 · inbound

MAPL: Multi-Objective Preference Learning for Robot Locomotion cites this paper.

MAPL: Multi-Objective Preference Learning for Robot Locomotion PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:06.258720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T21:11:00.753949Z digest=sha256:c54dc6fff7f0619e119480817e3553f70534fa33b156200e1c40d009b673f18c