Pith. sign in

Paper Citation Record · LEDGER

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models

As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2506.20061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20061 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:03:33.026420Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T00:46:28.503923Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee5ef227-57f1-48ca-94f1-05a33cc845a6 · outbound

This paper cites Universal value function approximators.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Universal value function approximators

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.816841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T23:03:29.373096Z digest=sha256:59abe5f5bf9e30ca3082aff6adbc2714ee242a568d0eb8987c4ccfee6634583e

Observation a7e33da3-9a98-4b9b-9623-c7aba574a393 · outbound

This paper cites Hindsight experience replay.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Hindsight experience replay

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:29.500447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:29.500447Z digest=sha256:4c730a71c81823ad05e5898d11c1e9ce6781b1908375e2f69c57d2c762e671c5

Observation f9e5e4fc-9e4b-413a-b71b-5f24f81619dc · outbound

This paper cites Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:29.615893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:29.615893Z digest=sha256:5916696262395106a18332301913d5dbb0577114ce5cab4c76f532f946451573

Observation ce835828-b91f-43ec-9f59-4485abcca790 · outbound

This paper cites Grounding language for transfer in deep reinforcement learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Grounding language for transfer in deep reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.565489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T23:03:29.777682Z digest=sha256:788f786f8dbb65cda0a8c1521e349ec27c49ac7c467e4406de46981d053fb6f6

Observation 37532636-43af-4b52-8434-53dd47ee7afa · outbound

This paper cites Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:29.905488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:29.905488Z digest=sha256:5727ef1d5e70ff83d0d45ecadf90facd6c2a0b74681e89a3ef1234bdde994d02

Observation 7974a871-e96a-47a5-9f66-4fe6ab15f670 · outbound

This paper cites Goal-Conditioned Reinforcement Learning: Problems and Solutions.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.107695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.107695Z digest=sha256:59f9d85cad7d15b8ccb2879a4fb37c8be2630e8916339d6e2ed6054a4629048d

Observation b101b800-f6c9-4d69-9320-e99fe1209ddf · outbound

This paper cites Maximum entropy gain exploration for long horizon multi-goal reinforcement learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Maximum entropy gain exploration for long horizon multi-goal reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.295432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T23:03:30.241905Z digest=sha256:6b3ff8bd163b6b63ecd08f5599f193dca698ac908e53f188f5e6c8144e5986c8

Observation ef5a2712-587e-4d1b-b38e-843eac71b0c4 · outbound

This paper cites Curriculum-guided hindsight experience replay.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Curriculum-guided hindsight experience replay

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.378167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.378167Z digest=sha256:dc7c210375c17f1e58df340c1f8a1da53d9d205d4b4aabdc3e82b02b0b6cc398

Observation 38706882-630a-49cd-93d4-47fb0f010657 · outbound

This paper cites Exploration via hindsight goal generation.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Exploration via hindsight goal generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.546500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.546500Z digest=sha256:3f0d459079326e5b37fb13aeec97a28cc1843c403f6738b4d11be764fcb3aca5

Observation cb3ba943-5387-44ad-bab6-55dd54832e3f · outbound

This paper cites Visual reinforcement learning with imagined goals.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Visual reinforcement learning with imagined goals

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.719316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.719316Z digest=sha256:fde51b561c84e5ab621bcdfec79c0255ae2856a1807d3f475283557d08f94676

Observation 0676bdcf-9d3c-4064-a8a6-e31cf92c2201 · outbound

This paper cites Unsupervised Control Through Non-Parametric Discriminative Rewards.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Unsupervised Control Through Non-Parametric Discriminative Rewards

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.883300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.883300Z digest=sha256:7321aa71f42da3438de6d082e074390409725820a5bac5fce50114050ef93078

Observation 9d332c08-ac5d-4255-8d10-a5b3c3b91e42 · outbound

This paper cites A Survey of Reinforcement Learning Informed by Natural Language.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models A Survey of Reinforcement Learning Informed by Natural Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:31.066702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:31.066702Z digest=sha256:a96a613135fc28237371b3a70f7e845df5a3a408c4ab7c5900d87c3f2b4dff74

Observation cfa8ee4d-b0b4-4fb2-b154-95b72e688eac · outbound

This paper cites The wisdom of hindsight makes language models better instruction followers.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models The wisdom of hindsight makes language models better instruction followers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.061171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.203145Z digest=sha256:a9532d22c03da2d7d99a71a2fb09243e5f0038d3edd5a8ebd397903ad4532a05

Observation 6d3a108f-8727-45c9-92fd-7970ba17c401 · outbound

This paper cites Analyzing the Stance of Facebook Posts on Abortion Considering State-level Health and Social Compositions.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Analyzing the Stance of Facebook Posts on Abortion Considering State-level Health and Social Compositions

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T23:03:33.440232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.331751Z digest=sha256:064ef3e1f6bd0a39369227f54bd1cb040906d0138d844d2aaa6b2c5189dd3d44

Observation 154d05e9-42a3-44d3-a35f-73f261b2fe91 · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:31.451743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:31.451743Z digest=sha256:89a26771ff2d7d3407522456d72cb8d097b17411b4975e0fe887a74841b262ba

Observation ce92ba1f-c190-465b-bf2f-deb45477acf8 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:34.726003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.573213Z digest=sha256:2df30f657036dd292c2ecc3153d91deecf50a53c5e10bb947b0fb2278d3eada6

Observation 87ef1c50-298c-4ceb-91aa-0af950e38492 · outbound

This paper cites Inferring Rewards from Language in Context.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Inferring Rewards from Language in Context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:31.747052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:31.747052Z digest=sha256:8fd072005ca57a4d2b24ed31de0db10293496bb3625904772be5340394722f90

Observation 472cdb42-28bf-4cde-9f5a-e3c525469997 · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Guiding pretraining in reinforcement learning with large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:34.435119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.894792Z digest=sha256:63fc96950cf40c5b78ceb322fa758cbb8325b856a4adc101bfd3e61d6c5e29df

Observation d4cb4566-d3f4-470e-87b5-8ecfd58751f8 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models React: Synergizing reasoning and acting in language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.011774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.011774Z digest=sha256:24b32b5131046c5d7165f84cfb2c488f71dcbc85a51a05c628894134559080de

Observation 125b9fc7-cde6-42f3-bf25-12de208fc593 · outbound

This paper cites Minedojo: Building open-ended embodied agents with internet-scale knowledge.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Minedojo: Building open-ended embodied agents with internet-scale knowledge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.159440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.159440Z digest=sha256:daf9925626a24d78813bd975b58c4ff4b0fd6e331bd5ca57a45b556a08ac63ef

Observation 4899dffe-6d59-40f4-b0e2-6883148a6f09 · outbound

This paper cites M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.302643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.302643Z digest=sha256:886f5bfc27cdcfdaa7dadf5fb0f5d2df2f8cb1c2c7af39636c72e68edd8fbd7d

Observation 68263a5e-58f2-4328-8d97-e32ffec8b761 · outbound

This paper cites Puterman.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Puterman

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:34.043969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T23:03:32.436169Z digest=sha256:79e1e2e634446b320e432e3200b5ef414b58ecaf398452fae9076c0b0fecb918

Observation 5d3772f7-f1c5-4ac3-82f1-a90506281224 · outbound

This paper cites Sutton and Andrew G.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Sutton and Andrew G

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.619630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.619630Z digest=sha256:fd00a7434df27430751ab150d835d846a560a78d2c06b701d684573e696f00c6

Observation c80792ea-c059-4d23-9320-953d70e826d4 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.776946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.776946Z digest=sha256:0556fc90cca36c1cb3e5921c97a06b92d6e118cbe8dadf343b73002a56eb88d6

Observation 625ffa4f-d64e-42a5-9af4-316d21545853 · outbound

This paper cites Simplifying Deep Temporal Difference Learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Simplifying Deep Temporal Difference Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.888434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.888434Z digest=sha256:bf58493c7951f4b47bdd8c9c34516795fec53c05031a1ac147cfe511ebf5252e

Observation 191974e1-80f5-4535-ae57-ebc654313a33 · outbound

This paper cites Prioritized level replay.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Prioritized level replay

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:33.846448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T23:03:33.026420Z digest=sha256:427be483951c0492d00ee1298aea5f8baa22e15fbe414719d15ac2fd21ef3fd8

Pith citing papers

Observation 0e39e0b9-01c2-4be0-9674-7f933d8fded8 · inbound

Learning More from Less: Reinforcement Learning from Hindsight cites this paper.

Learning More from Less: Reinforcement Learning from Hindsight Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T00:46:28.503923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:46:28.503923Z digest=sha256:d1eedd28a558273182a1f1943ba286ac8ce0d816bc42bde781ad2510f49f4b05