Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:03:33.026420Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2506.20061.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:03:33.026420Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T00:46:28.503923Z
A source-named dated measurement, never combined with another source.
Source: cited_works
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ee5ef227-57f1-48ca-94f1-05a33cc845a6 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Universal value function approximators
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a7e33da3-9a98-4b9b-9623-c7aba574a393 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Hindsight experience replay
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e5e4fc-9e4b-413a-b71b-5f24f81619dc · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce835828-b91f-43ec-9f59-4485abcca790 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Grounding language for transfer in deep reinforcement learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 37532636-43af-4b52-8434-53dd47ee7afa · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7974a871-e96a-47a5-9f66-4fe6ab15f670 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Goal-Conditioned Reinforcement Learning: Problems and Solutions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b101b800-f6c9-4d69-9320-e99fe1209ddf · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Maximum entropy gain exploration for long horizon multi-goal reinforcement learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ef5a2712-587e-4d1b-b38e-843eac71b0c4 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Curriculum-guided hindsight experience replay
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38706882-630a-49cd-93d4-47fb0f010657 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Exploration via hindsight goal generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb3ba943-5387-44ad-bab6-55dd54832e3f · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Visual reinforcement learning with imagined goals
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0676bdcf-9d3c-4064-a8a6-e31cf92c2201 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Unsupervised Control Through Non-Parametric Discriminative Rewards
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d332c08-ac5d-4255-8d10-a5b3c3b91e42 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models A Survey of Reinforcement Learning Informed by Natural Language
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa8ee4d-b0b4-4fb2-b154-95b72e688eac · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models The wisdom of hindsight makes language models better instruction followers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d3a108f-8727-45c9-92fd-7970ba17c401 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Analyzing the Stance of Facebook Posts on Abortion Considering State-level Health and Social Compositions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 154d05e9-42a3-44d3-a35f-73f261b2fe91 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Eureka: Human-Level Reward Design via Coding Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce92ba1f-c190-465b-bf2f-deb45477acf8 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 87ef1c50-298c-4ceb-91aa-0af950e38492 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Inferring Rewards from Language in Context
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 472cdb42-28bf-4cde-9f5a-e3c525469997 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Guiding pretraining in reinforcement learning with large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d4cb4566-d3f4-470e-87b5-8ecfd58751f8 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models React: Synergizing reasoning and acting in language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125b9fc7-cde6-42f3-bf25-12de208fc593 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Minedojo: Building open-ended embodied agents with internet-scale knowledge
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4899dffe-6d59-40f4-b0e2-6883148a6f09 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68263a5e-58f2-4328-8d97-e32ffec8b761 · outbound
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d3772f7-f1c5-4ac3-82f1-a90506281224 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Sutton and Andrew G
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c80792ea-c059-4d23-9320-953d70e826d4 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 625ffa4f-d64e-42a5-9af4-316d21545853 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Simplifying Deep Temporal Difference Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191974e1-80f5-4535-ae57-ebc654313a33 · outbound
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Prioritized level replay
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e39e0b9-01c2-4be0-9674-7f933d8fded8 · inbound
Learning More from Less: Reinforcement Learning from Hindsight Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.