Pith. sign in

Paper Citation Record · LEDGER

Training Agents using Upside-Down Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:1912.02877.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.02877 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:30:24.270417Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:57:25.915173Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 80bf680b-5b57-4d18-8841-ae1b5076cacc · inbound

Decision Transformer: Reinforcement Learning via Sequence Modeling cites this paper.

Decision Transformer: Reinforcement Learning via Sequence Modeling Training Agents using Upside-Down Reinforcement Learning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:11.117837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T15:11:11.056013Z digest=sha256:6c3775601767544526aca1da2fa2fdbe7f3d209e96c41cfc043bf6c444de0913

Observation 1baf92be-02ec-4bbb-be46-cfa2500d5d2e · inbound

Is Conditional Generative Modeling all you need for Decision-Making? cites this paper.

Is Conditional Generative Modeling all you need for Decision-Making? Training Agents using Upside-Down Reinforcement Learning

Reference 209

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:35:10.858704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T15:35:10.593969Z digest=sha256:59156edaab08dfa17193336bcad7c7069889d0fcf3e8e30be25e9cfd02601489

Observation 38f45699-690a-442d-9483-3fb582ba6590 · inbound

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework cites this paper.

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework Training Agents using Upside-Down Reinforcement Learning

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:43:19.125664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T03:43:18.632292Z digest=sha256:25434bd6243af83ddbd13e7fc30a3a2aa7b20e6da45847d26f570a9bfb336c59

Observation fc5c38cc-235b-47fc-a44a-eabe153027c2 · inbound

A Provable Approach for End-to-End Safe Reinforcement Learning cites this paper.

A Provable Approach for End-to-End Safe Reinforcement Learning Training Agents using Upside-Down Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.270417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.270417Z digest=sha256:125d734320bff5c7ec3717bc039b847570ba643fe98b579153c6500fb0ca56f8

Observation 1e69c1d7-7926-44ab-b3ac-bb44ef679d6c · inbound

How to Provably Improve Return Conditioned Supervised Learning? cites this paper.

How to Provably Improve Return Conditioned Supervised Learning? Training Agents using Upside-Down Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:08.014844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:08.014844Z digest=sha256:5688dbfe763b32a19081cdfab913d29634d5995d47226210e2ce9cd41970e450

Observation 325a3a7e-8201-4ff1-971f-797eabf7a0e8 · inbound

Behavioral Exploration: Learning to Explore via In-Context Adaptation cites this paper.

Behavioral Exploration: Learning to Explore via In-Context Adaptation Training Agents using Upside-Down Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:46.817188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:15:46.817188Z digest=sha256:32abb52fb46c341a0a9cd54695208cb9d504d9ea0907648721d52e954a5f3eb6

Observation 715459ac-bfe6-476e-b793-aea24c7a4984 · inbound

Equivariant Goal Conditioned Contrastive Reinforcement Learning cites this paper.

Equivariant Goal Conditioned Contrastive Reinforcement Learning Training Agents using Upside-Down Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.543536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.543536Z digest=sha256:33a342b83f8a18fd1d70d4a33982ac4b91263b506b97f206e7e16c9307ec5fba

Observation 28fa6eef-fb21-4a98-897d-7278f3b6ea36 · inbound

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration cites this paper.

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration Training Agents using Upside-Down Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:14.448302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:14.448302Z digest=sha256:3fc4719ad5919ef12a3d4da11cf5cda7197f91620b40076d5bed1a2905460a0d

Observation 310d9c01-bbdb-4460-8592-6e9a51c52323 · inbound

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success cites this paper.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Training Agents using Upside-Down Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.544642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.544642Z digest=sha256:3ed9b65b101a1c3244bccaed5354e3bfee4685d48c078306fdaf3a516176548a

Observation 07d1701b-aded-4818-9ad6-1aa0f6fc72eb · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Training Agents using Upside-Down Reinforcement Learning

Reference 232

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:58.178293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:a6ae058feae7452b8e86df140340fabad5744c72d5c61e44a4f5fc16c0f414df

Observation b33bb334-208a-4f7b-b6bd-9a6d9280bba5 · inbound

Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies cites this paper.

Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies Training Agents using Upside-Down Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.916771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T19:23:32.966503Z digest=sha256:54f4d7d23a88a5f9a0263ce2f1c5032f9cfd81434b515eefd4de43f3681685b2

Observation fc21b51e-44c7-4011-86bc-1d869c0540fc · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Training Agents using Upside-Down Reinforcement Learning

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:14.105428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:14.105428Z digest=sha256:28cfda145416363e4c9b61bf5e8fb9a1a8deb524999297e3b539664bcc100a43

Observation 008ffa89-a4da-4946-90dd-a2c1f29c159c · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Training Agents using Upside-Down Reinforcement Learning

Reference 269

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:32.197052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:32.197052Z digest=sha256:372ef46283232378f835f749f80e81090ef016ddc1397d58be3eb7b5a32038e3