Pith. sign in

Paper Citation Record · LEDGER

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success

As of 7 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 0 inbound Pith citation observations for arXiv:2601.18175.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.18175 v2

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T08:14:02.691163Z

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09029daa-0b2c-4c2e-86cb-a3943bd44ca3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.610152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.610152Z digest=sha256:13a2f8f55835c09250c11d9b3471290a8d7028538bc30e5fde575b0f59afdde0

Observation 5c8376d4-1666-4d32-9355-6b519a11430a · outbound

This paper cites The first equality is a definition of Lπ0 (π+).

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success The first equality is a definition of Lπ0 (π+)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.691163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.691163Z digest=sha256:f94fc0af97d88eeda0eef683678eee74b2625295b4b43e402f8636d0a89504c9

Observation 62f79463-8cd1-407f-86c7-4435e5607a42 · outbound

This paper cites Bhandari and D.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Bhandari and D

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.023064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.023064Z digest=sha256:4723aefbd39973b14055b8ba346a34acf2455d4036691f527344f540c4b56e3c

Observation d529df4a-4268-4725-bf0c-dea6b014cf42 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.464290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.464290Z digest=sha256:55670e0ed68f71ee0089987abea0eb062e36202b1908c8ba59612a38f5930064

Observation 310d9c01-bbdb-4460-8592-6e9a51c52323 · outbound

This paper cites Training Agents using Upside-Down Reinforcement Learning.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Training Agents using Upside-Down Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.544642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.544642Z digest=sha256:3ed9b65b101a1c3244bccaed5354e3bfee4685d48c078306fdaf3a516176548a

Observation bcb313a7-679a-48e0-a15c-31d27690d890 · outbound

This paper cites an unresolved cited work.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Unresolved cited work

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.093983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.093983Z digest=sha256:8e115e6f59d35ea51c35a97380711e807d91d88863ca95713f44e12b5b4e3d8e

Observation b281e5ca-9f1b-4672-afe5-f200cffd2c3f · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.253123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.253123Z digest=sha256:ee692cab839521e9b22e87b99bf661d80af52f6b916f09d316a45639b442916c

Observation 221a88ec-2bdf-4443-ac4e-ba250853f5da · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.352724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.352724Z digest=sha256:0d03a3c4dabe4c459c305b8513d2f5e716b3d03911e6082b769a4f3acae0d739

Observation 58bcb1bc-0d35-4f65-8783-f27f9934e8e3 · outbound

This paper cites The Llama 3 Herd of Models.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success The Llama 3 Herd of Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.173867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.173867Z digest=sha256:b8b3e070512af4c30ce8c27b3d01ab64b9c682575a46be98687703fb6bf1741f

Pith citing papers

No inbound Pith citation observations are available.