Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T08:14:02.691163Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 0 inbound Pith citation observations for arXiv:2601.18175.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T08:14:02.691163Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
9 of 9 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 09029daa-0b2c-4c2e-86cb-a3943bd44ca3 · outbound
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8376d4-1666-4d32-9355-6b519a11430a · outbound
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success The first equality is a definition of Lπ0 (π+)
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f79463-8cd1-407f-86c7-4435e5607a42 · outbound
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Bhandari and D
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d529df4a-4268-4725-bf0c-dea6b014cf42 · outbound
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Proximal Policy Optimization Algorithms
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 310d9c01-bbdb-4460-8592-6e9a51c52323 · outbound
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Training Agents using Upside-Down Reinforcement Learning
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcb313a7-679a-48e0-a15c-31d27690d890 · outbound
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Unresolved cited work
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b281e5ca-9f1b-4672-afe5-f200cffd2c3f · outbound
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 221a88ec-2bdf-4443-ac4e-ba250853f5da · outbound
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58bcb1bc-0d35-4f65-8783-f27f9934e8e3 · outbound
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success The Llama 3 Herd of Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.