Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:1912.02875.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.02875 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:30:29.406729Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.082706Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c848b185-a977-43b6-96b0-762c8ebfcd70 · inbound

Is Conditional Generative Modeling all you need for Decision-Making? cites this paper.

Is Conditional Generative Modeling all you need for Decision-Making? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 243

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:35:10.864352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T15:35:10.593969Z digest=sha256:e8206acd8a24e107fe4dee2b1184d95975304b356a1cd4dfb442e825f334381c

Observation 935963c6-4d05-4f6a-9792-79c1ea9c2621 · inbound

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework cites this paper.

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 151

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:43:19.168451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T03:43:18.632292Z digest=sha256:16e89e78bb7a60f0b1a27e96f96b947767b6b1a5a99c64bd7650b6eded069c2e

Observation 90cddf7c-eeb2-4a7d-bc5b-4a7371aa3876 · inbound

A Provable Approach for End-to-End Safe Reinforcement Learning cites this paper.

A Provable Approach for End-to-End Safe Reinforcement Learning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.086262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.086262Z digest=sha256:34032d3e8e04ad3036bf0047fb40d074fa44185c3f47a19a4073155b545ce6f1

Observation 8788ce0a-e4a4-4fde-8217-1fade5dd987c · inbound

BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL cites this paper.

BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:29.406729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:30:29.406729Z digest=sha256:47d416b267f45fd07b10ed81ef76fb29b52acd542cdc5d4b8c402012f403ded5

Observation 5717991e-91ec-4bfa-ba3a-df3a467f88cf · inbound

How to Provably Improve Return Conditioned Supervised Learning? cites this paper.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:07.756599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:07.756599Z digest=sha256:0365370e9093a58ea3d484cdb2858eba527893e8a049564051cbacffaf777d75

Observation 69a88021-63dc-4679-9ef6-aef353c8e3b9 · inbound

Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning cites this paper.

Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:17:14.120398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T09:15:45.511104Z digest=sha256:ce5467747e893ba0597e8182d05a63dd524c6665119be23da90f95522b267ba0

Observation 28f66d78-9c97-40d7-8194-3a772848df4f · inbound

Single-pass Adaptive Image Tokenization for Minimum Program Search cites this paper.

Single-pass Adaptive Image Tokenization for Minimum Program Search Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:32.548703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:35:32.548703Z digest=sha256:b4ee482360b3413db840dc09824f7ab5a6063811ad9a70578a41b55ee4e75c21

Observation 159f7b58-4c8a-4f70-9822-cd365b443697 · inbound

Behavioral Exploration: Learning to Explore via In-Context Adaptation cites this paper.

Behavioral Exploration: Learning to Explore via In-Context Adaptation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:46.469662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:15:46.469662Z digest=sha256:088fce677f70622fe04547b39ab1cb48e051c2017f7cdf07199d3b2fa31e7f54

Observation 1e036314-dc6f-45b0-adea-608e157c7c2c · inbound

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration cites this paper.

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:14.428722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:14.428722Z digest=sha256:a9a4f833e2ed3dfc5db2679ae823753d0061bb8ad68fc731fdad15d3f0bb32ce

Observation 7542ae43-ba7b-45cf-b733-1f13a4991a15 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.408653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:83b0f5f2773416e1186d5f97b7ea396462d02d638272564d2ca66b759dc83581

Observation 221a88ec-2bdf-4443-ac4e-ba250853f5da · inbound

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success cites this paper.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.352724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.352724Z digest=sha256:0d03a3c4dabe4c459c305b8513d2f5e716b3d03911e6082b769a4f3acae0d739

Observation cf5bc580-a521-4071-ab1e-75788d2bc868 · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:58.446232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:e89cea4e682ebd351205da54e4a5bfefa8c4ac153a2e112faa3d89e0292f2cb5

Observation cc5ca21d-ace3-47b1-8197-7cc9f2486d35 · inbound

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning cites this paper.

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.084211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:19:49.473294Z digest=sha256:54e8a1a66b546b1ed788d7a4c483678957536210ebe522bd68463ddc67a20850

Observation 1748dc48-5e81-4e46-8503-0c24aed948ee · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:55:41.995959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T04:58:28.971536Z digest=sha256:758d73e4b521103fb3bc41c7e5695d044ab7114dc7324eafaa0aa3a80132d179

Observation f190c8be-e006-41af-abc7-8f35fd549430 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T16:55:18.028851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:55:18.028851Z digest=sha256:e44bc0b83af56ad6238d8293ec9a5b3d232a77895ee0c495e8f5a6ccf73ab4b9

Observation fa1c7e24-98bb-4437-b66c-88cf4272c7f4 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T02:32:20.110034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:32:20.110034Z digest=sha256:2ca5db64174c149ad7c2e758828f297e48682d76b81ee860a82e1ec106c00258

Observation 5b479fce-53c0-4401-a849-44e8ea6307e2 · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 175

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:14.175760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:14.175760Z digest=sha256:32c2143ce1602e826c17a0960c813e7fce6c5eab50ab7f6c829996dec793499f

Observation 23d15564-9343-4d9e-9703-70d5d4ee1517 · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 268

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:32.194056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:32.194056Z digest=sha256:b3ef3845d3c0c55b61285ed230ecebcf39e0568286aca1b80ce27fc7e3c9e82b