Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:1912.02875.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.02875 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:56:16.028239Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.082706Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c848b185-a977-43b6-96b0-762c8ebfcd70 · inbound

Is Conditional Generative Modeling all you need for Decision-Making? cites this paper.

Is Conditional Generative Modeling all you need for Decision-Making? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 243

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:35:10.864352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T15:35:10.593969Z digest=sha256:d605d4367e088c279ac748934b3d0ecd4efd830fba89699d535965ae8a6a91e6

Observation 935963c6-4d05-4f6a-9792-79c1ea9c2621 · inbound

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework cites this paper.

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 151

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:43:19.168451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T03:43:18.632292Z digest=sha256:0dcd8a01ed329788822aa168d683e6834ff1db149eebf37978b1d5b25b1eb25d

Observation 32f142d3-c417-4019-bb8a-747a4a27778c · inbound

Upside-Down Reinforcement Learning for More Interpretable Optimal Control cites this paper.

Upside-Down Reinforcement Learning for More Interpretable Optimal Control Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T18:35:59.762452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:35:59.762452Z digest=sha256:a04d39c2023ee04fa31a3a754f722c21022b5ac0c46c3b67cc9f530558b71a26

Observation 6f1de91c-4b5c-41ac-8752-b368bbd48aa7 · inbound

Heuristically Adaptive Diffusion-Model Evolutionary Strategy cites this paper.

Heuristically Adaptive Diffusion-Model Evolutionary Strategy Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T16:32:34.663693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:32:34.663693Z digest=sha256:1510957956565c75eb701ac1805efc65092ad3ed857508f8cae5c2ab542f028a

Observation 697d7ca2-eb1d-4c4d-a695-be48221c09c2 · inbound

GenPlan: Generative Sequence Models as Adaptive Planners cites this paper.

GenPlan: Generative Sequence Models as Adaptive Planners Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:51:01.263833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:51:01.263833Z digest=sha256:248b1b10b4c89f589317f4135df22fbf452c6865f239d94d467c00e46bd7a6cf

Observation c31d09c1-597f-45b4-a07a-4debc3ede1aa · inbound

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents cites this paper.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.198736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.198736Z digest=sha256:6779c1a95c00cbfe14e810d400d6343750dd98a6af64fcee16984f9063f91c0c

Observation 04f71c72-c48c-497c-9951-b5e7e7415a13 · inbound

Are Expressive Models Truly Necessary for Offline RL? cites this paper.

Are Expressive Models Truly Necessary for Offline RL? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:12:40.522227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:12:40.522227Z digest=sha256:6bd8cf4912483f3f0cfa11d8c4339c5f5780123bde243942adcc5aea3dd47076

Observation f9afbe39-e9df-4622-9ba3-76afe04cc0c1 · inbound

Latent Diffusion Planning for Imitation Learning cites this paper.

Latent Diffusion Planning for Imitation Learning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:56:16.028239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:56:16.028239Z digest=sha256:ae8a98a01defcf54c408400e5089e8c922854b9a3690fd1d7dc46fba73f4e54b

Observation ba071985-3d19-4433-9d0c-900d35d4913c · inbound

Directly Forecasting Belief for Reinforcement Learning with Delays cites this paper.

Directly Forecasting Belief for Reinforcement Learning with Delays Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-16T04:44:49.120695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:44:49.120695Z digest=sha256:0eee0918e1d20e664161383248f12a48d4b5dea691b3185ae49c4e255068bc4a

Observation c90fb6fe-a89a-4064-8914-eb02567d1ab3 · inbound

Beyond the Known: Decision Making with Counterfactual Reasoning Decision Transformer cites this paper.

Beyond the Known: Decision Making with Counterfactual Reasoning Decision Transformer Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:57.492172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:57.492172Z digest=sha256:2c7b3b9c82c953b7c1512d82f4c10017e2b4b801ce375c5198a006ea93f9aa91

Observation 28bbaeba-5c66-4e6f-9f28-eae008db609f · inbound

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes cites this paper.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.811361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.811361Z digest=sha256:73c63e213468b785cffaff53ae627d661a3c238015697e9df9aca86bfb32ae18

Observation 90cddf7c-eeb2-4a7d-bc5b-4a7371aa3876 · inbound

A Provable Approach for End-to-End Safe Reinforcement Learning cites this paper.

A Provable Approach for End-to-End Safe Reinforcement Learning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.086262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.086262Z digest=sha256:7612545983fa5f054ee923032d758202f96233579d99b3f2afc833d3fa5589cb

Observation 8788ce0a-e4a4-4fde-8217-1fade5dd987c · inbound

BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL cites this paper.

BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:29.406729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:30:29.406729Z digest=sha256:9b115ff96b2d55d265acd6ad02bb434644cf000d7c08674502a4223abc9b9f9f

Observation 5717991e-91ec-4bfa-ba3a-df3a467f88cf · inbound

How to Provably Improve Return Conditioned Supervised Learning? cites this paper.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:07.756599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:07.756599Z digest=sha256:499e06454c0a55b857405844ef3782b76b57b477d0d0f1c80c8589d1cf4a94dc

Observation 69a88021-63dc-4679-9ef6-aef353c8e3b9 · inbound

Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning cites this paper.

Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:17:14.120398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T09:15:45.511104Z digest=sha256:0935fdd1b0d6c8ec33c60d8a8eda467bc237ae5adfc0a6df6e13a00d3d26e7ee

Observation 38fcec82-510a-4a84-9c06-39722501636c · inbound

Zero-Shot Reinforcement Learning Under Partial Observability cites this paper.

Zero-Shot Reinforcement Learning Under Partial Observability Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.860626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.860626Z digest=sha256:ac9be83cbfef3bd4e3de5b835461483180ea71e848e2b50204f97cfe6c9f93f9

Observation 28f66d78-9c97-40d7-8194-3a772848df4f · inbound

Single-pass Adaptive Image Tokenization for Minimum Program Search cites this paper.

Single-pass Adaptive Image Tokenization for Minimum Program Search Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:32.548703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:35:32.548703Z digest=sha256:200c6b332223b49b4574932937a0f354aeacec8eb87e439ef177a2f21a820c6a

Observation 159f7b58-4c8a-4f70-9822-cd365b443697 · inbound

Behavioral Exploration: Learning to Explore via In-Context Adaptation cites this paper.

Behavioral Exploration: Learning to Explore via In-Context Adaptation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:46.469662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:15:46.469662Z digest=sha256:9a3771a352d3d109f5cf0e7eccf74ed1d2cb3f72f0db3bdac49dee40e2292398

Observation 1e036314-dc6f-45b0-adea-608e157c7c2c · inbound

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration cites this paper.

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:14.428722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:14.428722Z digest=sha256:29e07ac65bd8ab5a0d801096466aa1bf65f3c080caf3acbc9622621c41a6777f

Observation 7542ae43-ba7b-45cf-b733-1f13a4991a15 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.408653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:2db070530244cac0dd8f33ca7297d3910b076907baf79b99fc9347e6c42b75f2

Observation 221a88ec-2bdf-4443-ac4e-ba250853f5da · inbound

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success cites this paper.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.352724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.352724Z digest=sha256:0b9166573d93b2738b9ddd03563c2dd2a13235e87d70be24d57c1421a1347913

Observation cf5bc580-a521-4071-ab1e-75788d2bc868 · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:58.446232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:02e446ad2bed0cdab8506f093b5b1bf4e9b29d19eab5c628cb96f7606bbb632d

Observation cc5ca21d-ace3-47b1-8197-7cc9f2486d35 · inbound

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning cites this paper.

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.084211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T00:19:49.473294Z digest=sha256:0560c5c742fd270115117b65919e3238f7b7f44dc88d159e183eb8294eccee83

Observation 1748dc48-5e81-4e46-8503-0c24aed948ee · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:55:41.995959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T04:58:28.971536Z digest=sha256:e9b6665b3f90444bd9a8a1124ec4bd77847665dea6d393826c9abe14c97d95a3

Observation f190c8be-e006-41af-abc7-8f35fd549430 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T16:55:18.028851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:55:18.028851Z digest=sha256:bc908503baee5c38c5f8565a96ac0d398972a52383df35dec657c8801d74b110

Observation fa1c7e24-98bb-4437-b66c-88cf4272c7f4 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T02:32:20.110034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:32:20.110034Z digest=sha256:01e72c9db7589417c3143ec94ae90e248b6eebc4c68dc71176fb2b6d3a74477f

Observation 5b479fce-53c0-4401-a849-44e8ea6307e2 · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 175

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:14.175760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:14.175760Z digest=sha256:97926675c943a847a76c692beaed2df5e47d7d46d33e223c97272697d83d030b

Observation 23d15564-9343-4d9e-9703-70d5d4ee1517 · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 268

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:32.194056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:32.194056Z digest=sha256:7a112992af2e0194c72ae2ca182b906020399ddbb88540d72ff090fee8e2cebc