Pith. sign in

Paper Citation Record · LEDGER

Model-Based Reinforcement Learning under Random Observation Delays

As of 6 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2509.20869.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.20869 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T14:35:04.201252Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact10
  • verified fuzzy29
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 406c6330-e16c-4254-bc3d-d4e3e05296a3 · outbound

This paper cites write newline.

Model-Based Reinforcement Learning under Random Observation Delays write newline

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.733081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:f45763f5c621e000e4c37868c0639ec98809471e9762af8cea4cebf17df8076f

Observation 72534a24-05bf-4b51-8beb-8a1f3a953b5f · outbound

This paper cites A cerebellar-based solution to the nondeterministic time delay problem in robotic control.

Model-Based Reinforcement Learning under Random Observation Delays A cerebellar-based solution to the nondeterministic time delay problem in robotic control

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.767000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:4aba37f62e555d3bf9c0553242d19f29d6bc36c5dda5e3df9f6232279ac25e29

Observation e01764b1-1e9a-4854-9f58-fb98943a31bb · outbound

This paper cites Closed-loop control with delayed information.

Model-Based Reinforcement Learning under Random Observation Delays Closed-loop control with delayed information

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.762654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:02466e743e37d4c4289e15311aa3020ae726ca5df9b694b17b7d03223b96dd16

Observation ef87f7b6-4d46-48e4-b0af-fdcf08c586fb · outbound

This paper cites Update with out-of-sequence measurements in tracking: exact solution.

Model-Based Reinforcement Learning under Random Observation Delays Update with out-of-sequence measurements in tracking: exact solution

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.758344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:3af2d8cea9f20ade13e5a1b76b4b09c5cbd84e377a9da9225daf18b51a23503d

Observation 7bc3da22-b67c-43f6-8632-964cb02ecbcb · outbound

This paper cites Reinforcement learning with random delays.

Model-Based Reinforcement Learning under Random Observation Delays Reinforcement learning with random delays

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.654857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:eb289356197294860e1469bf3c865aad1dccc8ed380322828a1035e72d02c8fa

Observation b381dd62-2fac-4a1d-9bd0-317180f5b5bc · outbound

This paper cites Bayesian filtering: From kalman filters to particle filters, and beyond.

Model-Based Reinforcement Learning under Random Observation Delays Bayesian filtering: From kalman filters to particle filters, and beyond

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.754284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:3b98743a8a7a837fbc79b35ffa7362897e94d5644ccc60edf5833407b2dad9ff

Observation fa3ed8f9-681a-4d5f-b352-ee0340607649 · outbound

This paper cites Acting in Delayed Environments with Non-Stationary Markov Policies.

Model-Based Reinforcement Learning under Random Observation Delays Acting in Delayed Environments with Non-Stationary Markov Policies

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.716906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:f03cc01fb5d8461381c5e457dea4f99a7c86880bc061000d8bf0cb2b1de8401c

Observation 8be2a045-383f-418a-b260-d0a7cfb37400 · outbound

This paper cites Communication delay in uav missions: A controller gain analysis to improve flight stability.

Model-Based Reinforcement Learning under Random Observation Delays Communication delay in uav missions: A controller gain analysis to improve flight stability

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.750778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:c78b7ab867b26a1388319d8aa8ce67048b1f8d5b9aa784febdae309d09b5aa78

Observation 4e40442a-2c12-40dc-b7c9-dada1b99b725 · outbound

This paper cites Recurrent world models facilitate policy evolution.

Model-Based Reinforcement Learning under Random Observation Delays Recurrent world models facilitate policy evolution

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.747267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:c25b05a293fc0a22c85b4562148ab0cf86d8ce55aafd23df872d3465ca723f27

Observation 7bbcee7c-c512-4c9e-b4d7-ff8d076b055b · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Model-Based Reinforcement Learning under Random Observation Delays Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.743794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:c6d5aa47e6ef6158dfe30605a640fe4692e98efe872cbed7d5a7b0a4830f8e3f

Observation fa203662-7aeb-4f51-b27d-28624a6e409c · outbound

This paper cites Learning latent dynamics for planning from pixels.

Model-Based Reinforcement Learning under Random Observation Delays Learning latent dynamics for planning from pixels

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.740137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:1d0e782e8220c57fd295e604f02fedddb97fe13d2c11f7ec07bd879b756800e3

Observation 0b8b7500-655a-40b8-a17b-88c3cde4c777 · outbound

This paper cites Mastering Atari with Discrete World Models.

Model-Based Reinforcement Learning under Random Observation Delays Mastering Atari with Discrete World Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:36:28.665464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:b671833ec26e9b467d18f808bc27ddd47cd47f512a87ce0b8ea90872bfad316e

Observation c1e3cbb5-a00d-4b7b-9008-c9d334e38497 · outbound

This paper cites Mastering Diverse Domains through World Models.

Model-Based Reinforcement Learning under Random Observation Delays Mastering Diverse Domains through World Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:36:28.653506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:a56232cedd8abeff28366de00e40a61883314e3d0510958e797cf134e380bde9

Observation 56d1e66a-7a89-4d5f-b2e1-fd29dd86f723 · outbound

This paper cites Mastering diverse control tasks through world models.

Model-Based Reinforcement Learning under Random Observation Delays Mastering diverse control tasks through world models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.736667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:408d55a838b26a8f39c50e3b2c1017898ab72906c8f29ecce94cf0d86d7b565a

Observation b0bb3a76-b285-49d5-9221-c98dfe250bd5 · outbound

This paper cites Temporal Difference Learning for Model Predictive Control.

Model-Based Reinforcement Learning under Random Observation Delays Temporal Difference Learning for Model Predictive Control

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.673612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:a6852311c98a4ed8517e4128a4b6168a33a41e4b697bd54b8f14c00c72b371ad

Observation 2a7b6924-cb74-4409-8502-f3f51bf66c04 · outbound

This paper cites Deep variational reinforcement learning for pomdps.

Model-Based Reinforcement Learning under Random Observation Delays Deep variational reinforcement learning for pomdps

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.729376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:5a6cc3764d58319577c1f18e562928c245f96ae798400e1d84877124c2ee0594

Observation 84e91d7f-f57f-4b85-a5d4-51d70d9e83c3 · outbound

This paper cites When to trust your model: Model-based policy optimization.

Model-Based Reinforcement Learning under Random Observation Delays When to trust your model: Model-based policy optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.725374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:96ec7b667a3f9a70fb2cbd2b0a5d232e37c92b685181ba4b0073f4a9507c9abc

Observation d05b673c-0738-4295-b6e8-4cbec606981b · outbound

This paper cites Planning and acting in partially observable stochastic domains.

Model-Based Reinforcement Learning under Random Observation Delays Planning and acting in partially observable stochastic domains

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.719253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:11e467b80c4880d61cf1e9ee7989eb148d156cacbae8c1a507228a3238b878fe

Observation 483fdeb5-9396-4933-90e4-86a7b701633c · outbound

This paper cites Reinforcement Learning from Delayed Observations via World Models.

Model-Based Reinforcement Learning under Random Observation Delays Reinforcement Learning from Delayed Observations via World Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.711413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:996ffa26f1f625f342665228c94cf02ca96f50399d887b201e58e799b9c74575

Observation ed57e875-7fdb-474e-a745-8451d699d7a0 · outbound

This paper cites Markov decision processes with delays and asynchronous cost collection.

Model-Based Reinforcement Learning under Random Observation Delays Markov decision processes with delays and asynchronous cost collection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.715880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:19dc89df25939d7ad46ba1ca2612320338aeb020df85aa9046fadca8ac02fb4c

Observation 60b78f21-3d45-49f6-a1a0-919c31b3491b · outbound

This paper cites Belief projection-based reinforcement learning for environments with delayed feedback.

Model-Based Reinforcement Learning under Random Observation Delays Belief projection-based reinforcement learning for environments with delayed feedback

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.694161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:924ac2644203d3ed0a7c92bb6d4918d79d5d039ed9d6d5120427507c0c11b880

Observation 9c8707cc-49fd-46a9-9bd7-ebfc7d22173d · outbound

This paper cites A partially observable markov decision process with lagged information.

Model-Based Reinforcement Learning under Random Observation Delays A partially observable markov decision process with lagged information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.712594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:ccdc4e43fb33748b88f9310c2a2589af3bebb592771fc41fa340bb5ca2536c0f

Observation b9cdf566-2d78-41cc-b6b1-aed62d3fca77 · outbound

This paper cites Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model.

Model-Based Reinforcement Learning under Random Observation Delays Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.708711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:86eaea1895670da485befc0e475cd299a2acbcf23b678ccb61c66e89b07e50e9

Observation fe66291f-0461-4119-825e-f72ce8efdeba · outbound

This paper cites Learning a belief representation for delayed reinforcement learning.

Model-Based Reinforcement Learning under Random Observation Delays Learning a belief representation for delayed reinforcement learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.704463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:ae3ccb70e3a0a3ccdaf72a07a4b20bc25950e850aa502a55f75eb1d2055540eb

Observation ec3815fe-7798-4bc1-984f-757f1ab4db28 · outbound

This paper cites Delayed reinforcement learning by imitation.

Model-Based Reinforcement Learning under Random Observation Delays Delayed reinforcement learning by imitation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.659770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:a1e002c202c77c2dfc1bb8d8f58b41e83d2fb97c79f874a59c3bc7b899c8d15c

Observation bc6f3d72-ec48-4834-80fe-b06233236230 · outbound

This paper cites Particle filter recurrent neural networks.

Model-Based Reinforcement Learning under Random Observation Delays Particle filter recurrent neural networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.701154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:2344dbe4186bd617226e25f3939229085e4029566139834af0282c42f7eeca7a

Observation 2b80e932-e12c-4774-8eac-8d5495f6be05 · outbound

This paper cites Discriminative Particle Filter Reinforcement Learning for Complex Partial Observations.

Model-Based Reinforcement Learning under Random Observation Delays Discriminative Particle Filter Reinforcement Learning for Complex Partial Observations

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.686413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:e674f2bacbaf48fa34807a5d5323805e413fc451ad6e592ecf53766dd2e06c51

Observation 525a8490-4bcf-464d-89dd-97fffccd9277 · outbound

This paper cites Setting up a reinforcement learning task with a real-world robot.

Model-Based Reinforcement Learning under Random Observation Delays Setting up a reinforcement learning task with a real-world robot

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.697647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:718820e67150694d3395e3fd641b3e05a096d9f4c2f2d7e2b0e187d8a67b9683

Observation bd448093-7fd6-4800-8892-8cf16a760a76 · outbound

This paper cites Transformers are Sample-Efficient World Models.

Model-Based Reinforcement Learning under Random Observation Delays Transformers are Sample-Efficient World Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.699389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:2b4047e89a4feed0fe5e1840f127ea2a080cf225307bc4f8a7d46732475f98fd

Observation 9c266cee-3908-4181-af7d-56d628e7cb3a · outbound

This paper cites Control delay in reinforcement learning for real-time dynamic systems: A memoryless approach.

Model-Based Reinforcement Learning under Random Observation Delays Control delay in reinforcement learning for real-time dynamic systems: A memoryless approach

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.690664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:7e8e7276fce7e15e0b0ad2ed2628e7741157df39c50cbdc74e76791415ed3550

Observation e38cf315-f0b9-4ad4-a4d1-91e9bcfd6fcc · outbound

This paper cites Mujoco: A physics engine for model-based control.

Model-Based Reinforcement Learning under Random Observation Delays Mujoco: A physics engine for model-based control

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.686816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:fe6fe7ba8f8b56ffeec6f07c67dc051a3ef807b4342a2b9a4c8571bc4e8355ff

Observation 47ac839e-fa66-4f7c-9f3a-b135ebf7259a · outbound

This paper cites Tree Search-Based Policy Optimization under Stochastic Execution Delay.

Model-Based Reinforcement Learning under Random Observation Delays Tree Search-Based Policy Optimization under Stochastic Execution Delay

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.705749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:2640600aaafdd4a9d72b44257884183ff7324534cb4c014dfc717988ce9332b5

Observation 5860a670-62eb-4c3a-95b7-e0412edd8011 · outbound

This paper cites Planning and learning in environments with delayed feedback.

Model-Based Reinforcement Learning under Random Observation Delays Planning and learning in environments with delayed feedback

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.681878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:2b740ccc7d43870bd90d535b9fb1169b785aaca4f48b5d4c233d0d34df15da00

Observation 9bed5266-616d-44c8-aa8f-a47cb8b664df · outbound

This paper cites Addressing signal delay in deep reinforcement learning.

Model-Based Reinforcement Learning under Random Observation Delays Addressing signal delay in deep reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.677607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:6414747b253a8efa24b13a16ebb59d5130aec800ffb144070c4b4eba6cfbe448

Observation 1d332628-6ea9-4567-8e3d-35524f4700cd · outbound

This paper cites Variational delayed policy optimization.

Model-Based Reinforcement Learning under Random Observation Delays Variational delayed policy optimization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.673375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:3e28b7b9d6327bcd6873012773bcfbfe2035af1936c03cb42f4772df7683b0cf

Observation a142d5f8-0ee1-4fae-a18e-213b632459bf · outbound

This paper cites Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays.

Model-Based Reinforcement Learning under Random Observation Delays Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.680422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:9605ba082d55892e804f6f4ade31625214466a0492e6a0df8d12d4a32eb5bc0c

Observation b6fcf3af-f5b0-4693-8cf9-56201cff736d · outbound

This paper cites Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning.

Model-Based Reinforcement Learning under Random Observation Delays Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.692263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:1640ee0a4587af14be04be7b4148a1b37c1a7a74a518fb9bf688f870f768a15a

Observation 972d2d93-31b7-41dc-ab81-7e9785bf1c06 · outbound

This paper cites Storm: Efficient stochastic transformer based world models for reinforcement learning.

Model-Based Reinforcement Learning under Random Observation Delays Storm: Efficient stochastic transformer based world models for reinforcement learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.669270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:30e2dc8a7b1fc771facca9a5f3a5d8ace677adce0e2cb4ac36a48006c4381cbe

Observation 3ea9858e-dd34-4c34-8177-f92a7c2eaa79 · outbound

This paper cites @esa (Ref.

Model-Based Reinforcement Learning under Random Observation Delays @esa (Ref

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T14:36:29.664590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:bcd9df5e9d105f06a8a19b34c9ad3bef08d6164015838079e6c02508a681aea4

Observation fb5634f8-0d14-436a-a41e-bad2bc352143 · outbound

This paper cites an unresolved cited work.

Model-Based Reinforcement Learning under Random Observation Delays Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-18T14:36:29.649750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:66a5162fdbca99678e8b4fe87ee1b937aeaba113a012d4e78680219c0876caf7

Observation 4e0e15c9-4e47-4cb7-9b6e-aeddc37031be · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

Model-Based Reinforcement Learning under Random Observation Delays TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 41

Resolution
malformed identifier
local_arxiv, observed 2026-05-18T14:36:28.659427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:35:04.201252Z digest=sha256:4bd5824436bb49e3522ea1e5aa0a0285f4ff32f2bde81a4b791d8e32fe96d3fd

Pith citing papers

No inbound Pith citation observations are available.