Pith. sign in

Paper Citation Record · LEDGER

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.11153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11153 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:22.868400Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1868bd64-3169-486a-ac63-7bef0f6ee600 · outbound

This paper cites Reinforcement learning in game industry—review, prospects and challenges,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Reinforcement learning in game industry—review, prospects and challenges,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.706877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.596392Z digest=sha256:e2a85a1defd49e6dc6d5da7dea7a6a9768a2c52d18601e3fc167a255a07b56ef

Observation 5b36d62b-6933-4947-ad82-0c273075a833 · outbound

This paper cites Artificial intelligence, machine learning and deep learning in advanced robotics, a review,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Artificial intelligence, machine learning and deep learning in advanced robotics, a review,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.601706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.601706Z digest=sha256:6e38a1e6a30b0d63437bc7c5c58571f3c540857158cba32d8d689c9306a5af6c

Observation afb15a91-ee65-49f1-81ff-92498e50dd95 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Playing Atari with Deep Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.606744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.606744Z digest=sha256:81de385be75f43d933df3e093ff4e117969a89a5881e5159e4567a51ff80b421

Observation f0d16ee2-0918-4a41-9bc0-b48eeb8b0c50 · outbound

This paper cites Weakly coupled deep q-networks,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Weakly coupled deep q-networks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.682572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.611371Z digest=sha256:0629c399544ba573eaf6b5fe4769444cebec54adb90c7a3b529089cf183098d5

Observation 2ca25150-c058-45a0-950c-fe3891c86abf · outbound

This paper cites Tuning apex dqn: A reinforcement learning based deep q-network algorithm,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Tuning apex dqn: A reinforcement learning based deep q-network algorithm,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.668685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.616440Z digest=sha256:93c526dd6e61aa214261c0c92ef1a482120ce77e116dca93092231471af0e20a

Observation 0e111d3c-b700-45ee-8af1-df7aa143c6ae · outbound

This paper cites Safe reinforcement learning via shielding under partial observability,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Safe reinforcement learning via shielding under partial observability,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.654874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.620781Z digest=sha256:37f7b76f37ae918f0e0d5c8a602cf933f4cfa30ba20d770873e8b3dc7b12d6a5

Observation c0c88722-5cad-4be5-be71-14d9dad8da2e · outbound

This paper cites Integrated task and motion planning for safe legged navigation in partially observable environments,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Integrated task and motion planning for safe legged navigation in partially observable environments,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.641601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.625634Z digest=sha256:5509ccae28942343d61eff1e71215566afc2fbf661708f66663300f058682987

Observation 2a3f48f8-7f06-47b2-a650-a797f9fb2ba3 · outbound

This paper cites Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.629718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.629718Z digest=sha256:fd868c341e595f24946647c3bcaaa188135a788c49faa74874ba83d9dfb45227

Observation 95640671-922e-4ad1-b63e-a3f8ee8426f4 · outbound

This paper cites A definition of continual reinforcement learning,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes A definition of continual reinforcement learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.633844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.633844Z digest=sha256:0bb9d6b3d9d430f8e74a4c5451818cdcc063db0181fa70500731372e49d0b2d7

Observation 0d9b6466-e0f7-4ed8-a408-370001b6650b · outbound

This paper cites Deep reinforcement learning unleashing the power of ai in decision-making,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep reinforcement learning unleashing the power of ai in decision-making,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.610494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.638178Z digest=sha256:f647b287dae7d0b40f7217725e504c855ed6f418553c695c951f05ac775ff4af

Observation b554e27a-280c-45c4-89ba-5f6b8de6cda1 · outbound

This paper cites Exploration in deep reinforcement learning: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Exploration in deep reinforcement learning: A survey,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.642683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.642683Z digest=sha256:5b50558bc0c573afffb1a4ffd293789c1f73f1cb3cf699129816f6e2e1ecfd9f

Observation ca5fec4b-57dc-47e4-bf84-fa36857fee63 · outbound

This paper cites Memory gym: Partially observable challenges to memory-based agents,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Memory gym: Partially observable challenges to memory-based agents,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.646843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.646843Z digest=sha256:92a0a8c38dd74b9242ff335f6eca4f6b7cbca5cdd6ef4b133bc75385dd27a283

Observation 0657da13-58f0-4684-8e76-e35f23d2df8e · outbound

This paper cites Deep rein- forcement learning: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep rein- forcement learning: A survey,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.579178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.655825Z digest=sha256:da92e61e27ecb12583feb95582a120e0ccfe7546c096608fa15a51fb6bb29374

Observation 5f247310-6da1-4525-b271-3e4b7100aa90 · outbound

This paper cites Partially Observable Markov Decision Processes (POMDPs) and Robotics.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Partially Observable Markov Decision Processes (POMDPs) and Robotics

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.659923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.659923Z digest=sha256:a34e43dd3d4dc490df9aa1a00564b15d81a202feb56dcadf7823016fc710304b

Observation 179ec413-502f-4197-b9b0-e231b42dee0c · outbound

This paper cites Recurrent neural networks,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Recurrent neural networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.565191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.664501Z digest=sha256:995b2f143aecbf5c0ea41205eb2fa7a3bea94ac43dc36dfbfeea493b13c04de0

Observation 5508d405-6ba2-44ff-a703-bd4ae0ce2130 · outbound

This paper cites Long short-term memory,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Long short-term memory,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.669081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.669081Z digest=sha256:a9f206c388e73c9480b48bd8674abd6828de7f9bae09ed1bb9805bb8c3a8aa5d

Observation 6c3462ee-db52-4081-a500-2e68d82f52db · outbound

This paper cites Gate-variants of gated recurrent unit (gru) neural networks,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Gate-variants of gated recurrent unit (gru) neural networks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.543519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.673326Z digest=sha256:4cad2be1c4f58f248e8fb8ba06f6d8eacbdb4e0abab862603559bfa01e28b712

Observation 70c446ac-4926-4f2c-9ad1-327e144ea182 · outbound

This paper cites Recurrent prediction model for partially observable mdps,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Recurrent prediction model for partially observable mdps,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.529450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.677802Z digest=sha256:a3567dca56e73ce2af2d4d0fbdb4dc6c90468e4bf4eee89e46a6782eff621a57

Observation ef5001db-3314-4c8c-b216-d89e3637eeda · outbound

This paper cites Visualizing transformers for nlp: a brief survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Visualizing transformers for nlp: a brief survey,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.514757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.682088Z digest=sha256:e0dfbdebacfb7dfa3ce25a75b51a1d2f6855596eb645dffd052018aaef064052

Observation c37841f9-3099-4022-80b6-56f081478351 · outbound

This paper cites Transformers in vision: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformers in vision: A survey,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.686804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.686804Z digest=sha256:fc7ce781b0ea65d34ccddf3473bbff3282359b0c7949490d2def3f2a6e956690

Observation cc20339b-6200-4c54-9e2d-9b4618c91561 · outbound

This paper cites Windows deep transformer q-networks: an extended variance reduction architecture for partially observable reinforcement learning,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Windows deep transformer q-networks: an extended variance reduction architecture for partially observable reinforcement learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.490794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.691129Z digest=sha256:8e838c3cf482518db03549a9bca3394b38d6d0480681ef4698e82f720d60fea2

Observation c391dc60-aab2-49e3-89ab-0fe39343235a · outbound

This paper cites Deep Transformer Q-Networks for Partially Observable Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep Transformer Q-Networks for Partially Observable Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.695510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.695510Z digest=sha256:4010bc47f55e7cdfbc020cb4c34feca409beaa94427adb03996dc9f0da00ef53

Observation 3888cc66-5364-4345-93f6-9de049e8a60e · outbound

This paper cites Deep Recurrent Q-Learning for Partially Observable MDPs.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep Recurrent Q-Learning for Partially Observable MDPs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.701126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.701126Z digest=sha256:da17f763ce9a019f8799942c21761f89dc673391ab2ef8cde25414b3c4290766

Observation 72a74ba4-3b51-48b8-be89-0bedcb1171fc · outbound

This paper cites On improving deep reinforcement learning for pomdps,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes On improving deep reinforcement learning for pomdps,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.476678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.705893Z digest=sha256:3def19b59c8655b138348d45a6e74adcf330893eefaa557357f7092e5a3fe3ec

Observation 46461b7f-00a1-4fad-b3cd-85734e91a628 · outbound

This paper cites Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.714601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.714601Z digest=sha256:44057cdfc6e1ff64694c7ad792c945c174c29c60a43af9e3f77afcda9bd693c9

Observation bf1657e5-c0f3-4a7c-b00d-852e039ba788 · outbound

This paper cites Deep reinforcement learning with bidirectional recurrent neural networks for dynamic spectrum access,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep reinforcement learning with bidirectional recurrent neural networks for dynamic spectrum access,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.463274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.719144Z digest=sha256:627c51c8c061783817d855fdd49a97d3200223690b5a0d4127c0d21b68b53d36

Observation ae84f17b-566d-4557-b30f-08976440d674 · outbound

This paper cites On transforming reinforcement learning with transformers: The development trajectory,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes On transforming reinforcement learning with transformers: The development trajectory,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.450134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.723470Z digest=sha256:0f61c5e72c1dd538e7ce49a548a91035f528f2a2314b70340266fb93c427a47e

Observation a562b3df-f1db-4c05-8ff3-3066d67b3fb7 · outbound

This paper cites Transformer in reinforcement learning for decision-making: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformer in reinforcement learning for decision-making: A survey,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.437267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.727574Z digest=sha256:6e09ac8dd96ea194c95af42d276b16e953640725b26be4d4a0eeeafc5b2c6d51

Observation bc4f5edf-f3a8-418b-9dd3-a8428fb3d894 · outbound

This paper cites Attention Is All You Need.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Attention Is All You Need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.731661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.731661Z digest=sha256:615bc09d86eb0f86c6643b7525b3ec2aa0ef0b428165d5eff3981dc0a1972a1c

Observation dc739bf0-97d2-4002-b7ab-733b2d16428b · outbound

This paper cites Decision Transformer: Reinforcement Learning via Sequence Modeling.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Decision Transformer: Reinforcement Learning via Sequence Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.735633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.735633Z digest=sha256:e9f02fc5b4d9432dafba3f4f7091c75f7ce61e50e3ff06bb81280c24cb032f4b

Observation 4cb33299-03b4-4e57-a132-5d852f3352db · outbound

This paper cites Deep Attention Recurrent Q-Network.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep Attention Recurrent Q-Network

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:01:23.170540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.739898Z digest=sha256:f849377458ba7fb36ed2c75f5de79b55698951ea2cfe1d149c6b7cb933afd90a

Observation bdde2bf2-25da-4f13-825a-1a06159bdf88 · outbound

This paper cites Towards Interpretable Reinforcement Learning Using Attention Augmented Agents.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Towards Interpretable Reinforcement Learning Using Attention Augmented Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.744141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.744141Z digest=sha256:924bc56cd922264e62cf630e8fecbc3b0b1f5c30af8d547fd3fc6bdeccef4498

Observation 15e4264b-f93a-4ea3-9c95-268ffa9b797b · outbound

This paper cites Efficient Transformers in Reinforcement Learning using Actor-Learner Distillation.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Efficient Transformers in Reinforcement Learning using Actor-Learner Distillation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.748705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.748705Z digest=sha256:6ad0184529d0f79cedd57d735b91c3c85df8fcd75608d44c844ec56c4892a146

Observation c8c00b89-1879-40dc-97da-39507316d0b6 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.753113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.753113Z digest=sha256:b37511f8d6c6b418ace7c4f80fe67fdf9e2041ceb9c0474fd24fdd4985866fa7

Observation 95495ef3-5df6-49fb-82f3-4c1d5f2e8526 · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Offline reinforcement learning as one big sequence modeling problem,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.422707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.757384Z digest=sha256:7aa663447d043a29a825df3000b702e0f01b9e9411c39f5beac458810432a9f0

Observation cf540214-cd83-474a-abbf-f3ea2833c154 · outbound

This paper cites Online Decision Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Online Decision Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.761745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.761745Z digest=sha256:0c6a6feb2305fd6ae0c6a998d968e6bf2579c8830b0a256515a15afd9e64c701

Observation 1d2e5ca5-fe82-49a4-ab95-acd8ef48894f · outbound

This paper cites Structured State Space Models for In-Context Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Structured State Space Models for In-Context Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.766001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.766001Z digest=sha256:259bc7f8105344a21e14a9a95875f3d90dfe689f82d0636c1593b9f9aebb6b64

Observation 9266ec07-5a79-46c4-9b8d-d0c3732ee6af · outbound

This paper cites Mastering Memory Tasks with World Models.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Mastering Memory Tasks with World Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.770341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.770341Z digest=sha256:81d959a2d0bdfb85e6f2a147e5ab0182b2349d28b208eb726bf9934d5572faa4

Observation 53be7564-2bce-4acd-a07d-f7e7840a8e45 · outbound

This paper cites Mastering atari with discrete world models,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Mastering atari with discrete world models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.408184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.774501Z digest=sha256:0224c8ba5211fbb8e4304cdb6fc274f6e5e0588240e28849ae990e4dcab25fe6

Observation 95412e92-4c5f-4e87-961d-c47b4c45e61f · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.783296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.783296Z digest=sha256:2166f7925cccd7930ecaabe6696cbb1df513772567c32dd6f9a01c67d06b96dc

Observation 1fb5fdc0-918f-403d-803b-cb9169442c2c · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes RWKV: Reinventing RNNs for the Transformer Era

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.787610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.787610Z digest=sha256:d8f60b04df564f3f5cfb76a43fae84d7b8df170dcd03378d0e3c1bf4b0272b67

Observation c4b344d7-2a02-4d2c-8cbe-5b654b71387a · outbound

This paper cites Resurrecting Recurrent Neural Networks for Long Sequences.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Resurrecting Recurrent Neural Networks for Long Sequences

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.792509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.792509Z digest=sha256:61c93ebfb9f1792ed33fad672c394fb9fb9994386c30073780032f9c02c469df

Observation 16c939f0-da6f-4dd8-88c1-87f85629d068 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Efficiently modeling long sequences with structured state spaces,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.394090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.797991Z digest=sha256:9627f6a4d0a4e465407f108b114b74ea320fe5df4c6433090248eea38b0e3c23

Observation fec57d69-189a-4b98-b57d-14c3d8f1049d · outbound

This paper cites Simplified State Space Layers for Sequence Modeling.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Simplified State Space Layers for Sequence Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.806620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.806620Z digest=sha256:c291aa61d59e9facb0701bff8546cc10a2bebefec0b1968f47348e9989213f2a

Observation 28bbaeba-5c66-4e6f-9f28-eae008db609f · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.811361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.811361Z digest=sha256:4804a133a0d061d953edd420f8eedfac77d131e97f4518bf956c7566409f9870

Observation 30d1836c-1d8b-4175-bf0b-aa61e69e536f · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Efficiently Modeling Long Sequences with Structured State Spaces

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.802215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.802215Z digest=sha256:4e4847ff33882cc33840d2e248d9be4cdad9ab246ba382d077a18b8f0da8cf15

Observation 97ea078e-db2f-4e07-a3f1-83beccc4c463 · outbound

This paper cites Longformer: The Long-Document Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Longformer: The Long-Document Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.820257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.820257Z digest=sha256:c93014d701eda1dd4ab9b5490f3b6e4d2a78558d8187f8c9ac89df1343659491

Observation 83ca39f0-2267-4829-b533-4e45e76c6fca · outbound

This paper cites Transformer-XL: Attentive language models beyond a fixed-length context,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformer-XL: Attentive language models beyond a fixed-length context,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.380533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.824663Z digest=sha256:bda76007d7af91c818397f53e9337b65040da9acc6e9e4f634454613f93a915b

Observation 1e28d407-380d-484e-9aca-ba9d95119dfe · outbound

This paper cites Stabilizing Transformers for Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Stabilizing Transformers for Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.815814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.815814Z digest=sha256:842b8c21a8aff24d51c149018f3950f057bac9b2e2e8db898b0b3a35dfa90718

Observation bfcfe6a5-0e3c-4837-889d-77b8519d1689 · outbound

This paper cites Recurrent Memory Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Recurrent Memory Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.833253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.833253Z digest=sha256:866b9c482b837677c4f3dfbb6ca6d9158f67ecb004f4692dad60b89a4477bd03

Observation abedf9b8-7a34-40b0-b995-1881e58d5461 · outbound

This paper cites Reformer: The Efficient Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Reformer: The Efficient Transformer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.837715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.837715Z digest=sha256:4ec3d792ed786a49c12ab50b527e71d481877c482419895c927752b6ce73a5d6

Observation bdddb050-6020-4a30-b036-b7c46fca5a9f · outbound

This paper cites Linear transformers are secretly fast weight programmers,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Linear transformers are secretly fast weight programmers,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.366622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.829200Z digest=sha256:6b9ec8eb78c827ae59de74b92caa930833a745c028b9d15f2ee4a04d7ddac349

Observation c186d155-232d-4e77-ac6e-696a6e491d05 · outbound

This paper cites Investigating the Role of Feed-Forward Networks in Transformers Using Parallel Attention and Feed-Forward Net Design.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Investigating the Role of Feed-Forward Networks in Transformers Using Parallel Attention and Feed-Forward Net Design

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.846703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.846703Z digest=sha256:6b5cfa80fc025e6fac703239920d79777c694b8fa035db46bd7b9834636c82f4

Observation 2057806e-73f7-49fb-943d-4dc4d76b0a5e · outbound

This paper cites Rethinking transformers in solving pomdps,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Rethinking transformers in solving pomdps,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.351037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.851223Z digest=sha256:cb86ebef98ae39db6c71d89c205f3f033485d812729651b392a8b40e20dcf936

Observation a7a99566-99ec-4cd8-8365-799994a5d2cb · outbound

This paper cites Rethinking Attention with Performers.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Rethinking Attention with Performers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.842169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.842169Z digest=sha256:8225281abd033814dcccbd7a6021d12405d2d885737b65c5a9d7c5a86bfe177b

Observation 0d42377e-36e6-426d-a376-46d3fb013090 · outbound

This paper cites Pomdp robot domains,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Pomdp robot domains,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.322069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.859622Z digest=sha256:2f0177bd410b721dd931bffbb404226a7b5b107f15c72b98d3614eaa926f71cf

Observation 85361ea8-fcb7-41ea-82e3-c8634bc66ce6 · outbound

This paper cites Learning policies for partially observable environments: Scaling up,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Learning policies for partially observable environments: Scaling up,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.307086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.864054Z digest=sha256:20ef934404cae54a17a80de1815aacde61b6279b04daa0cdb67638000548ec2c

Observation 1f01bb95-6aeb-41f4-8076-af9d547f5e1c · outbound

This paper cites gym-gridverse: Gridworld domains for fully and partially observable reinforcement learning,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes gym-gridverse: Gridworld domains for fully and partially observable reinforcement learning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.337172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.855486Z digest=sha256:2095d8a8a23e08e0ec365084e1e26dfff673c38c36e5df59fe88000f3a0906f8

Observation 083fdfd6-0e13-46c0-a7d8-14ce4a391f09 · outbound

This paper cites Solving large pomdps using real time dynamic programming,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Solving large pomdps using real time dynamic programming,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.292669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:01:22.868400Z digest=sha256:353b6cdd8601c4f276a07abee1319bd98e94e78614b93f8f6a248d36d25a1a28

Observation 1933d1b8-cf1c-4040-9f44-594c111672b1 · outbound

This paper cites On Improving Deep Reinforcement Learning for POMDPs.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes On Improving Deep Reinforcement Learning for POMDPs

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.710231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.710231Z digest=sha256:b5b62c6fd49d7dc1babfd44e476e3f1d3dd800bd381dff7a477a0e8bcafd97e2

Observation 926b6d1d-17ff-4c89-b562-24cb62e6ba2a · outbound

This paper cites Mastering Atari with Discrete World Models.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Mastering Atari with Discrete World Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.778724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.778724Z digest=sha256:ec94c9152425e0df040d2b5d0fb27427d540eee1e24b0fe081e0a86e424bcfa0

Pith citing papers

No inbound Pith citation observations are available.