Pith. sign in

Paper Citation Record · LEDGER

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes

As of 17 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.11153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11153 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:22.868400Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1868bd64-3169-486a-ac63-7bef0f6ee600 · outbound

This paper cites Reinforcement learning in game industry—review, prospects and challenges,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Reinforcement learning in game industry—review, prospects and challenges,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.706877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.596392Z digest=sha256:23de0444f70158b4108bdbe030f0f4dbc1a65e14857125dc7fa88ea5b61a54e2

Observation 5b36d62b-6933-4947-ad82-0c273075a833 · outbound

This paper cites Artificial intelligence, machine learning and deep learning in advanced robotics, a review,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Artificial intelligence, machine learning and deep learning in advanced robotics, a review,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.601706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.601706Z digest=sha256:e9c64523ddaefd043fc02477e7551d6d7ca2afb26edd187c6994cb3f564af0b0

Observation afb15a91-ee65-49f1-81ff-92498e50dd95 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Playing Atari with Deep Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.606744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.606744Z digest=sha256:ff7802f5a1cdddda958c92ca35175fd65e12b2eff5a4663343bd8f5b833ab08f

Observation f0d16ee2-0918-4a41-9bc0-b48eeb8b0c50 · outbound

This paper cites Weakly coupled deep q-networks,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Weakly coupled deep q-networks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.682572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.611371Z digest=sha256:37464ba2d2382cbbe420ec0182bbd6bf869bf093b58847c80e0bd79a078f0363

Observation 2ca25150-c058-45a0-950c-fe3891c86abf · outbound

This paper cites Tuning apex dqn: A reinforcement learning based deep q-network algorithm,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Tuning apex dqn: A reinforcement learning based deep q-network algorithm,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.668685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.616440Z digest=sha256:aa8f70d6423a0b2c20600fb1be88cfa440b78a1988926f43d59c9a1f85b9b6b0

Observation 0e111d3c-b700-45ee-8af1-df7aa143c6ae · outbound

This paper cites Safe reinforcement learning via shielding under partial observability,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Safe reinforcement learning via shielding under partial observability,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.654874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.620781Z digest=sha256:cd305665acf8ba7199d0f5039bf654ec6962e1286903753a9517fa39d420c9eb

Observation c0c88722-5cad-4be5-be71-14d9dad8da2e · outbound

This paper cites Integrated task and motion planning for safe legged navigation in partially observable environments,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Integrated task and motion planning for safe legged navigation in partially observable environments,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.641601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.625634Z digest=sha256:0450a994fa4200a1df4f30c8bd6fec00436dfd0af4d0857927122db4e7721c8f

Observation 2a3f48f8-7f06-47b2-a650-a797f9fb2ba3 · outbound

This paper cites Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.629718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.629718Z digest=sha256:f5311eb31475dfc5403f298f85da4c776a115cddd5b30ae6ff527375063f55f7

Observation 95640671-922e-4ad1-b63e-a3f8ee8426f4 · outbound

This paper cites A definition of continual reinforcement learning,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes A definition of continual reinforcement learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.633844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.633844Z digest=sha256:ae673997de3e7245d742a6fdf928a044d618e019fa3d078f90fec42c01d8ccc3

Observation 0d9b6466-e0f7-4ed8-a408-370001b6650b · outbound

This paper cites Deep reinforcement learning unleashing the power of ai in decision-making,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep reinforcement learning unleashing the power of ai in decision-making,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.610494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.638178Z digest=sha256:9aed6fba823e47931b561499e8f8520534ea1c1ca49fef696071227fad34fe0e

Observation b554e27a-280c-45c4-89ba-5f6b8de6cda1 · outbound

This paper cites Exploration in deep reinforcement learning: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Exploration in deep reinforcement learning: A survey,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.642683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.642683Z digest=sha256:484cf2f897e59c4f5970397ad6144fc89e2bb693db57e5a74c434ec48d8fbb4c

Observation ca5fec4b-57dc-47e4-bf84-fa36857fee63 · outbound

This paper cites Memory gym: Partially observable challenges to memory-based agents,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Memory gym: Partially observable challenges to memory-based agents,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.646843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.646843Z digest=sha256:e1f51c40ddfed66aad96f6578598b289a33c8712f054d1dbb4258e2d072fe556

Observation 0657da13-58f0-4684-8e76-e35f23d2df8e · outbound

This paper cites Deep rein- forcement learning: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep rein- forcement learning: A survey,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.579178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.655825Z digest=sha256:b998584db934a4b5a452d464842a82aecd55a5bf09017c9ebc078dde797353f6

Observation 5f247310-6da1-4525-b271-3e4b7100aa90 · outbound

This paper cites Partially Observable Markov Decision Processes (POMDPs) and Robotics.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Partially Observable Markov Decision Processes (POMDPs) and Robotics

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.659923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.659923Z digest=sha256:336457e03f984cb7aead37bf80302895cc226cfa170040c1f2dd27b277112719

Observation 179ec413-502f-4197-b9b0-e231b42dee0c · outbound

This paper cites Recurrent neural networks,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Recurrent neural networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.565191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.664501Z digest=sha256:2ada278786cf92bc77d588960f49d0d4600a5cc9422be007c30bef894a1ff57a

Observation 5508d405-6ba2-44ff-a703-bd4ae0ce2130 · outbound

This paper cites Long short-term memory,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Long short-term memory,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.669081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.669081Z digest=sha256:1d6c2c7734a98827ad17f97cec0e63bdcee2c364cd4914937cf84337cfbe475c

Observation 6c3462ee-db52-4081-a500-2e68d82f52db · outbound

This paper cites Gate-variants of gated recurrent unit (gru) neural networks,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Gate-variants of gated recurrent unit (gru) neural networks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.543519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.673326Z digest=sha256:1b4ef080936ccdfef74b70df4a79c816dee25b6c53bbd6bfd0d09539e41bad2b

Observation 70c446ac-4926-4f2c-9ad1-327e144ea182 · outbound

This paper cites Recurrent prediction model for partially observable mdps,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Recurrent prediction model for partially observable mdps,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.529450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.677802Z digest=sha256:5ac5adc3e465d678f2183530bc258bdc913a0879138e9118736ce0ab5fae9d91

Observation ef5001db-3314-4c8c-b216-d89e3637eeda · outbound

This paper cites Visualizing transformers for nlp: a brief survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Visualizing transformers for nlp: a brief survey,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.514757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.682088Z digest=sha256:56ceb3a10c1452a0c163bc62060189a066c0e600d225ec6ada949e675660348b

Observation c37841f9-3099-4022-80b6-56f081478351 · outbound

This paper cites Transformers in vision: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformers in vision: A survey,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.686804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.686804Z digest=sha256:5d68af04050ac8e2a16e90c4929c16005da3bd6410dd273a9c9201ee4cde7ed2

Observation cc20339b-6200-4c54-9e2d-9b4618c91561 · outbound

This paper cites Windows deep transformer q-networks: an extended variance reduction architecture for partially observable reinforcement learning,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Windows deep transformer q-networks: an extended variance reduction architecture for partially observable reinforcement learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.490794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.691129Z digest=sha256:247bafed368ca76508b51032cd0e5918a5a6a54c739ba828f7a8af5f995dd661

Observation c391dc60-aab2-49e3-89ab-0fe39343235a · outbound

This paper cites Deep Transformer Q-Networks for Partially Observable Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep Transformer Q-Networks for Partially Observable Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.695510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.695510Z digest=sha256:197e9953ace4c77c61346eb84905b3ddc18e43addfb8e7b5ff7c59be02270a7e

Observation 3888cc66-5364-4345-93f6-9de049e8a60e · outbound

This paper cites Deep Recurrent Q-Learning for Partially Observable MDPs.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep Recurrent Q-Learning for Partially Observable MDPs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.701126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.701126Z digest=sha256:a927a2ee44cf4526b17971b8f0de75a90ec35075bb82f62d1b53300540a08706

Observation 72a74ba4-3b51-48b8-be89-0bedcb1171fc · outbound

This paper cites On improving deep reinforcement learning for pomdps,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes On improving deep reinforcement learning for pomdps,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.476678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.705893Z digest=sha256:97a23d399729ef5e16892843cdd8a54803715f57571857cf0154daaab151603a

Observation 46461b7f-00a1-4fad-b3cd-85734e91a628 · outbound

This paper cites Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.714601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.714601Z digest=sha256:95245fef25eb1e85bed68532e5185403db8d1d22cd944ec2f1858baf0311e36f

Observation bf1657e5-c0f3-4a7c-b00d-852e039ba788 · outbound

This paper cites Deep reinforcement learning with bidirectional recurrent neural networks for dynamic spectrum access,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep reinforcement learning with bidirectional recurrent neural networks for dynamic spectrum access,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.463274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.719144Z digest=sha256:d61d4d9a0cfad716d3e070e58dc9bb03f745e36ee8228f3de6850557a73f4647

Observation ae84f17b-566d-4557-b30f-08976440d674 · outbound

This paper cites On transforming reinforcement learning with transformers: The development trajectory,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes On transforming reinforcement learning with transformers: The development trajectory,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.450134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.723470Z digest=sha256:686b34da3e363298f045bba4e4d1602b1b450524278fa1a9a0d033a88bd2b629

Observation a562b3df-f1db-4c05-8ff3-3066d67b3fb7 · outbound

This paper cites Transformer in reinforcement learning for decision-making: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformer in reinforcement learning for decision-making: A survey,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.437267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.727574Z digest=sha256:5b92dcd8b97ef678708548200fbf0a4e487e567d154e7bbfda39fda6bd3234b3

Observation bc4f5edf-f3a8-418b-9dd3-a8428fb3d894 · outbound

This paper cites Attention Is All You Need.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Attention Is All You Need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.731661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.731661Z digest=sha256:4224f117a45715622f27cae2709fb39bf8ffd6d9c2a77b090ab554b548c21ec8

Observation dc739bf0-97d2-4002-b7ab-733b2d16428b · outbound

This paper cites Decision Transformer: Reinforcement Learning via Sequence Modeling.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Decision Transformer: Reinforcement Learning via Sequence Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.735633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.735633Z digest=sha256:104a6cd48b37be2c4d81b189f7ee6540a62dac09eb0443f19b9d8cb19e0935ac

Observation 4cb33299-03b4-4e57-a132-5d852f3352db · outbound

This paper cites Deep Attention Recurrent Q-Network.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep Attention Recurrent Q-Network

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:01:23.170540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.739898Z digest=sha256:f979eb083d4a744a8fff552718e0cf12824a846860d12e49490a2b02b606d40d

Observation bdde2bf2-25da-4f13-825a-1a06159bdf88 · outbound

This paper cites Towards Interpretable Reinforcement Learning Using Attention Augmented Agents.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Towards Interpretable Reinforcement Learning Using Attention Augmented Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.744141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.744141Z digest=sha256:8d30502dad2eae00104631fcc2ff05ccd727344cd3fbd1d48076be226f89b10e

Observation 15e4264b-f93a-4ea3-9c95-268ffa9b797b · outbound

This paper cites Efficient Transformers in Reinforcement Learning using Actor-Learner Distillation.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Efficient Transformers in Reinforcement Learning using Actor-Learner Distillation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.748705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.748705Z digest=sha256:0d253148cdbe6a7facdc1314dfce613ad8fa0da833684e2f85f48379804e532f

Observation c8c00b89-1879-40dc-97da-39507316d0b6 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.753113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.753113Z digest=sha256:4510a797729d5cb5854db6e18e71cfbb6b76ecd128129eb6d20c3237b0cd556e

Observation 95495ef3-5df6-49fb-82f3-4c1d5f2e8526 · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Offline reinforcement learning as one big sequence modeling problem,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.422707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.757384Z digest=sha256:a6a5a682354bd5bff53f166ddb90d2bf87ee868ed4908f929afe0d128719fa04

Observation cf540214-cd83-474a-abbf-f3ea2833c154 · outbound

This paper cites Online Decision Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Online Decision Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.761745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.761745Z digest=sha256:fa82ed349f2c2bbe0a073af3a614a04802fe5f80e524be6d9bec05a526dcae2e

Observation 1d2e5ca5-fe82-49a4-ab95-acd8ef48894f · outbound

This paper cites Structured State Space Models for In-Context Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Structured State Space Models for In-Context Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.766001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.766001Z digest=sha256:768627266b0647a1350dbcb1e72d80de5649e2d51660a31d84cbfdc059be2daf

Observation 9266ec07-5a79-46c4-9b8d-d0c3732ee6af · outbound

This paper cites Mastering Memory Tasks with World Models.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Mastering Memory Tasks with World Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.770341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.770341Z digest=sha256:533497bf9c816023493c01b677fbef14a9962362bc44f4c1951c272290ebbff1

Observation 53be7564-2bce-4acd-a07d-f7e7840a8e45 · outbound

This paper cites Mastering atari with discrete world models,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Mastering atari with discrete world models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.408184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.774501Z digest=sha256:3172a747e3c38aff8b74678c20d785b5e8692292cae6c579a3ac088716358db4

Observation 95412e92-4c5f-4e87-961d-c47b4c45e61f · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.783296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.783296Z digest=sha256:52d1b88beb079669287a69e379059bb0ccbec9f28f8524616aebb198610f472a

Observation 1fb5fdc0-918f-403d-803b-cb9169442c2c · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes RWKV: Reinventing RNNs for the Transformer Era

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.787610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.787610Z digest=sha256:1e58b3d5a68e7f9e57c1ced30224ecdbecae56ca99f5c382e93090146feea3e0

Observation c4b344d7-2a02-4d2c-8cbe-5b654b71387a · outbound

This paper cites Resurrecting Recurrent Neural Networks for Long Sequences.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Resurrecting Recurrent Neural Networks for Long Sequences

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.792509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.792509Z digest=sha256:c0c2fc05bc75b5fe5531685adf5d3ca9e7de498f21df5878836c428207801bc0

Observation 16c939f0-da6f-4dd8-88c1-87f85629d068 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Efficiently modeling long sequences with structured state spaces,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.394090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.797991Z digest=sha256:0ee1102271a11562ba5271f6da6f1e596ef1d6944ac40ebfb2563a10d48c2cf5

Observation fec57d69-189a-4b98-b57d-14c3d8f1049d · outbound

This paper cites Simplified State Space Layers for Sequence Modeling.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Simplified State Space Layers for Sequence Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.806620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.806620Z digest=sha256:1014d5aca35c75e4cd4aca356e38e850a4427dd9055d624f8d25dd8880580686

Observation 28bbaeba-5c66-4e6f-9f28-eae008db609f · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.811361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.811361Z digest=sha256:6480ec909ef7cb43744af2ae4059c45760cdb75cdaa478743d4b07c2eb3fee9f

Observation 30d1836c-1d8b-4175-bf0b-aa61e69e536f · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Efficiently Modeling Long Sequences with Structured State Spaces

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.802215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.802215Z digest=sha256:28c8e9e192f8d2b9c9f1a5eda5c6434639ed2fb5bc482b74b975173d068d293b

Observation 97ea078e-db2f-4e07-a3f1-83beccc4c463 · outbound

This paper cites Longformer: The Long-Document Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Longformer: The Long-Document Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.820257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.820257Z digest=sha256:1fb8459c148ed0a73a6563cb8003a9b8394ffa2d4581acaca8c003de4a5e5131

Observation 83ca39f0-2267-4829-b533-4e45e76c6fca · outbound

This paper cites Transformer-XL: Attentive language models beyond a fixed-length context,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformer-XL: Attentive language models beyond a fixed-length context,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.380533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.824663Z digest=sha256:a00069251a004f7ed4b2b346b3bafb1af4df4e326489ffb8fb85bbb126dff64d

Observation 1e28d407-380d-484e-9aca-ba9d95119dfe · outbound

This paper cites Stabilizing Transformers for Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Stabilizing Transformers for Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.815814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.815814Z digest=sha256:e382bf3eb1c9bd0192bd1f40f40db372fa22686bde9a71ed766359a83fce1454

Observation bfcfe6a5-0e3c-4837-889d-77b8519d1689 · outbound

This paper cites Recurrent Memory Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Recurrent Memory Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.833253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.833253Z digest=sha256:6a6dc03bbb20dcfdca595544bc4d1e6f2dde31c0cbfbe0b3c4e83f98586b2245

Observation abedf9b8-7a34-40b0-b995-1881e58d5461 · outbound

This paper cites Reformer: The Efficient Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Reformer: The Efficient Transformer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.837715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.837715Z digest=sha256:203565f6c29d291903259ef5f3b7303636e14299fc45db4e5b4d42e6df526d1b

Observation bdddb050-6020-4a30-b036-b7c46fca5a9f · outbound

This paper cites Linear transformers are secretly fast weight programmers,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Linear transformers are secretly fast weight programmers,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.366622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.829200Z digest=sha256:9f56e6cc550f27c9aa984577bf681e0f2874104e48332af4b99e08f116b385b0

Observation c186d155-232d-4e77-ac6e-696a6e491d05 · outbound

This paper cites Investigating the Role of Feed-Forward Networks in Transformers Using Parallel Attention and Feed-Forward Net Design.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Investigating the Role of Feed-Forward Networks in Transformers Using Parallel Attention and Feed-Forward Net Design

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.846703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.846703Z digest=sha256:c49c1fb9af991fea4b5f076efe77eca5ba59f942be87e96d4d1086c16218aed9

Observation 2057806e-73f7-49fb-943d-4dc4d76b0a5e · outbound

This paper cites Rethinking transformers in solving pomdps,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Rethinking transformers in solving pomdps,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.351037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.851223Z digest=sha256:19fd100c9db4df56186d109c2a81ee3d949fc1d2776a4b307ed9d8f11110e7e5

Observation a7a99566-99ec-4cd8-8365-799994a5d2cb · outbound

This paper cites Rethinking Attention with Performers.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Rethinking Attention with Performers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.842169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.842169Z digest=sha256:37cd169efd5c4495a5c132a95f612be03fb439e14d2a840be76e167c063d3493

Observation 0d42377e-36e6-426d-a376-46d3fb013090 · outbound

This paper cites Pomdp robot domains,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Pomdp robot domains,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.322069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.859622Z digest=sha256:f622f45ce8f1b45d188ddd37a6320a01e9175be927a903fa20c55cc832159888

Observation 85361ea8-fcb7-41ea-82e3-c8634bc66ce6 · outbound

This paper cites Learning policies for partially observable environments: Scaling up,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Learning policies for partially observable environments: Scaling up,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.307086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.864054Z digest=sha256:70c886e5607353c6b46957fa0c785b598572b1261c753005ff3e4d4a49b07cc7

Observation 1f01bb95-6aeb-41f4-8076-af9d547f5e1c · outbound

This paper cites gym-gridverse: Gridworld domains for fully and partially observable reinforcement learning,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes gym-gridverse: Gridworld domains for fully and partially observable reinforcement learning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.337172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.855486Z digest=sha256:67abd6e302c7bd022661af3b6fefa852104c0c4d77426f5f0fc4bae7b2e21fc0

Observation 083fdfd6-0e13-46c0-a7d8-14ce4a391f09 · outbound

This paper cites Solving large pomdps using real time dynamic programming,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Solving large pomdps using real time dynamic programming,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.292669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:01:22.868400Z digest=sha256:9914e8ef69605e876cb8b92e35d2b691cb7b58eb000a425fda7836bba29749bf

Observation 1933d1b8-cf1c-4040-9f44-594c111672b1 · outbound

This paper cites On Improving Deep Reinforcement Learning for POMDPs.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes On Improving Deep Reinforcement Learning for POMDPs

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.710231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.710231Z digest=sha256:dd06bde818f7d257d61118af70e3381f75fd06585724847fbeb28acc48b98579

Observation 926b6d1d-17ff-4c89-b562-24cb62e6ba2a · outbound

This paper cites Mastering Atari with Discrete World Models.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Mastering Atari with Discrete World Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.778724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.778724Z digest=sha256:366e71d79c06ce477bcbfecfe9c99fbdef6c0e5534b46aab0c8c8b5313fe1152

Pith citing papers

No inbound Pith citation observations are available.