Pith. sign in

Paper Citation Record · LEDGER

Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2110.05038.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.05038 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:48.448037Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:49:02.344960Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 69a54379-6bff-43dd-99dd-511513960743 · inbound

Decoupled Hierarchical Reinforcement Learning with State Abstraction for Discrete Grids cites this paper.

Decoupled Hierarchical Reinforcement Learning with State Abstraction for Discrete Grids Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:48.448037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:48.448037Z digest=sha256:88fb36720511ef9b76a0b21bf2d8ff73bcea1724fa4451bd0d4205d71e1dbb0b

Observation 1bba64b6-349b-4b40-a804-ee262217b0ec · inbound

GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation cites this paper.

GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T20:48:53.701392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:48:53.701392Z digest=sha256:9daa5fd4e26becb32d5a62dc15c6d759c18e8fe079beb4764e354ed5618253e2

Observation ac32d414-1f4a-4eb3-ab82-9f50314332d1 · inbound

Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives cites this paper.

Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T22:51:56.016926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:51:56.016926Z digest=sha256:aeaef8cd65d6c398f74c07ee0120a3105f1b721ec8696b8a6aaf34ed3964caa6

Observation 5225a48a-9bbf-4634-ae6d-deddcbac7be3 · inbound

MINT: Minimal Information Neuro-Symbolic Tree for Objective-Driven Knowledge-Gap Reasoning and Active Elicitation cites this paper.

MINT: Minimal Information Neuro-Symbolic Tree for Objective-Driven Knowledge-Gap Reasoning and Active Elicitation Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:17:30.678350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T07:13:42.705619Z digest=sha256:4c1c12356553e133e0adadeb071e86589b7867955c8f80c35b66e6ca07a3a477

Observation 957fdfbf-93a6-49c2-9e88-592239812121 · inbound

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent cites this paper.

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:46:35.740096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:46:15.275441Z digest=sha256:ed23fe4a6a528d7322aea4bf3bec43e9c4e34dcfb7c1709bf94f8062ee670061

Observation 24db73d1-e528-4a8c-b857-aa1dcdd88a04 · inbound

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games cites this paper.

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:30:17.383704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:30:17.383704Z digest=sha256:6f1c25dcc4b896a0dfcabe5e3c8994d96aba77ae0ce7ec3e991efce0f069f9aa

Observation ef71cc07-f3e7-44d3-8589-4e31c3b17b07 · inbound

Belief-State RWKV for Reinforcement Learning under Partial Observability cites this paper.

Belief-State RWKV for Reinforcement Learning under Partial Observability Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.937670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T21:56:09.642064Z digest=sha256:23e029785d8a16e0f0ebccc8c6eabc9fb62cdf9029685c5f8791f87187547e19

Observation 5ff5e595-5504-49ed-a64d-a8d53a613bdc · inbound

Recurrent Deep Reinforcement Learning for Chemotherapy Control under Partial Observability cites this paper.

Recurrent Deep Reinforcement Learning for Chemotherapy Control under Partial Observability Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:38.523134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:51:16.753971Z digest=sha256:bd433c0f18614e6a49bc4de0cdfd7993cd62ff3327a6004daa8daf541da10456

Observation 9a3c8b48-9bf3-44c3-8636-152afeb3f5c2 · inbound

Learning POMDP World Models from Observations with Language-Model Priors cites this paper.

Learning POMDP World Models from Observations with Language-Model Priors Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:57:52.987389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:57:28.642081Z digest=sha256:a26a4dffe2fd105c2ce1bf7e4fd9e5723556407672da16bf7e4d05c0d7fad7a4

Observation 1076a5da-7386-47bf-a13f-c3efd2ab27bb · inbound

Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets cites this paper.

Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:02.347183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:47:16.490817Z digest=sha256:004481dd14f60094d8b08c9bdfaf8cc96a7b0320a5d35347ddf00b41ecc0d9ff

Observation aade19e8-2102-4b5c-b4e0-7afadf6e6cfe · inbound

When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary cites this paper.

When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:40.649329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:40.649329Z digest=sha256:304997c81e6aaf46c1adc505b3cdee8a53ffd66b0d350f66a4cf43951e4bf2ff