Pith. sign in

Paper Citation Record · LEDGER

Convergence of regularized agent-state-based Q-learning in POMDPs

As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2508.21314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21314 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:31:38.795970Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33b9f85e-83ba-409e-91d2-23c77c4164d8 · outbound

This paper cites Optimal control of Markov processes with incomplete state information I,.

Convergence of regularized agent-state-based Q-learning in POMDPs Optimal control of Markov processes with incomplete state information I,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.937314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:34.726619Z digest=sha256:af5c2bb54df05e45e4499c601c4a0bfd61753eb564d3abd2de974411e64a127c

Observation dbbb163e-c714-4370-a962-f4a5485fde5b · outbound

This paper cites The optimal control of partially observable Markov processes over a finite horizon,.

Convergence of regularized agent-state-based Q-learning in POMDPs The optimal control of partially observable Markov processes over a finite horizon,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.834699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:34.806924Z digest=sha256:dde0806990db62633404ebf5ba76b596590fddc9437ff43e9fbee8efff3a7c15

Observation d77d49cb-4264-422f-b0f9-d3a53976a919 · outbound

This paper cites Approximate information state for approximate planning and reinforcement learning in partially observed systems,.

Convergence of regularized agent-state-based Q-learning in POMDPs Approximate information state for approximate planning and reinforcement learning in partially observed systems,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.697682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:34.892758Z digest=sha256:d9e928d93e8d94fdd463fd112d0ccb0832af13067deee1e84f4d0b59f575337a

Observation e999cc6a-1015-4d36-921d-d4476f78ce11 · outbound

This paper cites Deep recurrent Q-learning for partially observable MDPs.

Convergence of regularized agent-state-based Q-learning in POMDPs Deep recurrent Q-learning for partially observable MDPs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.553437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:34.956485Z digest=sha256:d10dcfcfcb39a5e2a6b59f4a8db0780c2aefc0eda348a2959ac5245ad3df18ed

Observation 159ff1fe-54b4-4789-9574-1d09c12456cc · outbound

This paper cites Deep variational reinforcement learning for POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Deep variational reinforcement learning for POMDPs,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.378652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.023211Z digest=sha256:56ac70f592e080b03f9e3e4994ef53b2c397bf03e02677f261970db79a7cd4f9

Observation df196e73-64fd-43da-aa25-4a7373d4756a · outbound

This paper cites On Improving Deep Reinforcement Learning for POMDPs.

Convergence of regularized agent-state-based Q-learning in POMDPs On Improving Deep Reinforcement Learning for POMDPs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:35.087409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:35.087409Z digest=sha256:a23d97f6a48bf59111212ce67c81ce4787ecb81b8d95d521445e6722a896e2d6

Observation c3d8e18d-0db7-49a9-bafb-4eb16fa66946 · outbound

This paper cites Memory-based deep reinforcement learning for POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Memory-based deep reinforcement learning for POMDPs,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.181372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.154351Z digest=sha256:2be8cc8ef7e2fe0cf11ff4ef1e81fa60cbbf6fad72191805553a58d5876408bc

Observation 4a2642f5-2397-4d12-acc7-5abeb1c96664 · outbound

This paper cites Simple agent, complex envi- ronment: Efficient reinforcement learning with agent states,.

Convergence of regularized agent-state-based Q-learning in POMDPs Simple agent, complex envi- ronment: Efficient reinforcement learning with agent states,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.035610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.246813Z digest=sha256:394b03e1cd498158ee610b7578d1079d0da416605a40261db58ab5d9457e301e

Observation c9d00860-670b-49b2-8744-f1c80c843f50 · outbound

This paper cites Agent-state based policies in POMDPs: Beyond belief-state MDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Agent-state based policies in POMDPs: Beyond belief-state MDPs,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.868522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.313591Z digest=sha256:551c419ef0318a41dd78132f72f309a36f391c436c0e7ff77bdfa2846106cc8e

Observation c3688d29-c638-4798-aae7-069a54b5d3aa · outbound

This paper cites Reinforcement learning algorithm for partially observable Markov decision problems,.

Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning algorithm for partially observable Markov decision problems,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.655382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.394078Z digest=sha256:22537c72a0c3219d944300fe5bf82c67398026cca923f3d4f8e9e0142ff2d26b

Observation b638296f-62e1-47f6-945b-fb580ccef064 · outbound

This paper cites Convergence of finite memory Q learning for POMDPs and near optimality of learned policies under filter stability,.

Convergence of regularized agent-state-based Q-learning in POMDPs Convergence of finite memory Q learning for POMDPs and near optimality of learned policies under filter stability,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.505820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.465534Z digest=sha256:5257ba44e5a5b6f4a8468cd122761a11532ef7cb765e61f139f016e0a020e052

Observation 5f1c3ac5-5f2b-4344-9903-cbf5a49b5271 · outbound

This paper cites Periodic agent-state based Q- learning for POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Periodic agent-state based Q- learning for POMDPs,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.315477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.530699Z digest=sha256:8c1b467c88a990edab59a2f35e201da55e3d425170ae0150447ed8662328e929

Observation ede0e294-ec60-42b2-8ef0-664e944b95b9 · outbound

This paper cites Q-learning for stochastic control under general information structures and non-markovian environments,.

Convergence of regularized agent-state-based Q-learning in POMDPs Q-learning for stochastic control under general information structures and non-markovian environments,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.092609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.595069Z digest=sha256:d674ee90e295faba35c8e97d3dcd8dd8e795aa4e279e1de933e121a1484bf702

Observation 27d0ba31-a0cc-42fd-92bd-10fb265bd4f2 · outbound

This paper cites Reinforcement learning in non-Markovian environments,.

Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning in non-Markovian environments,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.943350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.677050Z digest=sha256:2ce2c868b56b46bdfca75c7bf708b1c250e147f9db3c0f9c17838f81ec8e04c6

Observation 4b90a2be-c0ae-463a-94f0-7b2268e34eb4 · outbound

This paper cites On actor-critic algorithms,.

Convergence of regularized agent-state-based Q-learning in POMDPs On actor-critic algorithms,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.692776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.765001Z digest=sha256:37709742cd6d6da42c11a7fc8038ba2fd332e8e732151763a945b4431b07fd08

Observation 54d6498f-8d3a-4649-9795-36c8460fc974 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Convergence of regularized agent-state-based Q-learning in POMDPs Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:35.857345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:35.857345Z digest=sha256:2b4ece232f43258f9b637aa078d2de27da1abae9cf71786d79306abd3ce7db89

Observation 6f54d08a-858c-4b4a-b90f-52c8076f127f · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

Convergence of regularized agent-state-based Q-learning in POMDPs Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.541911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:35.945532Z digest=sha256:6c3d71c29e15c4598fd11a86778fb35e24a2a1ea707acf83d67877c9a4409025

Observation a5c8d233-b142-4b9c-868f-66b4c9e65581 · outbound

This paper cites Under- standing the impact of entropy on policy optimization,.

Convergence of regularized agent-state-based Q-learning in POMDPs Under- standing the impact of entropy on policy optimization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.364516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.011272Z digest=sha256:14a7c717b803adca55f1cbdf9594b6a3128f3ec0654095318cbdfdd64eecc4fa

Observation df1c5df4-e07a-464b-9283-9e5d00a0a50c · outbound

This paper cites Deep reinforcement learning that matters,.

Convergence of regularized agent-state-based Q-learning in POMDPs Deep reinforcement learning that matters,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.141508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.109162Z digest=sha256:821e67cc5338d0d9d06d4f7f464d6723573c8172244fb8149f84a74b363f6b08

Observation 80329930-38b3-4885-89ff-c0b9b9139e3f · outbound

This paper cites Relative entropy policy search,.

Convergence of regularized agent-state-based Q-learning in POMDPs Relative entropy policy search,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.996455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.204865Z digest=sha256:8e8edbd07d7f6efbc99bc57a1835ea45352d7f27a8cbc47c271e8902e7e653af

Observation 2b2236a6-2014-46c0-9a35-d4773476fc40 · outbound

This paper cites Trust region policy optimization,.

Convergence of regularized agent-state-based Q-learning in POMDPs Trust region policy optimization,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.767473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.289776Z digest=sha256:75170d591d336bc4d18d003defbf53c9fee1eaae1e4c55f96606f296c5578db3

Observation d4ace202-8d60-4877-9cce-60afb331329e · outbound

This paper cites A unified view of entropy- regularized Markov decision processes,.

Convergence of regularized agent-state-based Q-learning in POMDPs A unified view of entropy- regularized Markov decision processes,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.583680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.379160Z digest=sha256:49913882c75dad25bb19fc3443799e0ea3d7653b01b47f46f9595dfb007616ca

Observation f8f17d5c-6661-4d4c-83be-9bea98b722b4 · outbound

This paper cites A theory of regularized Markov decision processes,.

Convergence of regularized agent-state-based Q-learning in POMDPs A theory of regularized Markov decision processes,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.369959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.429516Z digest=sha256:323370f76fbaf404895fe3df621013a7ba3d0a7045aa209a1e6aba1733f8f778

Observation e46cf0c7-e872-4fe8-a4e1-1c1fa766e711 · outbound

This paper cites Learning latent dynamics for planning from pixels,.

Convergence of regularized agent-state-based Q-learning in POMDPs Learning latent dynamics for planning from pixels,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.123291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.495999Z digest=sha256:d1ed43df52bbaf577aee65683d96a88ab1139e56239e6435b2134b87e537346c

Observation 21b2282f-babd-462a-9f4d-ed48a9b90bc4 · outbound

This paper cites Solar: Deep structured representations for model-based reinforcement learning,.

Convergence of regularized agent-state-based Q-learning in POMDPs Solar: Deep structured representations for model-based reinforcement learning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.900185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.562816Z digest=sha256:5a74586c734965da866b6d13f88e6154bb40d8b65f2e686ba0627e1868cb7300

Observation f8683189-dbd0-49f4-8ad1-d2f855b17ca7 · outbound

This paper cites Dream to control: Learning behaviors by latent imagination,.

Convergence of regularized agent-state-based Q-learning in POMDPs Dream to control: Learning behaviors by latent imagination,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.687181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.626433Z digest=sha256:3cade9d4aa4f9e52a7e2d38e7ad19b606ac0b9f48c7bc7b5271b64b8e110c511

Observation 544e148c-7e2a-44be-afec-00c55961a26b · outbound

This paper cites Bridging state and history representations: Understanding self-predictive RL,.

Convergence of regularized agent-state-based Q-learning in POMDPs Bridging state and history representations: Understanding self-predictive RL,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.482457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.695055Z digest=sha256:b02c58de736e2aca6f236eedf50120733f51631c44c3408f5f096b927b9cc7ac

Observation e597574b-5bf1-44e1-9e16-55694b7ce7b8 · outbound

This paper cites Entropy-regularized Point-based Value Iteration.

Convergence of regularized agent-state-based Q-learning in POMDPs Entropy-regularized Point-based Value Iteration

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:36.781916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:36.781916Z digest=sha256:77c94ae5d240baffedd844084179f82b7af32a866ef0d961b6b59f54b2b7e8e0

Observation b2bf9b10-891e-4cd9-bba0-2bff1acb2561 · outbound

This paper cites DESPOT: Online POMDP planning with regularization,.

Convergence of regularized agent-state-based Q-learning in POMDPs DESPOT: Online POMDP planning with regularization,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.322394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.867927Z digest=sha256:1a1297370296a5d4dc5980e92f818ce4532a4db3ceedbc8f750c7e64d330c672

Observation e19dd16f-0ff1-430d-a5d8-04bc53bb48a3 · outbound

This paper cites Smoother entropy for active state trajectory estimation and obfuscation in POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Smoother entropy for active state trajectory estimation and obfuscation in POMDPs,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.069938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:36.936524Z digest=sha256:3b42777358a85f09da334f4f6f9ae98a68379e249a071758a3ad3244581c3d8e

Observation a4468ba5-39a4-49b0-9778-870181cf0d07 · outbound

This paper cites The limits of pure exploration in POMDPs: When the observation entropy is enough,.

Convergence of regularized agent-state-based Q-learning in POMDPs The limits of pure exploration in POMDPs: When the observation entropy is enough,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.908793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.017812Z digest=sha256:c8a18a22404c676025c55911027a65f5bba951e40d008066aec034bb108e5c90

Observation c184507a-1e85-4b33-8a98-74a804bb044f · outbound

This paper cites an unresolved cited work.

Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:37.084415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:37.084415Z digest=sha256:19039d3238393e12768d783d7154301e8f5131862a83063236129a23207e0d63

Observation 18d25e0a-98c9-422f-947b-135b42902d99 · outbound

This paper cites Hiriart-Urruty and C.

Convergence of regularized agent-state-based Q-learning in POMDPs Hiriart-Urruty and C

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.707089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.159941Z digest=sha256:2027c63f482dd2c05a8e49ff9e0cb7f7ca9cedb3a7f05c10e2e115426476de58

Observation a55d2cd3-1b4c-4363-89e4-79d9a2276489 · outbound

This paper cites Differentiable dynamic programming for structured prediction and attention,.

Convergence of regularized agent-state-based Q-learning in POMDPs Differentiable dynamic programming for structured prediction and attention,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.503351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.237763Z digest=sha256:56751121c3d7a5e9ca25a710662d5d244b558d4fd63baae81e451479ad4f4c42

Observation 7cfe5ea7-ef0b-483e-b378-ccb4976e26bd · outbound

This paper cites Sequential decomposition of sequential dynamic teams: applications to real-time communication and networked control sys- tems,.

Convergence of regularized agent-state-based Q-learning in POMDPs Sequential decomposition of sequential dynamic teams: applications to real-time communication and networked control sys- tems,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.331616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.323069Z digest=sha256:2342c4184d2a39f806084905bd21941d9c50ab384aa28e0447f9bbfa284b1c9c

Observation 4de4b806-c3dd-4c6d-9232-57f89e0772f5 · outbound

This paper cites Reinforcement learning with deep energy-based policies,.

Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning with deep energy-based policies,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.068652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.412286Z digest=sha256:021a5aa0059c004ea8205ba5adf7b21ad2e42109bfaa328b12299c5aaa64e57d

Observation 040ab679-c9b4-4c4e-a0d3-edf9b0ad9552 · outbound

This paper cites A stochastic approximation method,.

Convergence of regularized agent-state-based Q-learning in POMDPs A stochastic approximation method,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.902865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.525757Z digest=sha256:05962f8b3e7790b08232c3fef43daef023ae1437fdd5215e1cec4c29ffba0bd0

Observation 3ab540f2-df98-4d98-92f6-2610b567d596 · outbound

This paper cites Q-learning,.

Convergence of regularized agent-state-based Q-learning in POMDPs Q-learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.716577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.615856Z digest=sha256:6592b462edfe64530ebc2bf02e9e10193c65137a68c626c7cce6ae8f2b56bd43

Observation 119f1e27-9cfe-4cb5-b36f-b3e55cf620c2 · outbound

This paper cites Asynchronous stochastic approximation and Q- learning,.

Convergence of regularized agent-state-based Q-learning in POMDPs Asynchronous stochastic approximation and Q- learning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.485248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.652967Z digest=sha256:0385c53ef868ae785342cd216ca1aed342b28ac55e786ce4f821cfb1814fe0cc

Observation 59d2640a-5d45-44aa-b9ff-1e4605571daf · outbound

This paper cites Learning without state- estimation in partially observable Markovian decision processes,.

Convergence of regularized agent-state-based Q-learning in POMDPs Learning without state- estimation in partially observable Markovian decision processes,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.237200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.681926Z digest=sha256:bb2237a0a9404a1388d98b971c15c5316860cd09e48b0d95f4bf7545ffc77210

Observation f091b0cf-ee18-4f4a-bb48-fc5ff88244e9 · outbound

This paper cites Gradient-based algorithms for zeroth- order optimization,.

Convergence of regularized agent-state-based Q-learning in POMDPs Gradient-based algorithms for zeroth- order optimization,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:40.858054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.802366Z digest=sha256:d4b49964a2062522b832def1900ebcc7c934cb2687a5c0658db7547907365320

Observation df21a8b9-1592-4b0e-a279-2b28286661cd · outbound

This paper cites X a∈A πt(a | Zt+1)Qt(Zt+1, a) − Ω(πt(· |Zt+1))− X a∈A π⋆(a | Zt+1)Qµ(Zt+1, a) + Ω(π⋆(· |Zt+1)) # (a) ≤ γ.

Convergence of regularized agent-state-based Q-learning in POMDPs X a∈A πt(a | Zt+1)Qt(Zt+1, a) − Ω(πt(· |Zt+1))− X a∈A π⋆(a | Zt+1)Qµ(Zt+1, a) + Ω(π⋆(· |Zt+1)) # (a) ≤ γ

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:40.652015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:37.911109Z digest=sha256:d678d923ceb4158c6ee139562d5a16cd8fd052dc62cf1f222d2416d65783b65d

Observation 8d962d91-cbdd-4605-aec2-ea7843db0b2c · outbound

This paper cites an unresolved cited work.

Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:31:40.410414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:38.053757Z digest=sha256:a24a1a9b51690e5867bd2ec51f0a9f34d69b4225652c9fbe6ac1e738f85df0eb

Observation 861e7c8a-51fa-45db-8e7c-2d55f55dcccb · outbound

This paper cites future states.

Convergence of regularized agent-state-based Q-learning in POMDPs future states

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:40.070206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:38.266308Z digest=sha256:6fae5fbb069805f8eaac6bd7604e36f2eaec7b7e79e60afefaf8d179b1aeb505

Observation c5a98ed7-d2b8-4cf5-b4cb-df55b3d781f7 · outbound

This paper cites X a∈A πℓ t (a | Zt+1)QJℓ+1K t (Zt+1, a) − Ω(πℓ t (· |Zt+1))− X a∈A πℓ,⋆(a | Zt+1)QJℓ+1K µ (Zt+1, a) + Ω(πℓ,⋆(· |Zt+1)) # (a) ≤ γ.

Convergence of regularized agent-state-based Q-learning in POMDPs X a∈A πℓ t (a | Zt+1)QJℓ+1K t (Zt+1, a) − Ω(πℓ t (· |Zt+1))− X a∈A πℓ,⋆(a | Zt+1)QJℓ+1K µ (Zt+1, a) + Ω(πℓ,⋆(· |Zt+1)) # (a) ≤ γ

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:39.710166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:38.420084Z digest=sha256:7e7d7a1a835e9fc4a0c6d955c5e19a17336e500d0420f844724012124fedf289

Observation 8f606a3a-2faf-468d-add3-72e98cb0c788 · outbound

This paper cites an unresolved cited work.

Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:31:39.414915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:38.616959Z digest=sha256:9651ac0c7f10ab685885b93e8564b12c61804d0f0da34c67fdf9d81910795513

Observation f9b99d8f-132e-4671-ad10-b449dc842323 · outbound

This paper cites These two cases must be considered separately.

Convergence of regularized agent-state-based Q-learning in POMDPs These two cases must be considered separately

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:39.077876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:31:38.795970Z digest=sha256:099954c718b4cc434fe0013d7284aaf6a7e103d7ff3fb95140623e8f77b8b3a1

Pith citing papers

No inbound Pith citation observations are available.