Pith. sign in

Paper Citation Record · LEDGER

Convergence of regularized agent-state-based Q-learning in POMDPs

As of 17 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2508.21314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21314 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:31:38.795970Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33b9f85e-83ba-409e-91d2-23c77c4164d8 · outbound

This paper cites Optimal control of Markov processes with incomplete state information I,.

Convergence of regularized agent-state-based Q-learning in POMDPs Optimal control of Markov processes with incomplete state information I,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.937314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:34.726619Z digest=sha256:d98baf59f2ec976b21c6cc8c9f5468eea4d42fb56ca91173c4019dcbcadb9766

Observation dbbb163e-c714-4370-a962-f4a5485fde5b · outbound

This paper cites The optimal control of partially observable Markov processes over a finite horizon,.

Convergence of regularized agent-state-based Q-learning in POMDPs The optimal control of partially observable Markov processes over a finite horizon,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.834699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:34.806924Z digest=sha256:f08192a49826072e28a2e587555856109dfb168da0bb3453b414913028e655d4

Observation d77d49cb-4264-422f-b0f9-d3a53976a919 · outbound

This paper cites Approximate information state for approximate planning and reinforcement learning in partially observed systems,.

Convergence of regularized agent-state-based Q-learning in POMDPs Approximate information state for approximate planning and reinforcement learning in partially observed systems,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.697682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:34.892758Z digest=sha256:459acac73b97032fabc76c66df7f5a195025cea8a35c89075b98b3e87bdd34aa

Observation e999cc6a-1015-4d36-921d-d4476f78ce11 · outbound

This paper cites Deep recurrent Q-learning for partially observable MDPs.

Convergence of regularized agent-state-based Q-learning in POMDPs Deep recurrent Q-learning for partially observable MDPs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.553437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:34.956485Z digest=sha256:e788635bf423f5d0d17d76ad8b77001084caf6a23c882b3700708b49b725644f

Observation 159ff1fe-54b4-4789-9574-1d09c12456cc · outbound

This paper cites Deep variational reinforcement learning for POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Deep variational reinforcement learning for POMDPs,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.378652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.023211Z digest=sha256:0313613defe135d8f2c90fc24dd134ffeff1fc340ab771e3f86f0df5a77a0077

Observation df196e73-64fd-43da-aa25-4a7373d4756a · outbound

This paper cites On Improving Deep Reinforcement Learning for POMDPs.

Convergence of regularized agent-state-based Q-learning in POMDPs On Improving Deep Reinforcement Learning for POMDPs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:35.087409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:35.087409Z digest=sha256:dda3b67b2f119bdb2a48309b46baa6c7f56f3de196ae5aaf33e651a930d12dd2

Observation c3d8e18d-0db7-49a9-bafb-4eb16fa66946 · outbound

This paper cites Memory-based deep reinforcement learning for POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Memory-based deep reinforcement learning for POMDPs,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.181372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.154351Z digest=sha256:390acf56d75bd68c2be89e2276e21998cf32ba6f70c11d5f00c24b039bc88fa1

Observation 4a2642f5-2397-4d12-acc7-5abeb1c96664 · outbound

This paper cites Simple agent, complex envi- ronment: Efficient reinforcement learning with agent states,.

Convergence of regularized agent-state-based Q-learning in POMDPs Simple agent, complex envi- ronment: Efficient reinforcement learning with agent states,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.035610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.246813Z digest=sha256:43abea4b133d99ab4f8db23d74c8b8e0cade426b92788a97355ae7347c9f9e7a

Observation c9d00860-670b-49b2-8744-f1c80c843f50 · outbound

This paper cites Agent-state based policies in POMDPs: Beyond belief-state MDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Agent-state based policies in POMDPs: Beyond belief-state MDPs,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.868522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.313591Z digest=sha256:11d467c0c7fc621d1c934cdfb95047e065335235a2f0c772ee9d027824d9f93c

Observation c3688d29-c638-4798-aae7-069a54b5d3aa · outbound

This paper cites Reinforcement learning algorithm for partially observable Markov decision problems,.

Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning algorithm for partially observable Markov decision problems,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.655382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.394078Z digest=sha256:7de4c36a404886680fa022fd2a85f4c65f7a1750825e526b8ac37a0ad2dbf504

Observation b638296f-62e1-47f6-945b-fb580ccef064 · outbound

This paper cites Convergence of finite memory Q learning for POMDPs and near optimality of learned policies under filter stability,.

Convergence of regularized agent-state-based Q-learning in POMDPs Convergence of finite memory Q learning for POMDPs and near optimality of learned policies under filter stability,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.505820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.465534Z digest=sha256:94841954aaee989a2be805f6ac4528f5d4ba01bd6fb904f46794e591e4ceaf1b

Observation 5f1c3ac5-5f2b-4344-9903-cbf5a49b5271 · outbound

This paper cites Periodic agent-state based Q- learning for POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Periodic agent-state based Q- learning for POMDPs,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.315477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.530699Z digest=sha256:f7aa2d081a6c113cabcf2239f4d803dc5c4bbffa6830684e27a702880880a307

Observation ede0e294-ec60-42b2-8ef0-664e944b95b9 · outbound

This paper cites Q-learning for stochastic control under general information structures and non-markovian environments,.

Convergence of regularized agent-state-based Q-learning in POMDPs Q-learning for stochastic control under general information structures and non-markovian environments,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.092609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.595069Z digest=sha256:1310171a5600c96c8532f7538995c95fa20107e55a803d0da16b182972db0f8e

Observation 27d0ba31-a0cc-42fd-92bd-10fb265bd4f2 · outbound

This paper cites Reinforcement learning in non-Markovian environments,.

Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning in non-Markovian environments,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.943350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.677050Z digest=sha256:2b8ee4c88f909a6dcfa3ec69e8b03d890897c5302e355afc7a2f942b3142db62

Observation 4b90a2be-c0ae-463a-94f0-7b2268e34eb4 · outbound

This paper cites On actor-critic algorithms,.

Convergence of regularized agent-state-based Q-learning in POMDPs On actor-critic algorithms,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.692776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.765001Z digest=sha256:8fa0c9d27d27df75cf2819e9da6b4f42ba3c1d30c1a1c125d5bbd41c44449cf2

Observation 54d6498f-8d3a-4649-9795-36c8460fc974 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Convergence of regularized agent-state-based Q-learning in POMDPs Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:35.857345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:35.857345Z digest=sha256:d5eaf881e7a73e50de5d90cde3254e0cd49aea00d803444298d76cb1c721993f

Observation 6f54d08a-858c-4b4a-b90f-52c8076f127f · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

Convergence of regularized agent-state-based Q-learning in POMDPs Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.541911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:35.945532Z digest=sha256:d282442f67daf818d158dfe68924f724a668f1a39bd68b4401e9bb11d9a753a9

Observation a5c8d233-b142-4b9c-868f-66b4c9e65581 · outbound

This paper cites Under- standing the impact of entropy on policy optimization,.

Convergence of regularized agent-state-based Q-learning in POMDPs Under- standing the impact of entropy on policy optimization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.364516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.011272Z digest=sha256:fa968db8c90abf146516c7d1b5442f1d3f095ca22951f9268da6bab64653d408

Observation df1c5df4-e07a-464b-9283-9e5d00a0a50c · outbound

This paper cites Deep reinforcement learning that matters,.

Convergence of regularized agent-state-based Q-learning in POMDPs Deep reinforcement learning that matters,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.141508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.109162Z digest=sha256:c9ea499d78e38d9a74c3b210c7a8f353d42836dc099f7894b67c1c705e2b7a82

Observation 80329930-38b3-4885-89ff-c0b9b9139e3f · outbound

This paper cites Relative entropy policy search,.

Convergence of regularized agent-state-based Q-learning in POMDPs Relative entropy policy search,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.996455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.204865Z digest=sha256:0f6cfc8a91c94bcfa21cdea9cdb0770614d546a85337830a3d164807f32cb662

Observation 2b2236a6-2014-46c0-9a35-d4773476fc40 · outbound

This paper cites Trust region policy optimization,.

Convergence of regularized agent-state-based Q-learning in POMDPs Trust region policy optimization,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.767473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.289776Z digest=sha256:ab2cd22ce78c52c7edf129fc3af5203d8217551fc32898042ff588edc387676d

Observation d4ace202-8d60-4877-9cce-60afb331329e · outbound

This paper cites A unified view of entropy- regularized Markov decision processes,.

Convergence of regularized agent-state-based Q-learning in POMDPs A unified view of entropy- regularized Markov decision processes,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.583680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.379160Z digest=sha256:1ffd647e0bfbbb472fea902857ef563aba3036d7e519b9b1777bb2b125e4a117

Observation f8f17d5c-6661-4d4c-83be-9bea98b722b4 · outbound

This paper cites A theory of regularized Markov decision processes,.

Convergence of regularized agent-state-based Q-learning in POMDPs A theory of regularized Markov decision processes,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.369959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.429516Z digest=sha256:6f4770fc5818db52a63b15f22476f2b498d804a7ada1900e5d83afd52e89bdc5

Observation e46cf0c7-e872-4fe8-a4e1-1c1fa766e711 · outbound

This paper cites Learning latent dynamics for planning from pixels,.

Convergence of regularized agent-state-based Q-learning in POMDPs Learning latent dynamics for planning from pixels,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.123291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.495999Z digest=sha256:c006bdddabc23a8c1592183204e22a84436a55fd30b197eac1ce762371941924

Observation 21b2282f-babd-462a-9f4d-ed48a9b90bc4 · outbound

This paper cites Solar: Deep structured representations for model-based reinforcement learning,.

Convergence of regularized agent-state-based Q-learning in POMDPs Solar: Deep structured representations for model-based reinforcement learning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.900185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.562816Z digest=sha256:776dc273e50782bab13022f448d52b6b0012ca54beecd4509e482191d4262eee

Observation f8683189-dbd0-49f4-8ad1-d2f855b17ca7 · outbound

This paper cites Dream to control: Learning behaviors by latent imagination,.

Convergence of regularized agent-state-based Q-learning in POMDPs Dream to control: Learning behaviors by latent imagination,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.687181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.626433Z digest=sha256:1047fe1e7bc60bd04075401154249d2930901dbffbbcddf32b043d654105084c

Observation 544e148c-7e2a-44be-afec-00c55961a26b · outbound

This paper cites Bridging state and history representations: Understanding self-predictive RL,.

Convergence of regularized agent-state-based Q-learning in POMDPs Bridging state and history representations: Understanding self-predictive RL,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.482457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.695055Z digest=sha256:6fa4c9e5f54e2fb4f3c0c12ad7f2cc353b7fc8f5c82ab59bdfe3ab031a1a6d75

Observation e597574b-5bf1-44e1-9e16-55694b7ce7b8 · outbound

This paper cites Entropy-regularized Point-based Value Iteration.

Convergence of regularized agent-state-based Q-learning in POMDPs Entropy-regularized Point-based Value Iteration

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:36.781916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:36.781916Z digest=sha256:86833061fd48cec63905e721a7eac1f89eddfed5eee830ae90e3edc12c0ae7ae

Observation b2bf9b10-891e-4cd9-bba0-2bff1acb2561 · outbound

This paper cites DESPOT: Online POMDP planning with regularization,.

Convergence of regularized agent-state-based Q-learning in POMDPs DESPOT: Online POMDP planning with regularization,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.322394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.867927Z digest=sha256:9372a7d403bf0dc03343d4585033db9e5b70ed9d0fe5fbff5fb3b96da22b37f4

Observation e19dd16f-0ff1-430d-a5d8-04bc53bb48a3 · outbound

This paper cites Smoother entropy for active state trajectory estimation and obfuscation in POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Smoother entropy for active state trajectory estimation and obfuscation in POMDPs,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.069938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:36.936524Z digest=sha256:04d7625a7969bab502cf986a531d07281928fc0191b0de63b617b92f7cd7334d

Observation a4468ba5-39a4-49b0-9778-870181cf0d07 · outbound

This paper cites The limits of pure exploration in POMDPs: When the observation entropy is enough,.

Convergence of regularized agent-state-based Q-learning in POMDPs The limits of pure exploration in POMDPs: When the observation entropy is enough,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.908793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.017812Z digest=sha256:7ac78edf04890166d0463aa07f4524a8dd132740ac117b693b8892d1669f228a

Observation c184507a-1e85-4b33-8a98-74a804bb044f · outbound

This paper cites an unresolved cited work.

Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:37.084415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:37.084415Z digest=sha256:fad151241020dd34b903df386ed3af901dba6ffbcab9c70da02ed19a4755cda2

Observation 18d25e0a-98c9-422f-947b-135b42902d99 · outbound

This paper cites Hiriart-Urruty and C.

Convergence of regularized agent-state-based Q-learning in POMDPs Hiriart-Urruty and C

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.707089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.159941Z digest=sha256:39ab2ba9522b3fb02219d75f0ebae6f52d53d336d8d6ba559c7ad5b2f3025652

Observation a55d2cd3-1b4c-4363-89e4-79d9a2276489 · outbound

This paper cites Differentiable dynamic programming for structured prediction and attention,.

Convergence of regularized agent-state-based Q-learning in POMDPs Differentiable dynamic programming for structured prediction and attention,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.503351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.237763Z digest=sha256:8beeded02058eceb5ea4ce7f800bff3e11f92c4b2db40ee6fd5411c6dc200897

Observation 7cfe5ea7-ef0b-483e-b378-ccb4976e26bd · outbound

This paper cites Sequential decomposition of sequential dynamic teams: applications to real-time communication and networked control sys- tems,.

Convergence of regularized agent-state-based Q-learning in POMDPs Sequential decomposition of sequential dynamic teams: applications to real-time communication and networked control sys- tems,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.331616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.323069Z digest=sha256:1f83c31a1a7fcddfb9c8b7331c361e3d30c5b7c814bcac8304c66a13a452c7b6

Observation 4de4b806-c3dd-4c6d-9232-57f89e0772f5 · outbound

This paper cites Reinforcement learning with deep energy-based policies,.

Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning with deep energy-based policies,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.068652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.412286Z digest=sha256:c0e493b442b3cffa892bfd279017fc330396c5d0da2b0b8e1fc2e666bedf5437

Observation 040ab679-c9b4-4c4e-a0d3-edf9b0ad9552 · outbound

This paper cites A stochastic approximation method,.

Convergence of regularized agent-state-based Q-learning in POMDPs A stochastic approximation method,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.902865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.525757Z digest=sha256:d42a0e29a8f0587d04acef537a2d31137a9d782de8892816659aaddcdda94277

Observation 3ab540f2-df98-4d98-92f6-2610b567d596 · outbound

This paper cites Q-learning,.

Convergence of regularized agent-state-based Q-learning in POMDPs Q-learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.716577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.615856Z digest=sha256:ecbadeca70ceaeeead72149d64a7b13fd5c74c9305ddda2523ea7ac74faa0392

Observation 119f1e27-9cfe-4cb5-b36f-b3e55cf620c2 · outbound

This paper cites Asynchronous stochastic approximation and Q- learning,.

Convergence of regularized agent-state-based Q-learning in POMDPs Asynchronous stochastic approximation and Q- learning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.485248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.652967Z digest=sha256:98cd5902381718d4d1ba0b1b5705e28cc0d2cfed5b98b26a10962ccc8d726e69

Observation 59d2640a-5d45-44aa-b9ff-1e4605571daf · outbound

This paper cites Learning without state- estimation in partially observable Markovian decision processes,.

Convergence of regularized agent-state-based Q-learning in POMDPs Learning without state- estimation in partially observable Markovian decision processes,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.237200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.681926Z digest=sha256:745375bf2864d24e2f9d58e158cbb5d571041a6cf63fc9f1590ffb7c424d3a0c

Observation f091b0cf-ee18-4f4a-bb48-fc5ff88244e9 · outbound

This paper cites Gradient-based algorithms for zeroth- order optimization,.

Convergence of regularized agent-state-based Q-learning in POMDPs Gradient-based algorithms for zeroth- order optimization,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:40.858054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.802366Z digest=sha256:e858e3d7db3a7852a0ce347f2819090c4c6574af701ad026e165d02fcb09f800

Observation df21a8b9-1592-4b0e-a279-2b28286661cd · outbound

This paper cites X a∈A πt(a | Zt+1)Qt(Zt+1, a) − Ω(πt(· |Zt+1))− X a∈A π⋆(a | Zt+1)Qµ(Zt+1, a) + Ω(π⋆(· |Zt+1)) # (a) ≤ γ.

Convergence of regularized agent-state-based Q-learning in POMDPs X a∈A πt(a | Zt+1)Qt(Zt+1, a) − Ω(πt(· |Zt+1))− X a∈A π⋆(a | Zt+1)Qµ(Zt+1, a) + Ω(π⋆(· |Zt+1)) # (a) ≤ γ

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:40.652015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:37.911109Z digest=sha256:779dc99533fa29d128e0aef33f4c9d258aac185029c1b9445881befbb0c8c5aa

Observation 8d962d91-cbdd-4605-aec2-ea7843db0b2c · outbound

This paper cites an unresolved cited work.

Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:31:40.410414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:38.053757Z digest=sha256:cc7b2a4f02319957e8dd682deb6722d4117c0f3cc1f59583b42d59f68d97fa58

Observation 861e7c8a-51fa-45db-8e7c-2d55f55dcccb · outbound

This paper cites future states.

Convergence of regularized agent-state-based Q-learning in POMDPs future states

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:40.070206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:38.266308Z digest=sha256:2c8d692ac7976eaca569802c9d2944fc7e86433264b7b392cb5350949b52614f

Observation c5a98ed7-d2b8-4cf5-b4cb-df55b3d781f7 · outbound

This paper cites X a∈A πℓ t (a | Zt+1)QJℓ+1K t (Zt+1, a) − Ω(πℓ t (· |Zt+1))− X a∈A πℓ,⋆(a | Zt+1)QJℓ+1K µ (Zt+1, a) + Ω(πℓ,⋆(· |Zt+1)) # (a) ≤ γ.

Convergence of regularized agent-state-based Q-learning in POMDPs X a∈A πℓ t (a | Zt+1)QJℓ+1K t (Zt+1, a) − Ω(πℓ t (· |Zt+1))− X a∈A πℓ,⋆(a | Zt+1)QJℓ+1K µ (Zt+1, a) + Ω(πℓ,⋆(· |Zt+1)) # (a) ≤ γ

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:39.710166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:38.420084Z digest=sha256:cf6beebc1aa37bd8764cf9ba47a189297f58b65c3ec8f0015ee43807aff02312

Observation 8f606a3a-2faf-468d-add3-72e98cb0c788 · outbound

This paper cites an unresolved cited work.

Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:31:39.414915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:38.616959Z digest=sha256:29cb0a827b1bc68b042ef2650b4049d3c18df62bf1cdc89a7a9aa8a1d3d14bff

Observation f9b99d8f-132e-4671-ad10-b449dc842323 · outbound

This paper cites These two cases must be considered separately.

Convergence of regularized agent-state-based Q-learning in POMDPs These two cases must be considered separately

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:39.077876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:31:38.795970Z digest=sha256:a494cbfa53ec5bf075fe3492ad9c80ed69197e4d48c6a810f919ebccb7d6a0f6

Pith citing papers

No inbound Pith citation observations are available.