Pith. sign in

Paper Citation Record · LEDGER

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms

As of 22 July 2026, this Paper Citation Record lists 60 of 60 outbound references and 2 inbound Pith citation observations for arXiv:2404.14442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.14442 v7

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-24T02:16:04.596951Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-22T06:31:00.163083+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:45:48.543849Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T23:36:38.356022Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact6
  • verified fuzzy54
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28cc24c6-c65a-486a-85fa-197fe053a6ce · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.397653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:ff8d5f2f698b3b041afd669fdad46ffec68183ae4f337bc9d7c5b1add518e526

Observation 97aacc98-5917-43ab-8f17-ce5e51f28664 · outbound

This paper cites Human-level control through deep rein- forcement learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Human-level control through deep rein- forcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.292830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:2dead17e98bfc13fd194ef4b0457cddbbf0b02c311d1fe847e323bc8ae010390

Observation 3cc5354f-bf5f-4605-92aa-a3b3884f803e · outbound

This paper cites Q-learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Q-learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.296427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:0ac0755f90596842fc30b5c7faf7bc863cec35a9a2247acc00ec24e5e436b979

Observation af840059-9b62-4580-bd77-db791cf081e8 · outbound

This paper cites Convergence of st ochastic iterative dynamic programming algorithms.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Convergence of st ochastic iterative dynamic programming algorithms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.300384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:9936aff5d1f0dbde735857677632e8fc00499607f12be69754e34e640985e0ef

Observation e79e4f8b-0f33-47c4-a390-8f4b462ea87f · outbound

This paper cites The ODE method for convergence of stochastic approximation and reinforcement learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms The ODE method for convergence of stochastic approximation and reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.311443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:7edf209f81c036fa9ddb3f65b24f16b18cf9278bca4af36aae59a4aae8ec3c75

Observation 6f1771e3-8a63-461b-8875-d4a6458d8872 · outbound

This paper cites A unified switching system perspective and con vergence analysis of Q- learning algorithms.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms A unified switching system perspective and con vergence analysis of Q- learning algorithms

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.281535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:4d74955db8ef0cc0100e475f50f34b3276fa21d87394ea71f9706fd498439565

Observation 940e2db8-f018-4452-a4d3-d18c1d92a81c · outbound

This paper cites Asynchronous stochastic approximation and Q- learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Asynchronous stochastic approximation and Q- learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.285313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:3ce8db0796d9d38669d2b7b9a0fa49d2090bdd0be743c8a3ac416bf2723eb212

Observation 183a7283-d48a-4457-8366-773752a8d3e1 · outbound

This paper cites Conv ergence results for single-step on-policy reinforcement-learning algorithms.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Conv ergence results for single-step on-policy reinforcement-learning algorithms

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.270691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:85d07a0858b2dbde0e169424af60a0cb33e94fbca339233c3606c9a02d0a29aa

Observation f7574daf-02c6-4a4f-b65f-4bfe96c4dde6 · outbound

This paper cites The asymptotic convergence-rate of Q-lear ning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms The asymptotic convergence-rate of Q-lear ning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.274225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:0300a0de258b4a5e3638047485a524fd690ff15bee84c2ea2ace701ddc6858b2

Observation ec2fd1a0-1e56-461f-92f9-b5246cd4e3c8 · outbound

This paper cites Learning rates for Q-learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Learning rates for Q-learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.265922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:4d410843157299e16a4f5d092dcde8381d2ad70751d76f50dd3d0fc18edb6ad0

Observation 143ef549-f556-4ede-bd22-6b3cb1548ee3 · outbound

This paper cites Error bounds for constant step- size Q-learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Error bounds for constant step- size Q-learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.363059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:5c3b27bf4f0de6c18269f106f7db59ededffcb873fd14fc87e9279acee6d7656

Observation cab0cf58-eaa2-4bbc-8c10-2ea6b677ea25 · outbound

This paper cites Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-24T02:18:44.883030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:434f179172aeba28f19c7db7e6d90f5d079a8353e2a8b174f7c282d4b97164e8

Observation 09b93c36-6021-4eba-827f-4b394ae47756 · outbound

This paper cites Finite-Time Analysis of Asynchronous Stochastic Approximation and $Q$-Learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Finite-Time Analysis of Asynchronous Stochastic Approximation and $Q$-Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:18:44.877485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:7f0f0306e750798523d5c5f0e8f6fe1d55f77955e63a11d994668d1c8f7878e7

Observation 3480e319-29f3-457a-a428-e5f7bc10fc2b · outbound

This paper cites Sample Complexity of Asynchronous Q-Learning: Sharper Analysis and Variance Reduction.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Sample Complexity of Asynchronous Q-Learning: Sharper Analysis and Variance Reduction

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:18:44.857935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:916e8f1ba409c7029b4c5846d8b63e46662700bf2c7774ca096f91ff7246957d

Observation 886d6936-c725-4e7f-8730-96008331033e · outbound

This paper cites A Lyapunov Theory for Finite-Sample Guarantees of Asynchronous Q-Learning and TD-Learning Variants.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms A Lyapunov Theory for Finite-Sample Guarantees of Asynchronous Q-Learning and TD-Learning Variants

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:18:44.890270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:e6ca2fb499a219753198a1ca83051f539dcf52adf430b3cef360f3f7efa984db

Observation 476ec9ca-3f23-4922-9344-d2f89aa69f06 · outbound

This paper cites Finite-sample analysis of contractive stochastic approxim ation using smooth convex envelopes.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Finite-sample analysis of contractive stochastic approxim ation using smooth convex envelopes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.367176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:55467aed4f515a6a4432f6cd0fd3b68f85b4ffcaba764f6d3c276362c777a11e

Observation abfdc2e7-f26b-4d39-870d-914537ff166b · outbound

This paper cites Final iteration convergence of Q-learning: Switching s ystem approach.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Final iteration convergence of Q-learning: Switching s ystem approach

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.348068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:64d9436ae34ace169e50ad815979d89e025946dd51df64f059cad12c3b9c3511

Observation 5cabdf23-b006-4da3-9391-af2b85ac67a8 · outbound

This paper cites A convergento(n) temporal-difference algorithm for off-policy learning with linear function approximation.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms A convergento(n) temporal-difference algorithm for off-policy learning with linear function approximation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.250979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:bb53f007d93fb45d57fa9e85b7bec2be55ed01b11d360afdb6fe498eb06c7daa

Observation 7a3deb8f-347d-4773-bf42-85f8de6abaa5 · outbound

This paper cites Fast gradient-descent methods for temporal-difference learnin g with linear function approx- imation.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Fast gradient-descent methods for temporal-difference learnin g with linear function approx- imation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.255130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:69cfcf734f4cc0810c37a8246b88c1e76ddc1f1bfcf34f5498803bfdc718f20d

Observation 693dfff2-9f50-4540-880a-b40f038321cf · outbound

This paper cites Gradient temporal- difference learning with regularized corrections.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Gradient temporal- difference learning with regularized corrections

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.328433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:e656140539cc4dc77563ac2af51729f6a19c4965af3d3619de65567d5d5c7f4e

Observation 204f6c31-5ca5-4b2d-b4b6-72f465bb9dee · outbound

This paper cites New versions of gradien t temporal difference learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms New versions of gradien t temporal difference learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.243220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:7d68adab56d3415092d695c8020fbcf2194b5b08783a04d29f51924e61493ce9

Observation 4aba900d-c1bd-4a37-9a71-398796c44c6e · outbound

This paper cites An analysis of reinforc ement learning with function approximation.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms An analysis of reinforc ement learning with function approximation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.343943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:6aa1d9354d231b84b28dc8f24a5d9342e070de408c01805e14b5f94c5244393c

Observation 6d5195b9-accc-45cb-8d72-6cac39780f78 · outbound

This paper cites Bhatnagar, H.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Bhatnagar, H

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.340034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:1d36f2c34bd9555b063c62e5f7a6b928c1af67114dfedcce5e37ce652104e72b

Observation b73ef3ca-4540-4b9d-accf-a535f093520c · outbound

This paper cites Reinforceme nt learning with deep energy- based policies.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Reinforceme nt learning with deep energy- based policies

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.239204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:fb1c4caf474dcb15636b4e11bf4a49f459522789375c7706d781257d7568177d

Observation caac22d1-c0e6-4d43-8d40-de4bfbb0653e · outbound

This paper cites Revisiting the softmax Bellman o perator: New benefits and new perspective.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Revisiting the softmax Bellman o perator: New benefits and new perspective

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.246923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:cb806100c28925eea5a00d90367583d3fe050139edd39f8e8bf4755bbe2681c9

Observation 17a14c53-3214-4ba0-8949-8cb9d795d29f · outbound

This paper cites Reinforcemen t learning with dynamic Boltzmann softmax updates.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Reinforcemen t learning with dynamic Boltzmann softmax updates

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.258916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:447032f15f63bb9460a5683f0232b032291eb98e4de1bed731ffb261606867ca

Observation 1433e224-7f39-4771-875b-1805dec87a05 · outbound

This paper cites An alternative softmax operator for reinforcement learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms An alternative softmax operator for reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.289210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:0d116e62c4ba2283be5fe7c0363036c013725b63c4080b20fe10562d57a4a36a

Observation c9cbe175-75fe-4ad1-9b5a-0807d9c9e361 · outbound

This paper cites Smoothed Q-learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Smoothed Q-learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:18:44.864713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:a21349f3d7b2fce8be93fda6d45298a1459ea2d635e077018825e1d3b770475d

Observation c3c3ef14-6946-4d41-a43a-5a3413f9668f · outbound

This paper cites An analog scheme for fixed point computation. i. theory.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms An analog scheme for fixed point computation. i. theory

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.223591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:a8edfe3d985d90ed392b2c917abbed05fc49b86d43633ca5ca4d88fe87a9ca41

Observation 4ccd7d8d-3fc9-4caf-8f24-dc6aee59cf7a · outbound

This paper cites Liberzon, Switching in systems and control.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Liberzon, Switching in systems and control

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.227607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:e36a5b9b253c5d4a2b63baf1bbd4686ae50792579e5d53e6407a52142f940d7b

Observation c4216a61-52e0-47cb-99a4-797ce582ce96 · outbound

This paper cites Unified finite-time error analysis of soft q -learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unified finite-time error analysis of soft q -learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.231158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:c64eb9b39a7f71840f3e45faeed003e35c67250118f205baafb13cd418eee20f

Observation cd0296a3-0135-48af-8b76-e08200c65e47 · outbound

This paper cites SBEED: Convergent reinforcement learning with nonlinear function approximation.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms SBEED: Convergent reinforcement learning with nonlinear function approximation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.393955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:e39c6248d4a6636a1616ab7ec992cddf8fcee373315e71296a19c391d268d6ca

Observation e7867490-3ee2-4be7-bba1-08f01f5863cf · outbound

This paper cites On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-24T02:18:44.871063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:ba24cf8b1bf88a6a7e38b85317ecc2fd313dab8ec89017915dbbe04667f129d5

Observation c7906621-92a0-4863-9e1e-4e29722253d2 · outbound

This paper cites Nonlinear systems.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Nonlinear systems

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.215869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:2528abd915611c5af69b50436da87aa07a93131fa0ebdbc012d5758391b2bb8a

Observation 735754e0-fd8d-4e28-9998-1128d45754de · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.207860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:70bc64f8431e8ec7e55402f4961c6bfcf561cd31c60b153a9a87a5ab1e0ffa24

Observation 24dc0017-51c2-4ee4-b328-425d730c6cd4 · outbound

This paper cites Finite-time analysis of asynchronous q-lea rning under diminishing step-size from control-theoretic view.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Finite-time analysis of asynchronous q-lea rning under diminishing step-size from control-theoretic view

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.378810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:eae0d3311c81cb13e82f2b77bea16583dd3c34b0aca8a02a5007d1f37def5b5a

Observation 050e7beb-d5ec-420e-9067-70bae24944e8 · outbound

This paper cites Kushner and G.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Kushner and G

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.262508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:e49b7658bf7ae88171d913e819f15e5c820b6359ffcbfc90f36309960e4829d2

Observation 32c53f3e-5a1b-4ad3-b662-a2f586b033a8 · outbound

This paper cites A stochastic approximation method.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms A stochastic approximation method

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.277565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:6c81bff92c14f2478231db9b9771a94076edc1e78ae8f61b585a7eb71fbcf1bc

Observation c172d941-e6d5-42d4-87ed-2fa7b06a56d8 · outbound

This paper cites Note on the derivatives with respect to a param eter of the solutions of a system of differential equations.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Note on the derivatives with respect to a param eter of the solutions of a system of differential equations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.382288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:b1bbb0580abe7216a5b67f939e623e65a8a772c0bc04cb48ec172e2bf51e9050

Observation 95d2fe58-6c54-4e9e-a64a-e294ae96110a · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.390422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:0cb5e710cf8848ecde35fee9434513e5b4e7158c83698a884e6c6288215e949d

Observation cef9e683-cc38-4edf-8937-083424958d00 · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.375125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:3d1766251e165330ba731dd373b61e4e8b165eec64b4f12602194d65283c1fb7

Observation 57f7574a-3fbe-4ece-924f-b63212d35649 · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.371153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:9ca42bac2896130d616ff87211b39080856cc9c3e06ee2881b27e0dce4efe8e8

Observation 0d6a051b-712a-4517-a989-a9d2cc61b94b · outbound

This paper cites In addition, there exists a constant C0 < ∞ such that for any initial θ0 ∈ Rn, we have E[∥εk+1∥2 2|Gk] ≤ C0(1 + ∥θk∥2 2), ∀k ≥ 0.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms In addition, there exists a constant C0 < ∞ such that for any initial θ0 ∈ Rn, we have E[∥εk+1∥2 2|Gk] ≤ C0(1 + ∥θk∥2 2), ∀k ≥ 0

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.386309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:5be21a617afd63b46d6cd2737ca15b13d6bf452504eb06d873d033364b020aff

Observation 2a9b36a8-47c3-4aaa-950e-ec246c9e1213 · outbound

This paper cites (17) Lemma 2 ( [5, Borkar and Meyn theorem]).

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms (17) Lemma 2 ( [5, Borkar and Meyn theorem])

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.356212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:7a06a6df7f37d69a9c4d51f6ea6bb7f6b48666771d1252ecd4a30dfa2eecebf1

Observation 17ef970e-76f9-4dba-9839-10f99c0c24c1 · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.332290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:ad181cc2a098f3dcaf61a0acdfa09bc50fe2ec2c3ba84007f4dbab117c76149e

Observation bd502002-6b8c-4ffb-93ce-801bda678d54 · outbound

This paper cites e., xt → H as t → ∞.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms e., xt → H as t → ∞

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.351917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:81545b64a7b3e8189d6ab6155cd55e3bc0f794bdaf57f7016a2db27584b1ba8a

Observation 953d76c0-123b-4329-9b10-d072fc4b4707 · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.318953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:7ca812d3b82970257c7b599f585d03ac727f41c05af04b716647d85679c9439f

Observation 847b5ecc-aada-4ca9-a4a9-3bfd396e8cfb · outbound

This paper cites In addition, there exists a constant C0 < ∞ such that for any initial θ0 ∈ Rn, we have E[∥εk+1∥2 2|Gk] ≤ C0(1 + ∥θk∥2 2), ∀k ≥ 0 with probability one.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms In addition, there exists a constant C0 < ∞ such that for any initial θ0 ∈ Rn, we have E[∥εk+1∥2 2|Gk] ≤ C0(1 + ∥θk∥2 2), ∀k ≥ 0 with probability one

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.323053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:b77a1f2b223b3cb37ef70215e852deee01510abc6ca482a2af1cbc801d259db9

Observation fedb40c7-9b3e-4a52-a6c6-962792dfecda · outbound

This paper cites Lemma 3 ( [38, Robbins and Monro theorem]).

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Lemma 3 ( [38, Robbins and Monro theorem])

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.304156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:390f6f668725ba8206e133be72b921f6d17bd13aed75f5bdd268a113c9dda25e

Observation bf4b9f89-163f-47a2-b39c-5d7cd9189a5d · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.307533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:907759bb678f36835772d1def0180b49be75a2bbf0a4e59dcc8f842e331dd8e0

Observation bab10916-60a6-4e11-8ef4-d5546f33dcde · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.211937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:052b716f119eef748f8b4f2fed92b7ec08770935b7db03b70a6090d08eda39c2

Observation 6ab92562-9ef0-4b49-a867-6ad80717bbd4 · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.219677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:caf1c2521b171c0fcff2599b35e2cf7ed1a93917baf44a5df311c8460f34a243

Observation 81f93788-b4a2-4b42-bc07-c7a8f24104bc · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.235369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:ccb48cbad8a0f38db390c6f05ef0cc1324e1b4798854d4a582735f1608eec74c

Observation 957cb86b-a636-4770-9002-2d478f663253 · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.359694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:7b86d420fdeb6db0071e4037618007c9600cc08ac235d563636dad57754cf166

Observation ba9ee95d-defb-4058-8cb7-b920533667ff · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.199127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:43d156dca0b47150de71260e164f19c8cfec263aadc0925d26959c5123d4a919

Observation 88507e15-fe43-4157-9df0-a16ac61a0deb · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.203427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:fe69916c7b848ef4f064877f0d9925e3b9a0276198f0219ce9899c8ac7f69edd

Observation c3e5cecb-3e8e-476f-9c74-c433b45fc8f4 · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.193553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:b17843faf6eeaf9577dae6ce1341277e49b0186c589a293b44ce57bc67e75979

Observation 1cebaad9-9ca5-4902-849a-8696c85bcc4a · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.336276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:cbaafb8a5555b4adcf45379a95b1dd1634158be53bb0d1cb1b7b431abdda3e5b

Observation 0d384e2e-326c-4254-8c83-ed913515de1b · outbound

This paper cites an unresolved cited work.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Unresolved cited work

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.189893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:58003f8357a280410f25cb1dc949a18750bf8ecd161d8848e0840904ad714f3d

Observation eebe205d-ebd2-4f70-bb88-8a39c3398e8c · outbound

This paper cites Similar to the ODE case, the first three rows (max, L SE, and mellowmax) show that the trajectories converge toward their fixe d points as proved in Theorem 5.

Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms Similar to the ODE case, the first three rows (max, L SE, and mellowmax) show that the trajectories converge toward their fixe d points as proved in Theorem 5

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:18:46.315604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-24T02:16:04.596951Z digest=sha256:9aad9de28b84e09109647ad60ffea746deb3f62c433f52117886b81ed323b8df

Pith citing papers

Observation eeed0fd3-3546-4dcc-8add-dfbd564138a0 · inbound

Contraction-Aligned Analysis of Soft Bellman Residual Minimization with Weighted Lp-Norm for Markov Decision Problem cites this paper.

Contraction-Aligned Analysis of Soft Bellman Residual Minimization with Weighted Lp-Norm for Markov Decision Problem Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:41:43.130250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-10T18:45:48.543849Z digest=sha256:f93d36f102971568caedf687f99d852586ee084c5b2807a14c217ca69f67893e

Observation 624a9226-b504-4f17-9350-74df4165231b · inbound

Safe-Support Q-Learning: Learning without Unsafe Exploration cites this paper.

Safe-Support Q-Learning: Learning without Unsafe Exploration Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:41:43.130250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-07T16:36:35.746034Z digest=sha256:e7e7dbf28e38ab71b145bb02199bbb21c7572379aaa5b05ad828f8a8e2c2b13d