Pith. sign in

Paper Citation Record · LEDGER

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL

As of 14 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2505.19923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19923 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:47.489929Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy5
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 893f724a-cc24-4fe6-b9fd-1cd62d354734 · outbound

This paper cites The mean-wise best results among algorithms are highlighted in bold.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL The mean-wise best results among algorithms are highlighted in bold

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:14:47.818399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:47.489929Z digest=sha256:cab0f9add24730128afc05b19b2a458ea1c420b4ac0bdb148769806bf6599d8a

Observation 7dfa7fc1-3245-47dd-9d6d-545b04214591 · outbound

This paper cites RvS: What is Essential for Offline RL via Supervised Learning?.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL RvS: What is Essential for Offline RL via Supervised Learning?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.370476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.370476Z digest=sha256:b64d31399d550ca8372c0191f11d321036630792e6e1a69e999bccaabf7e1a3d

Observation 58fe6bdf-e92c-4148-80d2-f214ac4ea121 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.440249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.440249Z digest=sha256:2225461fc095167df86ac6c17eb7837af5d83f93ec7236a06dd05862637199b1

Observation cf2ca248-7508-4b23-8b16-305983013a7a · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Planning with Diffusion for Flexible Behavior Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.772661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.772661Z digest=sha256:38460c105f6a8084f4eed7ca21df6e11eac71fd5defcaddcafe660b591322056

Observation 771cf11e-2c94-4336-a504-57a71276bbda · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Offline Reinforcement Learning with Implicit Q-Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.838651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.838651Z digest=sha256:111e8f0f5ca6dafd939fc5490ba95eaf87b17f37a802c26a892ad02606e556a0

Observation 56770b7a-9f6a-4c27-a856-48e770a0cb14 · outbound

This paper cites Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:14:48.307564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:45.930801Z digest=sha256:e8a957e1bbb055d7cac03adf40da9fb9541d2e96a2c5663efa8a85b460f8dd6c

Observation 9e633d48-ad30-4a98-b584-8c13f7e19791 · outbound

This paper cites S., Ghadirzadeh, A., Chen, X., and Finn, C.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL S., Ghadirzadeh, A., Chen, X., and Finn, C

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:50.287789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:46.017316Z digest=sha256:d5b9bb91a03e47f653c682c77081ba84b2b73126265a7206c5b5a829a8afa947

Observation 4b5041e9-8924-418a-a0bb-aa938abaf12b · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.091351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.091351Z digest=sha256:13d8d159795b5e776ed4bc68b9be9445f9b5c8fe3429474fefc4fbfd14b83418

Observation 4abe55e6-075f-41d3-b5de-45378bc73e6d · outbound

This paper cites Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.138209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.138209Z digest=sha256:72e2b9cfe71847945dcf1615deeff747eff1dc31014f8f791fed7041f4d263d5

Observation 4f965292-2506-4043-9daa-6e2b00cb5c13 · outbound

This paper cites Hyperparameter Selection for Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Hyperparameter Selection for Offline Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.200212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.200212Z digest=sha256:2f52ba2105b53439cebdd89bdabef63d4d54e5e33174bf2be21bb242709f2db6

Observation 8dbaed76-6543-4524-a810-85d8bc47c469 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.292149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.292149Z digest=sha256:1a4aa579b3cb39269e946c4d4b70d95d1230d496d3e9567feb7ff9a4221d4848

Observation 7a33a587-91c8-495d-a991-4a2389b37c28 · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.345569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.345569Z digest=sha256:ea39c7f8286dda00112e18ed0cc16c0e546e07ffee68d19a1c46d2127e9386b1

Observation 5795690e-4847-4996-ba18-6a6adf07d008 · outbound

This paper cites Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.427182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.427182Z digest=sha256:dc01bc4e21fbeb80f2b76fd576648fda65b0b5734cd9295a475d09f4984aa003

Observation d018c99c-319e-49f4-b0e0-a31927ab0fdf · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.513693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.513693Z digest=sha256:8220f4a9f0e8224d18b8d4c7d7e9ee62ecc17d91645d44a74ce7fdc0f5c89644

Observation 5cd06e8b-3fc1-415c-8cfe-fe4537e8cdc8 · outbound

This paper cites ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.601651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.601651Z digest=sha256:a3d7fcb9abbff066fb3046dfda261adfdc06360a51cc65946053ad77ea5ad4ff

Observation 9772347e-cd21-4d34-833d-ee2ba45afc40 · outbound

This paper cites Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.698844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.698844Z digest=sha256:c4353401df590039d3e5a19f3204d4275526819afec0f42eb9705b05a5c11348

Observation e191725b-0487-48e2-a856-c8c480519051 · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:49.903609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:46.935078Z digest=sha256:3bd3ea13bf10de103e65cf366769d06c97988ab531d336bd2a075ed1f6f95e8e

Observation 9d17ed07-74aa-4a4a-80e6-e2d3388bac09 · outbound

This paper cites Recent works focus on explicit policy constraints for stochastic policies (Wu et al., 2022; Nair et al.,.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Recent works focus on explicit policy constraints for stochastic policies (Wu et al., 2022; Nair et al.,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:49.457619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:47.118568Z digest=sha256:0f691ba0f16abf60b9bbd7b85f670a76c59820c6bc53e5b698493c950b230b48

Observation 68cea97c-3af2-4c50-b1e8-c77a93ac0da6 · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:49.244395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:47.219318Z digest=sha256:d6888be54ae7838351383b3c301c47e9c4c0e6dc25e6603462aaf832587a9f34

Observation 521b3951-9399-427c-96d8-c02efa20145c · outbound

This paper cites Moreover, the coefficients are updated by maximizing Q-values, lacking the interpretability offered by our method.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Moreover, the coefficients are updated by maximizing Q-values, lacking the interpretability offered by our method

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:48.972316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:47.318561Z digest=sha256:8b84c5d87914f0d1b8bac3fd79c28438f657476e1f364ceec22eed388656ff45

Observation cd96905c-98d3-4dec-860b-916782727a32 · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:48.759894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:47.389710Z digest=sha256:f92bb23678372d1f486c2fa8e597f08643abf4ae0d750afe8caded8f1f02a3a0

Observation 41d6a292-3654-4631-86a1-08c3242581f0 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Off-policy deep reinforcement learning without exploration

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:50.543479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:45.559317Z digest=sha256:3c2af6ad9a4a929e3bdb11adb59baed0d6f8a731fa081052bf8c2ba9ce97d648

Observation 202fe6ac-af03-4756-8236-36e4ef3b38c3 · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.632354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.632354Z digest=sha256:664eac09d55e5ba4366d91938dd7a789f61b4e8529b1996c8925b48d3183c112

Observation 05b7370f-d2b3-493b-ab10-8219106b87b3 · outbound

This paper cites Various CQL variants adjust the constraints or modify the regularizer to avoid excessive pessimism (Lyu et al., 2022; Nakamoto et al., 2024; Mao et al., 2024; Yu et al., 2021).

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Various CQL variants adjust the constraints or modify the regularizer to avoid excessive pessimism (Lyu et al., 2022; Nakamoto et al., 2024; Mao et al., 2024; Yu et al., 2021)

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:49.690918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:47.021956Z digest=sha256:69f537388c952f8027cb9c4afe7de291724fa21b3605ff875e7073f08fa43ffb

Observation e614c14e-8003-4cc3-a08d-0ac9dc9f223a · outbound

This paper cites Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.120978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.120978Z digest=sha256:479476dd5d322d90344778551e14b301dc46190497693ef194abd22bf3decb42

Observation 9ca01919-e7c7-4a62-9b55-695b0f199b9b · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.188595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.188595Z digest=sha256:f7fc1a309885bbb80a9c7c95b16b2c4569e0a22e65735b15b71ae4f873eb8390

Observation 563b4579-5f45-48bf-9108-6c2930b0c1e2 · outbound

This paper cites Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.273426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.273426Z digest=sha256:ef7f4706eddda24955bfa9032128d57ab3e184cb2f5212d6f8f2f3d7f41307de

Observation 4c8ab753-1a18-465f-a148-48665a553261 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.701272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.701272Z digest=sha256:94cc57f0edd5c13e2339dc874eecfd41aef46f832fc00db6faa716664e5e8c05

Observation 28a30096-fa6c-4025-92f7-9aca61919fcd · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:50.114624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:14:46.808770Z digest=sha256:864c62ac8cfa45aaf70156d3202507a062dc3987da76f3dbf43102957b52c06a

Pith citing papers

No inbound Pith citation observations are available.