Pith. sign in

Paper Citation Record · LEDGER

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL

As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2505.19923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19923 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:47.489929Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy5
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 893f724a-cc24-4fe6-b9fd-1cd62d354734 · outbound

This paper cites The mean-wise best results among algorithms are highlighted in bold.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL The mean-wise best results among algorithms are highlighted in bold

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:14:47.818399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:47.489929Z digest=sha256:d06fd5d2c82672048ca07b9eb85b7e18ef59ea9363cc33de1adf7beeb3e59c86

Observation 7dfa7fc1-3245-47dd-9d6d-545b04214591 · outbound

This paper cites RvS: What is Essential for Offline RL via Supervised Learning?.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL RvS: What is Essential for Offline RL via Supervised Learning?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.370476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.370476Z digest=sha256:fc08ad1aad56aa915359bb706f5c140ef5dd0d249b5215bd98145f61cb10b474

Observation 58fe6bdf-e92c-4148-80d2-f214ac4ea121 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.440249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.440249Z digest=sha256:f53d6351debf15ba286048b0c7ef5792a644d915baaced0037db92109e65b119

Observation cf2ca248-7508-4b23-8b16-305983013a7a · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Planning with Diffusion for Flexible Behavior Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.772661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.772661Z digest=sha256:eb0740da78d4d9ed2f8e83d1401261c27e209612839a1c60913ceeff18f3057c

Observation 771cf11e-2c94-4336-a504-57a71276bbda · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Offline Reinforcement Learning with Implicit Q-Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.838651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.838651Z digest=sha256:4a00fd23cbee4747f57a0812727e3111a3d37b41e0055f4b230e9d590a117db7

Observation 56770b7a-9f6a-4c27-a856-48e770a0cb14 · outbound

This paper cites Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:14:48.307564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:45.930801Z digest=sha256:60e41f7c8cf9836387ed2738b952694900337bcad044511e3ff460213f1750a4

Observation 9e633d48-ad30-4a98-b584-8c13f7e19791 · outbound

This paper cites S., Ghadirzadeh, A., Chen, X., and Finn, C.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL S., Ghadirzadeh, A., Chen, X., and Finn, C

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:50.287789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:46.017316Z digest=sha256:430efd29467c64be91de2580c61808cc9395e72d4799d2ac327c6c38ba6a63d4

Observation 4b5041e9-8924-418a-a0bb-aa938abaf12b · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.091351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.091351Z digest=sha256:7f7bd4d89e7c8b1d64698922fb3bcc1ee358fb630fd42aa7f084676a2506ef34

Observation 4abe55e6-075f-41d3-b5de-45378bc73e6d · outbound

This paper cites Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.138209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.138209Z digest=sha256:ed541e48ff7919584b2926ea8fbfb2f4706ea54d19915b9682d841251fcdf72e

Observation 4f965292-2506-4043-9daa-6e2b00cb5c13 · outbound

This paper cites Hyperparameter Selection for Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Hyperparameter Selection for Offline Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.200212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.200212Z digest=sha256:dc961b7f3a493c1d0efee31a2ccbe808e7015abc9304a814a1819b861c3c4674

Observation 8dbaed76-6543-4524-a810-85d8bc47c469 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.292149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.292149Z digest=sha256:b162ea51063c1b25c6f82e44569c5c7cbac394004a807b4b6f435deee8d5fbc4

Observation 7a33a587-91c8-495d-a991-4a2389b37c28 · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.345569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.345569Z digest=sha256:8aa1237894e1a790c0399c9d537f8cd7b65dcc56c358e0d6e4767591b29e9ca6

Observation 5795690e-4847-4996-ba18-6a6adf07d008 · outbound

This paper cites Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.427182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.427182Z digest=sha256:10a0a82c464095c41518dd3f456c381af1e97f1389ee514c09316d41fc7d794c

Observation d018c99c-319e-49f4-b0e0-a31927ab0fdf · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.513693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.513693Z digest=sha256:9ff296615f63a7439887dcc55f9dba545141da119e495820a6530e101d897ba8

Observation 5cd06e8b-3fc1-415c-8cfe-fe4537e8cdc8 · outbound

This paper cites ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.601651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.601651Z digest=sha256:6af2f7f43d049ebf704577d35a49b80e1cef9f33e8c5df0b5f28049be0334d1e

Observation 9772347e-cd21-4d34-833d-ee2ba45afc40 · outbound

This paper cites Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.698844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.698844Z digest=sha256:883338511b474d7439fc88a2688ad58cc67bda11b99da043d3bfd998d912b5ee

Observation e191725b-0487-48e2-a856-c8c480519051 · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:49.903609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:46.935078Z digest=sha256:01e023029fcbc732380a1aec2a9b127cf2773ccb28524b5d163a6b7b3186ddcf

Observation 9d17ed07-74aa-4a4a-80e6-e2d3388bac09 · outbound

This paper cites Recent works focus on explicit policy constraints for stochastic policies (Wu et al., 2022; Nair et al.,.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Recent works focus on explicit policy constraints for stochastic policies (Wu et al., 2022; Nair et al.,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:49.457619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:47.118568Z digest=sha256:caeffa08c9bbf5c1dd7cbe5624392ea5cf9f44ada49b663e110c78553098c1bd

Observation 68cea97c-3af2-4c50-b1e8-c77a93ac0da6 · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:49.244395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:47.219318Z digest=sha256:2936fbbc0fff6cc59aa9e9cb570a843dcf7ed0c3d4e693d7551a4153ba6306b5

Observation 521b3951-9399-427c-96d8-c02efa20145c · outbound

This paper cites Moreover, the coefficients are updated by maximizing Q-values, lacking the interpretability offered by our method.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Moreover, the coefficients are updated by maximizing Q-values, lacking the interpretability offered by our method

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:48.972316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:47.318561Z digest=sha256:bed03f9705ec3cf8af5e091b0d1cf7a6e135c0006a6cc4f406c269c7a0bb57c1

Observation cd96905c-98d3-4dec-860b-916782727a32 · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:48.759894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:47.389710Z digest=sha256:ccf54e0632f9ccccda8a0683c1f8dbd471b397f413373db8d6d910570b2cf14f

Observation 41d6a292-3654-4631-86a1-08c3242581f0 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Off-policy deep reinforcement learning without exploration

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:50.543479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:45.559317Z digest=sha256:feccbe638a2d960d068be5448af4a701d31037cc3026c575407a57d49172adf7

Observation 202fe6ac-af03-4756-8236-36e4ef3b38c3 · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.632354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.632354Z digest=sha256:7f0ad93ec2c2ec6a607ca781d39f19d7737b90e8e0fbc603cdfad6f2796353c4

Observation 05b7370f-d2b3-493b-ab10-8219106b87b3 · outbound

This paper cites Various CQL variants adjust the constraints or modify the regularizer to avoid excessive pessimism (Lyu et al., 2022; Nakamoto et al., 2024; Mao et al., 2024; Yu et al., 2021).

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Various CQL variants adjust the constraints or modify the regularizer to avoid excessive pessimism (Lyu et al., 2022; Nakamoto et al., 2024; Mao et al., 2024; Yu et al., 2021)

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:49.690918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:47.021956Z digest=sha256:d750fcc32754b1e23cf8a3df1fbf71bb88bd95d1b9ef2958f035f2cdf12f72d5

Observation e614c14e-8003-4cc3-a08d-0ac9dc9f223a · outbound

This paper cites Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.120978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.120978Z digest=sha256:40dd75564d595f6b2f28a797944b37ac81841062c6e4b8a642f963fce30011e9

Observation 9ca01919-e7c7-4a62-9b55-695b0f199b9b · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.188595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.188595Z digest=sha256:a255c3399e050af2f98f41e0b4c26cd751c9ea505369a2c05fe136810d780175

Observation 563b4579-5f45-48bf-9108-6c2930b0c1e2 · outbound

This paper cites Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.273426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.273426Z digest=sha256:783f86dec9381ce1e1e53b4d27568568caea326dd87a3d4597526a6ebb298cb9

Observation 4c8ab753-1a18-465f-a148-48665a553261 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.701272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.701272Z digest=sha256:2dd106f4875b89d1da4fad0f23c4e04e5307cff35d59f26892694ebfe901ae07

Observation 28a30096-fa6c-4025-92f7-9aca61919fcd · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:50.114624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:14:46.808770Z digest=sha256:1728c41e1a09d4c587f7eaa53acb3232cef3c5c47e92a29a5d17306acac3060c

Pith citing papers

No inbound Pith citation observations are available.