Pith. sign in

Paper Citation Record · LEDGER

Behavior Regularized Offline Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 59 inbound Pith citation observations for arXiv:1911.11361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1911.11361 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T15:19:21.367628Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:30:59.112579Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact14
  • verified fuzzy3
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 06ae7e58-308f-472a-8d03-edbfd6b3e6ae · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

Behavior Regularized Offline Reinforcement Learning Maximum a Posteriori Policy Optimisation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.384956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:f2a8d88c76876ea71c27df7bf27163b083bff948098b05cf56593fde4c50b48e

Observation b92ebeef-a629-4390-b315-c341af45c056 · outbound

This paper cites An Optimistic Perspective on Offline Reinforcement Learning.

Behavior Regularized Offline Reinforcement Learning An Optimistic Perspective on Offline Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.390202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:ed5981da44baa79e42991e829983504173d85d67541602920f14ce749e8cb98a

Observation 252d7f97-2309-4331-9cf7-f6e58e14ef8d · outbound

This paper cites Residual algorithms: Reinforcement learning with function approximation.

Behavior Regularized Offline Reinforcement Learning Residual algorithms: Reinforcement learning with function approximation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:19:21.445329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:29d057b1bf66e96425997af81c1c6fda274ee3862c87d63c179383caa237d484

Observation 3407c222-eb82-4d02-ae71-0fd10e08a173 · outbound

This paper cites OpenAI Gym.

Behavior Regularized Offline Reinforcement Learning OpenAI Gym

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T15:19:21.414803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:32c971d43eb933969a4dc206d9643d37a96a48eb30ef2e72c5473064adc1dd52

Observation c1acbe81-2593-43f4-9dc4-2db4eea566ba · outbound

This paper cites Diagnosing Bottlenecks in Deep Q-learning Algorithms.

Behavior Regularized Offline Reinforcement Learning Diagnosing Bottlenecks in Deep Q-learning Algorithms

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T23:25:51.131241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:ab0b115aad914202388bb8d9af522934531893e915d16bd3bbb0d9539ae04f4b

Observation b6b042a1-08d5-4666-8bde-5a1126682347 · outbound

This paper cites Off-Policy Deep Reinforcement Learning without Exploration.

Behavior Regularized Offline Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.423129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:73226576d9d0cc7a9b2ef94d07f1fc857b0626e67059ce7af42333b3aa391b86

Observation 45e0609b-6d5d-45c7-ad4e-df10c522ee28 · outbound

This paper cites an unresolved cited work.

Behavior Regularized Offline Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:19:21.455180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:2b5db975c7bf9576d791b40331be25947c4ba816182bc6e6332fe0eec69812a8

Observation 83f3ecad-3210-405a-9a0b-51dca9688678 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Behavior Regularized Offline Reinforcement Learning Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T15:19:21.427156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:90c90fe7aa7d1e4ffcb9e502f4c78943804000c2b2641da5b51d4d7da16418c1

Observation 35458c85-07ee-40b9-a1ce-6a8346badca0 · outbound

This paper cites Off-Policy Evaluation via Off-Policy Classification.

Behavior Regularized Offline Reinforcement Learning Off-Policy Evaluation via Off-Policy Classification

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.431260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:864c2b7227b7bccb01ab41bde4ef43d56795a04f082ce037eec79c86ddc7ce11

Observation c3c973c9-edca-4681-8eec-d787f2c36444 · outbound

This paper cites Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog.

Behavior Regularized Offline Reinforcement Learning Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.435201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:562fe9b94aa4147532777c83877966fa997af38f22c082c7b8a801800c41f9fe

Observation 3e80bcb5-9896-41f8-b0e7-abe10bae736d · outbound

This paper cites Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction.

Behavior Regularized Offline Reinforcement Learning Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:19:21.438740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:9824442f34d95229e01190e738b9a1c7f481591d02abe254559c5e8a78f7e85b

Observation 9690e71a-dc55-4ea6-b2d9-353c1ada4f71 · outbound

This paper cites Safe Policy Improvement with Baseline Bootstrapping.

Behavior Regularized Offline Reinforcement Learning Safe Policy Improvement with Baseline Bootstrapping

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T23:41:48.517026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:2e4f8a49efaadcec2b6d740ddd7d844ca5c9baa1ed1afbd99f7911fdc59bff21

Observation d7cdf11a-cf36-4e90-ae59-688703bd6cf0 · outbound

This paper cites Continuous control with deep reinforcement learning.

Behavior Regularized Offline Reinforcement Learning Continuous control with deep reinforcement learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T15:19:21.394456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:5890613b43de18f863a9c4001ebcadb6895422b48823207b8c3b08a775f2af33

Observation 2e0a7bfb-6e1d-49b1-8409-69b6be9f60f8 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Behavior Regularized Offline Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-13T15:19:21.398489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:65ad38aeb4dc70eb50673884e7943550544fb1ccab47efe67ea4d8ae86354ad4

Observation d7727fdf-2026-46b1-9439-40732d01decc · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Behavior Regularized Offline Reinforcement Learning Asynchronous methods for deep reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:19:21.450328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:bfe9fc377cd5535ac3d8f0d7fbf707eb7fc72df11dd8e02258431dd004e5ba1b

Observation 9ee4315b-5939-41f4-a9b2-aabf0ac682a0 · outbound

This paper cites Trust-PCL: An Off-Policy Trust Region Method for Continuous Control.

Behavior Regularized Offline Reinforcement Learning Trust-PCL: An Off-Policy Trust Region Method for Continuous Control

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T22:34:15.110658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:11eafdecd103bc8535b59055d4a60022dad208ff5923f23576ae080a2d9f4515

Observation 4ba809e1-f758-431f-ab64-282d13f4702a · outbound

This paper cites DualDICE: Behavior-Agnostic Estimation of Discounted Stationary Distribution Corrections.

Behavior Regularized Offline Reinforcement Learning DualDICE: Behavior-Agnostic Estimation of Discounted Stationary Distribution Corrections

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.407189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:7b1fb4445aaf07db5631d0b2e40d6eefd7658f262d75269c920bcee4c966a76b

Observation 30fef271-52dd-4251-ad54-d8d0b55315bc · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Behavior Regularized Offline Reinforcement Learning Deep Reinforcement Learning and the Deadly Triad

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.411238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:b191936c76d03864cab7f77512eb00f0ee83ccd68da66ea5be69927ab2f57ea7

Observation 585f8302-f87b-4414-a5bc-955498b65522 · outbound

This paper cites Each dataset contains 1 million transitions.

Behavior Regularized Offline Reinforcement Learning Each dataset contains 1 million transitions

Reference 19

Resolution
malformed identifier
raw_fallback, observed 2026-05-13T15:19:21.448121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:e99c48a0345d31fe56ff8e86c597340c3a564e0a2de178b689498ffd36a823c8

Observation ce37897d-5852-42be-9718-c7aeae0ccc61 · outbound

This paper cites Gradient penalty (one sided version of the penalty in Gulrajani et al.

Behavior Regularized Offline Reinforcement Learning Gradient penalty (one sided version of the penalty in Gulrajani et al

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:19:21.452481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:9fe71050fcffc78db55a8d03b0359acaa284e7219325fdf086cffb675a27f23e

Pith citing papers

Observation f23487c3-36f7-43bf-8abd-e58104d9bd1f · inbound

D4RL: Datasets for Deep Data-Driven Reinforcement Learning cites this paper.

D4RL: Datasets for Deep Data-Driven Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:19:17.322890Z digest=sha256:1ed8905413d263018b03e5f758e64d69e2615d18ace87097e8195bf2377998ad

Observation 0917520c-e628-477f-b86c-b7c6188dbb71 · inbound

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems cites this paper.

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems Behavior Regularized Offline Reinforcement Learning

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T11:33:20.892688Z digest=sha256:bd48baf68b019502b22b6aecd6100ccebc30dadc3fc740cf7531cb900bb65fa5

Observation 765bef19-fcc1-489c-acba-1172f6b4b3f3 · inbound

Decision Transformer: Reinforcement Learning via Sequence Modeling cites this paper.

Decision Transformer: Reinforcement Learning via Sequence Modeling Behavior Regularized Offline Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:11:11.319161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T15:11:11.056013Z digest=sha256:bb40ce3da5ad85acdff761632d53fca37e9fcf5434c515ff7670a0c0fe66ea0d

Observation 75f85732-0c04-47a5-b31f-69d2ca963c83 · inbound

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation cites this paper.

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation Behavior Regularized Offline Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T08:51:55.826747Z digest=sha256:e516c9e8582e3ed0996bc9c90d207ead053203ae90d31b598db7e9b8f63b6440

Observation f64fabc1-1a1d-408c-bda8-03bc7021c3f7 · inbound

Offline Reinforcement Learning with Implicit Q-Learning cites this paper.

Offline Reinforcement Learning with Implicit Q-Learning Behavior Regularized Offline Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:47:05.621624Z digest=sha256:e8adb36314d77c746cb8cd1defb3221b8af491afac8645eaea8f5718bc865621

Observation b5d510ac-e9e6-4f65-9da6-a99fdb5cef21 · inbound

Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning cites this paper.

Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:55:15.772701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T07:55:15.678348Z digest=sha256:53676bc1b1d986f8cf32640fa614dbbc36046b28339a316480d66cc453030530

Observation e2be11f9-e0d6-47a7-bd22-195289c344a3 · inbound

Is Conditional Generative Modeling all you need for Decision-Making? cites this paper.

Is Conditional Generative Modeling all you need for Decision-Making? Behavior Regularized Offline Reinforcement Learning

Reference 134

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T15:35:10.764277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T15:35:10.593969Z digest=sha256:9ca074fb3bd4087c52a236f56f736f9614ebb0386cbc862c0bdea44c902e5722

Observation 55ee2c2c-377b-40b0-b152-b073598958a3 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Behavior Regularized Offline Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:cb0f41a536201e36559bc5f24684ae42efd77ffa53dc82c9237edcd56402ca75

Observation 806917f9-2128-4109-904b-6806ed8b63a9 · inbound

A Review of Causal Decision Making cites this paper.

A Review of Causal Decision Making Behavior Regularized Offline Reinforcement Learning

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:47:26.303995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T02:47:21.582667Z digest=sha256:169bd46e36fd3e3ee41902644bcb62a61ec07c06ea6935bcbc399f4c7970fc78

Observation 42d5572d-c0b5-4403-b4d6-6f933d934507 · inbound

VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning cites this paper.

VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:35:03.974431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T19:34:33.263674Z digest=sha256:e20961c2e38d382213bb7f2667d8ef2181b2179a44007b38b5caec12e8728be8

Observation 08129b76-4f33-4ed9-9afa-a22c420908cf · inbound

Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning cites this paper.

Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:59.112579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:30:59.112579Z digest=sha256:d55e46163ba3139d91f8bfd935752f1e432a39839fc01ebda26174f6c5d8068b

Observation 1efc85fd-624a-4372-92a4-6a50001a3ea7 · inbound

Density-Ratio Weighted Behavioral Cloning: Learning Control Policies from Corrupted Datasets cites this paper.

Density-Ratio Weighted Behavioral Cloning: Learning Control Policies from Corrupted Datasets Behavior Regularized Offline Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:54:21.643027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T20:52:08.062450Z digest=sha256:142cee33bde91ef8c1036408158bfc2cd5b6c1c210fc1cbf935cd1e7f3eef800

Observation 495b7b13-f15e-4645-9999-fe494460d8c4 · inbound

Value Flows cites this paper.

Value Flows Behavior Regularized Offline Reinforcement Learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:34.967200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:34.967200Z digest=sha256:5d4804d14447842afee8b26bf841c45b72261d9564cc0ffc2a7f050f95ede434

Observation dbe55e0d-a0ad-4698-8b26-ff4c518e443e · inbound

Offline Reinforcement Learning with Generative Trajectory Policies cites this paper.

Offline Reinforcement Learning with Generative Trajectory Policies Behavior Regularized Offline Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T10:14:00.316126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:14:00.316126Z digest=sha256:5d0ebb694707aa1dbac36a58ba38bc948863a05bb1b8ffb7150db936cb87ae06

Observation f66d2dca-4c80-467d-9ed9-11657ba88217 · inbound

Dichotomous Diffusion Policy Optimization cites this paper.

Dichotomous Diffusion Policy Optimization Behavior Regularized Offline Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:31.749822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:31.749822Z digest=sha256:7114b33eea76b14e7a826f4d35b3ea38db450b74b06ff6e577f1479f6a72d3a7

Observation 933b222c-37b7-4b28-82b9-de8b415aa0f3 · inbound

On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage cites this paper.

On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage Behavior Regularized Offline Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T00:05:19.900233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:05:19.900233Z digest=sha256:3696272465ae341e2733810a439cd6b465e3599551e6b8f736a73456091ebd4c

Observation 824c354f-8c00-4d96-9370-37f55efa3522 · inbound

Conservative Equilibrium Discovery in Offline Game-Theoretic Multiagent Reinforcement Learning cites this paper.

Conservative Equilibrium Discovery in Offline Game-Theoretic Multiagent Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T20:01:01.373185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:01:01.373185Z digest=sha256:ce7d9a90c67b38b73a9ef9582305191c58dc65ffaf124714bf41319489aeffba

Observation 84689829-318c-4648-9085-d67229a035e0 · inbound

Hyperfastrl: Hypernetwork-based reinforcement learning for unified control of parametric chaotic PDEs cites this paper.

Hyperfastrl: Hypernetwork-based reinforcement learning for unified control of parametric chaotic PDEs Behavior Regularized Offline Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:09:00.155657Z digest=sha256:249085384936539c8cdcba40ca5a97e06d9e6cb1febe3be715ab6316abd94beb

Observation 148c999e-caa3-4570-8e21-210598083330 · inbound

Learning from Demonstration with Failure Awareness for Safe Robot Navigation cites this paper.

Learning from Demonstration with Failure Awareness for Safe Robot Navigation Behavior Regularized Offline Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T07:50:54.065679Z digest=sha256:55593a747255e0dc82ef68917e0276d4a17c3b3b1b33c1f11682c1e0868c7df8

Observation a99306da-60fc-411c-afc1-e3b49488e233 · inbound

Pessimism-Free Offline Learning in General-Sum Games via KL Regularization cites this paper.

Pessimism-Free Offline Learning in General-Sum Games via KL Regularization Behavior Regularized Offline Reinforcement Learning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T20:10:30.920114Z digest=sha256:9ecb3591ee8ceb417f3c7f58076a31a8dde38d94f515dbe1df77b55f6fd55ed4

Observation 932a20a2-733b-4978-8583-627793babb59 · inbound

Pessimism-Free Offline Learning in General-Sum Games via KL Regularization cites this paper.

Pessimism-Free Offline Learning in General-Sum Games via KL Regularization Behavior Regularized Offline Reinforcement Learning

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T23:39:13.528272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T23:34:21.091472Z digest=sha256:ac32938f45d00249a9e8106999eea7d1b6ed3d57d1154c9348b2358161a5fcaa

Observation 5616167d-c061-4e09-9067-bd1e7576691a · inbound

Fast Rates in $\alpha$-Potential Games via Regularized Mirror Descent cites this paper.

Fast Rates in $\alpha$-Potential Games via Regularized Mirror Descent Behavior Regularized Offline Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T19:27:15.002190Z digest=sha256:3c7710c64b16be29941885f82c535617da5a0d7e2ec271b8c1530800f2497b90

Observation 183ef916-8b87-4981-a677-326678c5b139 · inbound

Fast Rates in $\alpha$-Potential Games via Regularized Mirror Descent cites this paper.

Fast Rates in $\alpha$-Potential Games via Regularized Mirror Descent Behavior Regularized Offline Reinforcement Learning

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T00:39:18.022793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T00:38:53.592229Z digest=sha256:39db4b456efe4fd0375c310a08f251f4ff47272a4840f4015ce892cf29e44a59

Observation 2726a3a4-e724-4d5b-bfce-d5b7edadcba8 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Behavior Regularized Offline Reinforcement Learning

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:8e98ccb3e8b8d3f9a0e86416c7f0afbd5eeb91bc99ae20fa1adc1afa7acefcf8

Observation 8f3af840-3e14-47c8-becb-804337b8e558 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Behavior Regularized Offline Reinforcement Learning

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T00:45:11.401553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:8aaaa4f00c75879775095e22cd7602ff8ffb11aac9024f31edaffb4e9bcc5568

Observation 2e5f7374-7691-4479-9c16-8ad386480d2e · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Behavior Regularized Offline Reinforcement Learning

Reference 204

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:ba7db1b5dc609b52f4b821b33587237b755b4b98869785658b626502d3106fec

Observation ec13c5f2-ae10-4ada-a735-c9dfeb071674 · inbound

AdamO: A Collapse-Suppressed Optimizer for Offline RL cites this paper.

AdamO: A Collapse-Suppressed Optimizer for Offline RL Behavior Regularized Offline Reinforcement Learning

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T14:48:05.166453Z digest=sha256:868ac098a0c159bbc67f05c013226e133e4d8a1b6e8e78ad5fd5cb93040748bb

Observation e6345ba4-7702-4554-aedb-c55b7c1c4fc5 · inbound

An adaptive variance estimator for relative sparsity cites this paper.

An adaptive variance estimator for relative sparsity Behavior Regularized Offline Reinforcement Learning

Reference 101

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T19:35:46.097113Z digest=sha256:7458fbd3859bd94a3a59ed8a14d5f06e2cfc4e75b41e7c7d1b4687646e6b2ab4

Observation a630e76f-3544-4f10-a1ca-e95a4bf528b4 · inbound

On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization cites this paper.

On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization Behavior Regularized Offline Reinforcement Learning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T19:35:59.832868Z digest=sha256:db4d3ff345dd72ef99b10babcf03d28845a030becc71b92b8270cf4e1ca0391f

Observation a946cab4-1297-4ad0-9968-c016734422f5 · inbound

AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification cites this paper.

AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification Behavior Regularized Offline Reinforcement Learning

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T05:23:18.294600Z digest=sha256:043b74c7c55eb64cd576322d109d22bb10329eda08820620c28f0a64c0043ea4

Observation aefdd0e6-76fd-42db-80cb-e9c752e84518 · inbound

Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow cites this paper.

Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow Behavior Regularized Offline Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:16:36.302353Z digest=sha256:9071e5567f518216e56ba5dd3adc76eaba6787b4190d0eaa959050b3db12208c

Observation ffa3e126-e098-40bb-a5eb-f6fe201c8057 · inbound

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning cites this paper.

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T02:17:25.783688Z digest=sha256:893dda85b4c9b5e4dbb5eb54eae71c8b772ab796c9ad7a63d5ec0ad6edfa1645

Observation 63c5fded-be7b-4745-a725-c084af64f9b4 · inbound

Zero-shot Imitation Learning by Latent Topology Mapping cites this paper.

Zero-shot Imitation Learning by Latent Topology Mapping Behavior Regularized Offline Reinforcement Learning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:31:58.729048Z digest=sha256:5a9bafd97849e061338dfa8d50b5f18b3a8e24d1b568606fbedf3cf13348687d

Observation 3f3608e2-762a-4c0e-a06e-b4cf29459cf3 · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Behavior Regularized Offline Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:da7eb1bbaa91eb95e95723b68fe8b968572245da1d483ecc6222f67318305c1e

Observation 4362b89d-50ce-4625-8303-fd111d078b05 · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Behavior Regularized Offline Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:54:05.781779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:8a7be03e92dc103958bbaeb211c2ae4fe3ae5539bd88e6bd9a617b827be1fa21

Observation 6116149f-7a80-43da-b934-5cfb98534174 · inbound

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning cites this paper.

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning Behavior Regularized Offline Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T04:22:07.953637Z digest=sha256:ba1d387d6f976c86a7803deb8351c21eee9181548470be6e993269a63bc466cb

Observation 6ec9c9ed-582a-4a7a-bb3b-a8a250fdf450 · inbound

Aligning Flow Map Policies with Optimal Q-Guidance cites this paper.

Aligning Flow Map Policies with Optimal Q-Guidance Behavior Regularized Offline Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.456208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:04:22.603786Z digest=sha256:0609c046fa8a1c0f688b67564090b3f044d5dc0878e304d9dc517269150d8a68

Observation 62492d95-73c1-4e21-92a5-883943d68e7e · inbound

Bridging Domain Gaps with Target-Aligned Generation for Offline Reinforcement Learning cites this paper.

Bridging Domain Gaps with Target-Aligned Generation for Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:12:54.825226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:12:19.948073Z digest=sha256:2b3da947bf5d804adc535b69d026f946afd44be815275ad312a6d99cc1751a45

Observation 1d66c375-73cc-41e0-b836-f74a2e288b18 · inbound

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy cites this paper.

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy Behavior Regularized Offline Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:27:51.709195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:26:19.032362Z digest=sha256:4c2501add8222f7f756a69c1c2915316a72beae46cfd5454b416b6ecd3974d06

Observation eae99009-4523-4a96-931f-5d300b910bfa · inbound

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy cites this paper.

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy Behavior Regularized Offline Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:25:46.658353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:35:14.348849Z digest=sha256:af174d2016f7f05bcf4b2abe3d5777cbd4355415dd23954070bd7a4abddc4c56

Observation cc5e60bd-d702-417e-9a3b-29ca20e024c1 · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:25:45.864510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:356ab12f4852ba0006ac337776787bec77a524d29eee45d5e5167af42daf3d9e

Observation 54e9aa71-e3de-4721-a59d-dbf5d64f707f · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Behavior Regularized Offline Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:13:18.046411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:1b6a54ef02179af7156f08638506be32506d74eeef2de0fbd654402f7d916ee8

Observation 1d18ef97-ed01-4e62-b765-7cec0f68eb86 · inbound

Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference cites this paper.

Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference Behavior Regularized Offline Reinforcement Learning

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:23:54.043323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:17:17.682142Z digest=sha256:911e02434d1f17ed27d1f22056777881a60df43197c9263a3a9a7063a0da0267

Observation 674744f1-5159-4346-9947-96a6756aff4e · inbound

SPAR: Support-Preserving Action Rectification cites this paper.

SPAR: Support-Preserving Action Rectification Behavior Regularized Offline Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:43:30.999577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T14:34:39.094753Z digest=sha256:8b82c7fd6f3406ab3777b50b71440af8f4e31c122286ba7921354e80a1f02481

Observation 647f13be-8ac4-4238-b722-e56c9f627a42 · inbound

Moment Matching Q-Learning cites this paper.

Moment Matching Q-Learning Behavior Regularized Offline Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:03:29.519731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:59:15.129557Z digest=sha256:fb47c08ef26f7551a93776c994a2ddd00d069bd4926bd446a1764e795bf0df62

Observation 617cd559-579a-47b8-b821-8d623146afb3 · inbound

When Offline Selectors Cannot Beat the Best Single Model: A Diagnostic Study on edX Dropout Prediction cites this paper.

When Offline Selectors Cannot Beat the Best Single Model: A Diagnostic Study on edX Dropout Prediction Behavior Regularized Offline Reinforcement Learning

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:16:26.269126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T11:09:49.799311Z digest=sha256:2970bc1b7c45d67a82675648e1a0fde1f6531119214ab4f1a9403325acbcf401

Observation 3aab3e34-b56d-488c-bbc7-57e65a2b34ca · inbound

UNIQ: Conformal Calibration for Adaptive Conservatism in Offline Reinforcement Learning cites this paper.

UNIQ: Conformal Calibration for Adaptive Conservatism in Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:43:15.138490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T08:39:34.884726Z digest=sha256:ba34fb0cbbdf6ed3a83e778e6bf88c1bb3ff03d52d771b061d4b266dec884913

Observation d359f891-5a75-4e24-ae5a-4f8771172a50 · inbound

Counterfactual Transport Flows for Offline Conservative Trajectory Refinement cites this paper.

Counterfactual Transport Flows for Offline Conservative Trajectory Refinement Behavior Regularized Offline Reinforcement Learning

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:28.579196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:33:35.857240Z digest=sha256:49d7801f00d6f09b717c196f99e6c408cf642541c83d0a8d2298dfbba2442051

Observation d8a92057-29fe-4451-b238-7579148f1aea · inbound

Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning cites this paper.

Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning Behavior Regularized Offline Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:27:36.361729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:57:31.180928Z digest=sha256:1f6950172b642d3d8247ae4994c2f7f8fa36770640933cd5bfaec98101973bad

Observation 40d2b0d4-c6ac-4afa-9ba9-0956e525f26b · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:17:36.918438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:714d1d546f458ead4b5b2f52d03b63d9b2b9d42f4f84b1a87b4ae7e16f78a37b

Observation 550ffca1-95b7-4952-98c5-05a4076069f0 · inbound

Reversal Q-Learning cites this paper.

Reversal Q-Learning Behavior Regularized Offline Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:38:49.878345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T02:30:24.951689Z digest=sha256:5465bfd30e668867ec5cf4ea0f8127c2ff172eb7343e50360588075fe73c28bf

Observation 8064d5cf-36a9-49c7-a1b7-1f92296fc2e0 · inbound

Offline Reinforcement Learning for Warehouse SLAM Throughput Control cites this paper.

Offline Reinforcement Learning for Warehouse SLAM Throughput Control Behavior Regularized Offline Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:39:45.271472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T08:38:47.196692Z digest=sha256:ea4b8b5273b726b7e005b21f8a7538887ddb5a5057804c613f5764bf22601e88

Observation b904ea4d-485d-4915-9f8b-e12248217d44 · inbound

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors cites this paper.

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors Behavior Regularized Offline Reinforcement Learning

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:30:07.892023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T21:15:07.735336Z digest=sha256:803c05f6b5f3b6602b2ad8a7111d93f1800ff75fa550f646920a43ba49ea564c

Observation d4efaf69-9cd8-47f3-8099-325661fe148c · inbound

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience cites this paper.

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience Behavior Regularized Offline Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-01T18:25:58.759576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T01:57:04.293058Z digest=sha256:df64455c256ad7ae4c3936b2663093675d9da024cbda03548dc89789b4b6728a

Observation 333e00da-25ef-47d5-b638-b8b2166a6438 · inbound

Offline Reinforcement Learning for Fluid Controls: Data-based Multi-observational Policy Extraction cites this paper.

Offline Reinforcement Learning for Fluid Controls: Data-based Multi-observational Policy Extraction Behavior Regularized Offline Reinforcement Learning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:40.856236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T06:25:00.318818Z digest=sha256:ac2f2fdc6f75322f6b83259b5728a798afcc8682885f5f9973c0bc2d7f8a0742

Observation 405a6e6e-a03f-4800-8f5c-ed1fd658b44e · inbound

VINE: Taming Generative Control Policies for Reinforcement Learning cites this paper.

VINE: Taming Generative Control Policies for Reinforcement Learning Behavior Regularized Offline Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T12:17:04.321971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:17:04.321971Z digest=sha256:583984ed633a4f7bcd37acad5113df5d6d33b07de9f75e4d69da6b4538b6cb72

Observation d71456da-f613-4873-8330-e33d06acd1ce · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Behavior Regularized Offline Reinforcement Learning

Reference 206

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:16.846768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:16.846768Z digest=sha256:6ea174aa99b771d36e47d61a9406789859fd236509e6cd08facaf93114d606b6

Observation f7116211-d4db-4c72-baab-b38f3a602b08 · inbound

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation cites this paper.

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation Behavior Regularized Offline Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T13:15:14.620583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:15:14.620583Z digest=sha256:f5572d524d3d1e111e49ed821471358c42777e32c7a4c0cc2c8628d7dcd59c49

Observation 32df2db8-2384-4a95-ad7f-c374e6cba6e4 · inbound

Deep Reinforcement Learning: From First Principles to Reasoning Models cites this paper.

Deep Reinforcement Learning: From First Principles to Reasoning Models Behavior Regularized Offline Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.732902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.732902Z digest=sha256:1e984b98a887104bc71c429b03330782e6bcf288aa47d9967064f32f94b05388