Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T15:19:21.367628Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 59 inbound Pith citation observations for arXiv:1911.11361.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T15:19:21.367628Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:30:59.112579Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
20 of 20 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 06ae7e58-308f-472a-8d03-edbfd6b3e6ae · outbound
Behavior Regularized Offline Reinforcement Learning Maximum a Posteriori Policy Optimisation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b92ebeef-a629-4390-b315-c341af45c056 · outbound
Behavior Regularized Offline Reinforcement Learning An Optimistic Perspective on Offline Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 252d7f97-2309-4331-9cf7-f6e58e14ef8d · outbound
Behavior Regularized Offline Reinforcement Learning Residual algorithms: Reinforcement learning with function approximation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3407c222-eb82-4d02-ae71-0fd10e08a173 · outbound
Behavior Regularized Offline Reinforcement Learning OpenAI Gym
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c1acbe81-2593-43f4-9dc4-2db4eea566ba · outbound
Behavior Regularized Offline Reinforcement Learning Diagnosing Bottlenecks in Deep Q-learning Algorithms
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6b042a1-08d5-4666-8bde-5a1126682347 · outbound
Behavior Regularized Offline Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45e0609b-6d5d-45c7-ad4e-df10c522ee28 · outbound
Behavior Regularized Offline Reinforcement Learning Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 83f3ecad-3210-405a-9a0b-51dca9688678 · outbound
Behavior Regularized Offline Reinforcement Learning Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 35458c85-07ee-40b9-a1ce-6a8346badca0 · outbound
Behavior Regularized Offline Reinforcement Learning Off-Policy Evaluation via Off-Policy Classification
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3c973c9-edca-4681-8eec-d787f2c36444 · outbound
Behavior Regularized Offline Reinforcement Learning Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e80bcb5-9896-41f8-b0e7-abe10bae736d · outbound
Behavior Regularized Offline Reinforcement Learning Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9690e71a-dc55-4ea6-b2d9-353c1ada4f71 · outbound
Behavior Regularized Offline Reinforcement Learning Safe Policy Improvement with Baseline Bootstrapping
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7cdf11a-cf36-4e90-ae59-688703bd6cf0 · outbound
Behavior Regularized Offline Reinforcement Learning Continuous control with deep reinforcement learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e0a7bfb-6e1d-49b1-8409-69b6be9f60f8 · outbound
Behavior Regularized Offline Reinforcement Learning Playing Atari with Deep Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7727fdf-2026-46b1-9439-40732d01decc · outbound
Behavior Regularized Offline Reinforcement Learning Asynchronous methods for deep reinforcement learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9ee4315b-5939-41f4-a9b2-aabf0ac682a0 · outbound
Behavior Regularized Offline Reinforcement Learning Trust-PCL: An Off-Policy Trust Region Method for Continuous Control
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ba809e1-f758-431f-ab64-282d13f4702a · outbound
Behavior Regularized Offline Reinforcement Learning DualDICE: Behavior-Agnostic Estimation of Discounted Stationary Distribution Corrections
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 30fef271-52dd-4251-ad54-d8d0b55315bc · outbound
Behavior Regularized Offline Reinforcement Learning Deep Reinforcement Learning and the Deadly Triad
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 585f8302-f87b-4414-a5bc-955498b65522 · outbound
Behavior Regularized Offline Reinforcement Learning Each dataset contains 1 million transitions
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ce37897d-5852-42be-9718-c7aeae0ccc61 · outbound
Behavior Regularized Offline Reinforcement Learning Gradient penalty (one sided version of the penalty in Gulrajani et al
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f23487c3-36f7-43bf-8abd-e58104d9bd1f · inbound
D4RL: Datasets for Deep Data-Driven Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0917520c-e628-477f-b86c-b7c6188dbb71 · inbound
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems Behavior Regularized Offline Reinforcement Learning
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 765bef19-fcc1-489c-acba-1172f6b4b3f3 · inbound
Decision Transformer: Reinforcement Learning via Sequence Modeling Behavior Regularized Offline Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75f85732-0c04-47a5-b31f-69d2ca963c83 · inbound
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation Behavior Regularized Offline Reinforcement Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f64fabc1-1a1d-408c-bda8-03bc7021c3f7 · inbound
Offline Reinforcement Learning with Implicit Q-Learning Behavior Regularized Offline Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5d510ac-e9e6-4f65-9da6-a99fdb5cef21 · inbound
Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2be11f9-e0d6-47a7-bd22-195289c344a3 · inbound
Is Conditional Generative Modeling all you need for Decision-Making? Behavior Regularized Offline Reinforcement Learning
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55ee2c2c-377b-40b0-b152-b073598958a3 · inbound
IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Behavior Regularized Offline Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 806917f9-2128-4109-904b-6806ed8b63a9 · inbound
A Review of Causal Decision Making Behavior Regularized Offline Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42d5572d-c0b5-4403-b4d6-6f933d934507 · inbound
VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 08129b76-4f33-4ed9-9afa-a22c420908cf · inbound
Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1efc85fd-624a-4372-92a4-6a50001a3ea7 · inbound
Density-Ratio Weighted Behavioral Cloning: Learning Control Policies from Corrupted Datasets Behavior Regularized Offline Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 495b7b13-f15e-4645-9999-fe494460d8c4 · inbound
Value Flows Behavior Regularized Offline Reinforcement Learning
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe55e0d-a0ad-4698-8b26-ff4c518e443e · inbound
Offline Reinforcement Learning with Generative Trajectory Policies Behavior Regularized Offline Reinforcement Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66d2dca-4c80-467d-9ed9-11657ba88217 · inbound
Dichotomous Diffusion Policy Optimization Behavior Regularized Offline Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933b222c-37b7-4b28-82b9-de8b415aa0f3 · inbound
On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage Behavior Regularized Offline Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 824c354f-8c00-4d96-9370-37f55efa3522 · inbound
Conservative Equilibrium Discovery in Offline Game-Theoretic Multiagent Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84689829-318c-4648-9085-d67229a035e0 · inbound
Hyperfastrl: Hypernetwork-based reinforcement learning for unified control of parametric chaotic PDEs Behavior Regularized Offline Reinforcement Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 148c999e-caa3-4570-8e21-210598083330 · inbound
Learning from Demonstration with Failure Awareness for Safe Robot Navigation Behavior Regularized Offline Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a99306da-60fc-411c-afc1-e3b49488e233 · inbound
Pessimism-Free Offline Learning in General-Sum Games via KL Regularization Behavior Regularized Offline Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 932a20a2-733b-4978-8583-627793babb59 · inbound
Pessimism-Free Offline Learning in General-Sum Games via KL Regularization Behavior Regularized Offline Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5616167d-c061-4e09-9067-bd1e7576691a · inbound
Fast Rates in $\alpha$-Potential Games via Regularized Mirror Descent Behavior Regularized Offline Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 183ef916-8b87-4981-a677-326678c5b139 · inbound
Fast Rates in $\alpha$-Potential Games via Regularized Mirror Descent Behavior Regularized Offline Reinforcement Learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2726a3a4-e724-4d5b-bfce-d5b7edadcba8 · inbound
Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Behavior Regularized Offline Reinforcement Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f3af840-3e14-47c8-becb-804337b8e558 · inbound
Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Behavior Regularized Offline Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e5f7374-7691-4479-9c16-8ad386480d2e · inbound
QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Behavior Regularized Offline Reinforcement Learning
Reference 204
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec13c5f2-ae10-4ada-a735-c9dfeb071674 · inbound
AdamO: A Collapse-Suppressed Optimizer for Offline RL Behavior Regularized Offline Reinforcement Learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e6345ba4-7702-4554-aedb-c55b7c1c4fc5 · inbound
An adaptive variance estimator for relative sparsity Behavior Regularized Offline Reinforcement Learning
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a630e76f-3544-4f10-a1ca-e95a4bf528b4 · inbound
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization Behavior Regularized Offline Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a946cab4-1297-4ad0-9968-c016734422f5 · inbound
AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification Behavior Regularized Offline Reinforcement Learning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aefdd0e6-76fd-42db-80cb-e9c752e84518 · inbound
Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow Behavior Regularized Offline Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ffa3e126-e098-40bb-a5eb-f6fe201c8057 · inbound
Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 63c5fded-be7b-4745-a725-c084af64f9b4 · inbound
Zero-shot Imitation Learning by Latent Topology Mapping Behavior Regularized Offline Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f3608e2-762a-4c0e-a06e-b4cf29459cf3 · inbound
RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Behavior Regularized Offline Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4362b89d-50ce-4625-8303-fd111d078b05 · inbound
RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Behavior Regularized Offline Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6116149f-7a80-43da-b934-5cfb98534174 · inbound
TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning Behavior Regularized Offline Reinforcement Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ec9c9ed-582a-4a7a-bb3b-a8a250fdf450 · inbound
Aligning Flow Map Policies with Optimal Q-Guidance Behavior Regularized Offline Reinforcement Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 62492d95-73c1-4e21-92a5-883943d68e7e · inbound
Bridging Domain Gaps with Target-Aligned Generation for Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d66c375-73cc-41e0-b836-f74a2e288b18 · inbound
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy Behavior Regularized Offline Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eae99009-4523-4a96-931f-5d300b910bfa · inbound
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy Behavior Regularized Offline Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc5e60bd-d702-417e-9a3b-29ca20e024c1 · inbound
Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54e9aa71-e3de-4721-a59d-dbf5d64f707f · inbound
COOPO: Cyclic Offline-Online Policy Optimization Algorithm Behavior Regularized Offline Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d18ef97-ed01-4e62-b765-7cec0f68eb86 · inbound
Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference Behavior Regularized Offline Reinforcement Learning
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 674744f1-5159-4346-9947-96a6756aff4e · inbound
SPAR: Support-Preserving Action Rectification Behavior Regularized Offline Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 647f13be-8ac4-4238-b722-e56c9f627a42 · inbound
Moment Matching Q-Learning Behavior Regularized Offline Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 617cd559-579a-47b8-b821-8d623146afb3 · inbound
When Offline Selectors Cannot Beat the Best Single Model: A Diagnostic Study on edX Dropout Prediction Behavior Regularized Offline Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3aab3e34-b56d-488c-bbc7-57e65a2b34ca · inbound
UNIQ: Conformal Calibration for Adaptive Conservatism in Offline Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d359f891-5a75-4e24-ae5a-4f8771172a50 · inbound
Counterfactual Transport Flows for Offline Conservative Trajectory Refinement Behavior Regularized Offline Reinforcement Learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d8a92057-29fe-4451-b238-7579148f1aea · inbound
Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning Behavior Regularized Offline Reinforcement Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40d2b0d4-c6ac-4afa-9ba9-0956e525f26b · inbound
Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 550ffca1-95b7-4952-98c5-05a4076069f0 · inbound
Reversal Q-Learning Behavior Regularized Offline Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8064d5cf-36a9-49c7-a1b7-1f92296fc2e0 · inbound
Offline Reinforcement Learning for Warehouse SLAM Throughput Control Behavior Regularized Offline Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b904ea4d-485d-4915-9f8b-e12248217d44 · inbound
Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors Behavior Regularized Offline Reinforcement Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4efaf69-9cd8-47f3-8099-325661fe148c · inbound
Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience Behavior Regularized Offline Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 333e00da-25ef-47d5-b638-b8b2166a6438 · inbound
Offline Reinforcement Learning for Fluid Controls: Data-based Multi-observational Policy Extraction Behavior Regularized Offline Reinforcement Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 405a6e6e-a03f-4800-8f5c-ed1fd658b44e · inbound
VINE: Taming Generative Control Policies for Reinforcement Learning Behavior Regularized Offline Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d71456da-f613-4873-8330-e33d06acd1ce · inbound
Reinforcement Learning: From Algorithms To Foundation Models Behavior Regularized Offline Reinforcement Learning
Reference 206
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7116211-d4db-4c72-baab-b38f3a602b08 · inbound
Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation Behavior Regularized Offline Reinforcement Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32df2db8-2384-4a95-ad7f-c374e6cba6e4 · inbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Behavior Regularized Offline Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.