Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:30:49.187659Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2412.18855.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:30:49.187659Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3be00b87-80c9-4e0d-83fb-77693bf1b357 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Constrained policy optimization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e8779d-b6d3-40c8-8977-23a777fb2748 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reinforcement learning: Theory and algorithms
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c1eb6107-27ea-47da-b206-13c9cf8efe86 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reincarnating reinforcement learning: Reusing prior computation to accelerate progress
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aeba181c-5585-4ce2-88a1-953ce5a8e638 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Efficient Online Reinforcement Learning with Offline Data
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f57d225d-8fdb-48d5-9a7b-b1e140eff677 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 931c0e32-c4c3-48b9-a8da-f1bc7699846c · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Decision transformer: Reinforcement learning via sequence modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d97be98-45b7-48a3-8e26-6d41cd6f76c7 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Safe Exploration in Continuous Action Spaces
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a0bb77-79d9-41f5-bf6e-e2a850b4f462 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uncertainty-aware model-based offline reinforcement learning for automated driving
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1be9be94-cf64-49a2-a916-7783e452b959 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b822ef4-6f86-4a48-95c5-9a5f7ddfd70f · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A minimalist approach to offline reinforcement learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26167c32-88bf-4702-b62e-2c88e1012874 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Addressing function approximation error in actor-critic methods
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e884ace-2ae7-4635-b915-da21ba897fbc · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Off-policy deep reinforcement learning without exploration
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94ccdcad-0bf9-47a6-ab8c-1e72494dfeba · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Extreme Q-Learning: MaxEnt RL without Entropy
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1078a651-dda5-442a-99a1-292e83de04f2 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Open and real-world human-ai coordination by heterogeneous training with communication
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 791cd18e-9715-4197-8547-41090e436c30 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A simple unified uncertainty-guided framework for offline-to-online reinforcement learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a59874-9d4a-42a2-8e17-a87128268627 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reinforcement learning with deep energy-based policies
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 680d33f4-22a6-4f31-8ff5-940f5489d238 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f919f741-e582-43b1-9b2d-b1242342896c · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 516b2635-e169-418f-987d-e79741e2376e · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uncertainty-driven pessimistic q-ensemble for offline-to- online reinforcement learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 64af4ca2-64c0-40f6-9c16-df69ed7322cc · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline reinforcement learning with implicit q-learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6746a8bb-fd32-4958-a3e1-2deb673aa3ab · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Conservative q-learning for offline reinforcement learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17fb3d7f-9744-436b-8fe3-685d92aaf3ba · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Batch policy learning under constraints
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c00cb2d8-d4c6-46e5-b035-6c9dc9e02144 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 391ffc41-f4b9-468d-a5cc-564f956456eb · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b857664-28ff-4841-a166-1eed0a315c88 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82fd9920-06d9-4eac-9227-6fc5976b813a · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53a18ba4-5f14-4d3c-8f35-9dd8a1d5391c · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Mildly conservative q-learning for offline reinforcement learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9dfaeb79-4d21-4164-9f6e-3c6368a1aa84 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL MOORe: Model-based Offline-to-Online Reinforcement Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fbd3e492-d978-437d-938e-9749de752dea · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Supported trust region optimization for offline reinforcement learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1281f0c8-a005-4ff5-b632-a1173f4816a4 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Fine-tuning offline policies with optimistic action selection
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5d037d44-40f3-4ef0-96c3-9441caeb1c8b · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 87cd898a-545a-4798-8d34-cfe6b3d63c7c · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e2f4430-3cb6-4f79-ba2e-f0e4e1694785 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77df57cc-3353-487f-a5cf-429697734043 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16c7240c-2801-434c-9ca7-525d0eecf8fa · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Trust region policy optimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1115b8-aee4-4409-a4d4-0693a3fdc85d · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b2e1149-fdc4-4d38-8b3e-976e37b5566f · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Proximal Policy Optimization Algorithms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88530470-1a24-44bc-9604-bad92fa4e204 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Sutton and AndrewG
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 601c7bc8-c6a0-41ef-9765-2c427eebd7dc · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Lever- aging factored action spaces for efficient offline reinforcement learning in healthcare
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a778597d-eb97-48f1-a963-2c861bc24650 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL CORL: Research-oriented deep offline reinforcement learning library
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 01c0eaf4-3bea-4485-811b-50926caddbbc · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reward Constrained Policy Optimization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2f83805-5139-4b0d-a893-8b0679a025e8 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Jump-start reinforcement learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d77c9109-c886-4784-b58a-52e7e42c26cf · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Train once, get a family: State-adaptive balances for offline-to- online reinforcement learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7cafb6e8-c05c-48d4-ba66-1963ea3253e7 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Supported policy optimization for offline reinforcement learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d7ec5233-1756-45db-b3f8-afee46f24507 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7759f3da-00c4-4d39-b472-4b77da252b4d · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A policy-guided imitation approach for offline reinforcement learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 20d6d736-a044-495b-8099-7d0881f00018 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bafb258-e4a9-481d-8b7f-b1ff5d9a276c · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Actor-critic alignment for offline-to-online reinforcement learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4058ab69-5d44-424e-977b-1bd44d22b625 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f05126e9-b6dd-412d-9435-c2f8ec460bd6 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Policy Expansion for Bridging Offline-to-Online Reinforcement Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36075a83-cc30-4203-aa73-eade42be6aea · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1c5647e-6306-4562-ab2d-7b441ae543d8 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 652d11f2-dc55-47c1-92c3-2f48f7e03ee6 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Online decision transformer
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6a610778-6fdc-45ac-94b5-8d919abe8ac1 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL However, our O2SAC still outperforms it with less computational cost during online fine-tuning and less requirements for offline policy
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2de3f5ca-53a8-465d-90c2-41cd7ee63c22 · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL For model-based O2O RL, [31] explores regions with high uncertainty and returns in learned model
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 292accf8-fd76-46a6-88f4-c3b9448ede9b · outbound
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Moreover, [4] find that LayerNorm is favourable for efficient online RL with offline data
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.