Pith. sign in

Paper Citation Record · LEDGER

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL

As of 12 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2412.18855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18855 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:30:49.187659Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact3
  • verified fuzzy22
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3be00b87-80c9-4e0d-83fb-77693bf1b357 · outbound

This paper cites Constrained policy optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Constrained policy optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.935676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.935676Z digest=sha256:0b93ffe2b0da54a5a07a6567c198fbd007077b1cd02e0e05f96a2e82f18fae5f

Observation b8e8779d-b6d3-40c8-8977-23a777fb2748 · outbound

This paper cites Reinforcement learning: Theory and algorithms.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reinforcement learning: Theory and algorithms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.087909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:48.941106Z digest=sha256:cc30c2aa57561fde7a2e749c27c76958473699ef5de6169ae340b7b51b17b3a9

Observation c1eb6107-27ea-47da-b206-13c9cf8efe86 · outbound

This paper cites Reincarnating reinforcement learning: Reusing prior computation to accelerate progress.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reincarnating reinforcement learning: Reusing prior computation to accelerate progress

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.071758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:48.946054Z digest=sha256:206da8619a5e10cd8e9d8cb278d334790a4c8104e8127355798195d66c2f5509

Observation aeba181c-5585-4ce2-88a1-953ce5a8e638 · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Efficient Online Reinforcement Learning with Offline Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.950682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.950682Z digest=sha256:0a0154d5f0ad40f4e100905c78e8cfc41d3b48d8d133ac667e3e71cfa6916473

Observation f57d225d-8fdb-48d5-9a7b-b1e140eff677 · outbound

This paper cites Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.955499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.955499Z digest=sha256:9266a32f2a89885f7ac0877c6a57f323f38f4d2a3b1377cbae56913ea115e4b9

Observation 931c0e32-c4c3-48b9-a8da-f1bc7699846c · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Decision transformer: Reinforcement learning via sequence modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.960378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.960378Z digest=sha256:f0803caebe553b93f3f6e9493653cbcecf80d732cccfa7153f7ee457db6c8051

Observation 8d97be98-45b7-48a3-8e26-6d41cd6f76c7 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Safe Exploration in Continuous Action Spaces

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.965286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.965286Z digest=sha256:05467df92d2e6701f740252db8764a557079a8c7249d7c81160edce4165655e7

Observation 16a0bb77-79d9-41f5-bf6e-e2a850b4f462 · outbound

This paper cites Uncertainty-aware model-based offline reinforcement learning for automated driving.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uncertainty-aware model-based offline reinforcement learning for automated driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.047352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:48.969948Z digest=sha256:61a3757eff4928149e68f6252c0260f5ed0262a29cf4991e52fe817fa7ac959c

Observation 1be9be94-cf64-49a2-a916-7783e452b959 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.974180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.974180Z digest=sha256:396c85d984e1cf97ef2292e00fec767386c5fa672374490da7f9a5355b511ed6

Observation 3b822ef4-6f86-4a48-95c5-9a5f7ddfd70f · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A minimalist approach to offline reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.978973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.978973Z digest=sha256:3b4be487506085793c6dc907407bc03a6b9c09f539a124f4521fa609fed26b25

Observation 26167c32-88bf-4702-b62e-2c88e1012874 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Addressing function approximation error in actor-critic methods

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.983533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.983533Z digest=sha256:be710b2f0db1e40fa92e3f1d16c4e9915e3a71f8e6f15d4fefebf9b1ea02d000

Observation 3e884ace-2ae7-4635-b915-da21ba897fbc · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Off-policy deep reinforcement learning without exploration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.987827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.987827Z digest=sha256:c6bda67070a2941139c450872cb6b26e74a303f7eb675446e18317901144d18e

Observation 94ccdcad-0bf9-47a6-ab8c-1e72494dfeba · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.991946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.991946Z digest=sha256:07de8bb464172ad1de125083965d29e1352d1b06c3f64a0a1184c86c1eeaca2a

Observation 1078a651-dda5-442a-99a1-292e83de04f2 · outbound

This paper cites Open and real-world human-ai coordination by heterogeneous training with communication.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Open and real-world human-ai coordination by heterogeneous training with communication

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.001318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:48.997192Z digest=sha256:2b4e02b0d6ec4b2a2956c4d5e0fff9f7259ec7bc5e14bb4066dbb0e5c87ef517

Observation 791cd18e-9715-4197-8547-41090e436c30 · outbound

This paper cites A simple unified uncertainty-guided framework for offline-to-online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A simple unified uncertainty-guided framework for offline-to-online reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.001554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.001554Z digest=sha256:18ca06050d5b0aed2f090be27cf2efb7f031e78520d415187891caf185f34ab7

Observation a2a59874-9d4a-42a2-8e17-a87128268627 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reinforcement learning with deep energy-based policies

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.005983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.005983Z digest=sha256:ae7eefe5ce608a9e4ed650b3e76e6c5bec386a3071ea79e69c2a2072ac9d76e8

Observation 680d33f4-22a6-4f31-8ff5-940f5489d238 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.010070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.010070Z digest=sha256:5bbc19e6d16ffc50d90d743f79da3530b0166ecb86a54a005a0cf5ba7c3a5596

Observation f919f741-e582-43b1-9b2d-b1242342896c · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.014463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.014463Z digest=sha256:61e6d35184044b303bd2d7997175460b37243158a1db8ba7c1c694b50fc9b2e5

Observation 516b2635-e169-418f-987d-e79741e2376e · outbound

This paper cites Uncertainty-driven pessimistic q-ensemble for offline-to- online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uncertainty-driven pessimistic q-ensemble for offline-to- online reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.969216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.019462Z digest=sha256:0561e84718b0b9dfeb43af87570699f2c6f4e42c1b10362b67903344154aa996

Observation 64af4ca2-64c0-40f6-9c16-df69ed7322cc · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline reinforcement learning with implicit q-learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.023652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.023652Z digest=sha256:f8e63a031136cd03bfcca268cd18dcc9f926c7015c81d66b3d9dcebf05120679

Observation 6746a8bb-fd32-4958-a3e1-2deb673aa3ab · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Conservative q-learning for offline reinforcement learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.027727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.027727Z digest=sha256:f1568780b6a176ceb5e5658049f54ef5d0e131e08abbb7bc0677e75354954613

Observation 17fb3d7f-9744-436b-8fe3-685d92aaf3ba · outbound

This paper cites Batch policy learning under constraints.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Batch policy learning under constraints

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.935505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.031806Z digest=sha256:2f163333921045319a62c761b0e065480fde1da3755af87e568c869f072bdb54

Observation c00cb2d8-d4c6-46e5-b035-6c9dc9e02144 · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.919106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.035910Z digest=sha256:e031f48dd825fe7dfe22347d23306385a1f58e6932d6a35b0393101fb268d943

Observation 391ffc41-f4b9-468d-a5cc-564f956456eb · outbound

This paper cites Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.039985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.039985Z digest=sha256:f86e48d0d2ce9efa34a2ed7d0203e2b1b316b28769504bc8b0eeb23add63eee7

Observation 7b857664-28ff-4841-a166-1eed0a315c88 · outbound

This paper cites PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.044417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.044417Z digest=sha256:4319fc34f07ea4d86c16118ccdb249e60f3973636eed812f30e133be67028c07

Observation 82fd9920-06d9-4eac-9227-6fc5976b813a · outbound

This paper cites Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.048977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.048977Z digest=sha256:4c9cb418fdd28906e46297f26744cdf01ed7944ee30f9f4353b2f97bdeba64e0

Observation 53a18ba4-5f14-4d3c-8f35-9dd8a1d5391c · outbound

This paper cites Mildly conservative q-learning for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Mildly conservative q-learning for offline reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.903853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.053729Z digest=sha256:8715f8cbe52c39a927b9cf39a3a64ae4001cd1079da57f4bf434121bc9399c73

Observation 9dfaeb79-4d21-4164-9f6e-3c6368a1aa84 · outbound

This paper cites MOORe: Model-based Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL MOORe: Model-based Offline-to-Online Reinforcement Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:30:49.411742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.057905Z digest=sha256:f1ff310f497e1ffc352dece39e5100c079802733230da3232d93a6a0449bc059

Observation fbd3e492-d978-437d-938e-9749de752dea · outbound

This paper cites Supported trust region optimization for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Supported trust region optimization for offline reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.887854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.062641Z digest=sha256:c41db73a133ad986b38a73d07c33c5baf33b0f701ab71f0e7d0afc6eb75ff239

Observation 1281f0c8-a005-4ff5-b632-a1173f4816a4 · outbound

This paper cites Fine-tuning offline policies with optimistic action selection.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Fine-tuning offline policies with optimistic action selection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.872845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.066875Z digest=sha256:d2215d18704ec723741461475a3f4893da3b86d1397836377e12413214e94ed9

Observation 5d037d44-40f3-4ef0-96c3-9441caeb1c8b · outbound

This paper cites Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:30:49.391005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.071160Z digest=sha256:725faca8c0903bc8fce85f963c24660f1da63e3dd9d342eaea8020a49ea2efb9

Observation 87cd898a-545a-4798-8d34-cfe6b3d63c7c · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.075601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.075601Z digest=sha256:ab0b4088ee9fc12afdd2c3389509fb46315067145100af6438b2409b7776b937

Observation 2e2f4430-3cb6-4f79-ba2e-f0e4e1694785 · outbound

This paper cites Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.080017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.080017Z digest=sha256:f0cb210e1c471cb11eb6cb1fb8144d57667c86acc67a62becb886000b161f791

Observation 77df57cc-3353-487f-a5cf-429697734043 · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Bridging offline reinforcement learning and imitation learning: A tale of pessimism

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.858375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.084656Z digest=sha256:d05d0fb9a9986456f3d1087695de145c6b07236ca41f9a4cb33e0124ecb7ca4a

Observation 16c7240c-2801-434c-9ca7-525d0eecf8fa · outbound

This paper cites Trust region policy optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Trust region policy optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.089965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.089965Z digest=sha256:f7092ecfaa2fdc569d32435bcbc84f2ce217ebd97a6783f02961ed3e609cb838

Observation 3f1115b8-aee4-4409-a4d4-0693a3fdc85d · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.094279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.094279Z digest=sha256:fabd34a000c4c3909708a897efd4d26944b044a8aad5b53f4d5b56bfc1913faf

Observation 3b2e1149-fdc4-4d38-8b3e-976e37b5566f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.098890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.098890Z digest=sha256:f7555dd16d3df0b88846d771d79c19820ed88be5d8bbffa0eb5003194f417ce6

Observation 88530470-1a24-44bc-9604-bad92fa4e204 · outbound

This paper cites Sutton and AndrewG.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Sutton and AndrewG

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.831044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.103172Z digest=sha256:c1ac84fa7146c6e9c9e5668abce6cac78fd591b07c212a208d5b87d9288976a3

Observation 601c7bc8-c6a0-41ef-9765-2c427eebd7dc · outbound

This paper cites Lever- aging factored action spaces for efficient offline reinforcement learning in healthcare.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Lever- aging factored action spaces for efficient offline reinforcement learning in healthcare

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.815625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.107419Z digest=sha256:9398510299e0a0ba17b8de13dc083fd30f1de3a88d995cb8b8c3185a6ad0a917

Observation a778597d-eb97-48f1-a963-2c861bc24650 · outbound

This paper cites CORL: Research-oriented deep offline reinforcement learning library.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL CORL: Research-oriented deep offline reinforcement learning library

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.801226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.111884Z digest=sha256:f5090df4c844fb901967743e50075f5ef09fc9f05fcd7125b5fa4969f8462b1e

Observation 01c0eaf4-3bea-4485-811b-50926caddbbc · outbound

This paper cites Reward Constrained Policy Optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reward Constrained Policy Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.116005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.116005Z digest=sha256:8ae3391bc5f6d8e22333cc1b6c5785a5e7a1e13a3591bfa8ec43abf3132de898

Observation d2f83805-5139-4b0d-a893-8b0679a025e8 · outbound

This paper cites Jump-start reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Jump-start reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.786819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.120558Z digest=sha256:e815622b7cc923a9271c07ab6b550b53f8d0d2966c5e1cf115e5a9a9eca6455c

Observation d77c9109-c886-4784-b58a-52e7e42c26cf · outbound

This paper cites Train once, get a family: State-adaptive balances for offline-to- online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Train once, get a family: State-adaptive balances for offline-to- online reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.772852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.125050Z digest=sha256:279eeb9a2470979090ce222c4bf497767e24d2632ac452511c8b0ee84fd65ce6

Observation 7cafb6e8-c05c-48d4-ba66-1963ea3253e7 · outbound

This paper cites Supported policy optimization for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Supported policy optimization for offline reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.758387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.129229Z digest=sha256:82e43dba798cbf4fe7f2a18a696ac22c3f32eaa529f5eff3ed8ef0d4f9421f89

Observation d7ec5233-1756-45db-b3f8-afee46f24507 · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.133453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.133453Z digest=sha256:e9c0b6200858b20c4097e8839c02d816eee17398ea1619af1cc11f505a8da89f

Observation 7759f3da-00c4-4d39-b472-4b77da252b4d · outbound

This paper cites A policy-guided imitation approach for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A policy-guided imitation approach for offline reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.733631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.137747Z digest=sha256:6e3735113b3a3bdae17b866bb93e3dd62822c3c3e6c2076a726a38be742930b0

Observation 20d6d736-a044-495b-8099-7d0881f00018 · outbound

This paper cites Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.142024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.142024Z digest=sha256:27494a118934b7f73c5b636b63bd443f5ea036fa02cbd78908189379d2e1094f

Observation 7bafb258-e4a9-481d-8b7f-b1ff5d9a276c · outbound

This paper cites Actor-critic alignment for offline-to-online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Actor-critic alignment for offline-to-online reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.717995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.146683Z digest=sha256:ebfa22638102a61560aa4f5adb6a1736bca7e7e7f7924d2edfd90b57aa27cfc8

Observation 4058ab69-5d44-424e-977b-1bd44d22b625 · outbound

This paper cites Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:30:49.280625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.151220Z digest=sha256:926d9cd09b09cf96d92d6a50f93d30d1ff2292f23ecc9d9ef8e9e297c9700cef

Observation f05126e9-b6dd-412d-9435-c2f8ec460bd6 · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.155903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.155903Z digest=sha256:d0b1a385e7deb51a03345b3d27be5b575ba88060709a639683f1282953cece5b

Observation 36075a83-cc30-4203-aa73-eade42be6aea · outbound

This paper cites ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.160669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.160669Z digest=sha256:f3d26c18eff0e623a4004f71ac42e60a6e7d5868d7247a9f110ed8fdeea3ac3a

Observation f1c5647e-6306-4562-ab2d-7b441ae543d8 · outbound

This paper cites Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.165225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.165225Z digest=sha256:7f35498e37030f879f4b3218983975bdf312f4d0c67dd1d4d50e516a9d575a7c

Observation 652d11f2-dc55-47c1-92c3-2f48f7e03ee6 · outbound

This paper cites Online decision transformer.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Online decision transformer

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:30:49.702843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.170538Z digest=sha256:c1d07feb898d6d5ffdfd3a112409a6137618dc59a72bf886eae4a41a0afdbd55

Observation 6a610778-6fdc-45ac-94b5-8d919abe8ac1 · outbound

This paper cites However, our O2SAC still outperforms it with less computational cost during online fine-tuning and less requirements for offline policy.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL However, our O2SAC still outperforms it with less computational cost during online fine-tuning and less requirements for offline policy

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.688112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.175429Z digest=sha256:9abc3a0e7905bad944eb4d6cfed00d9f69f63d275955d5730c09ce7c0b56c762

Observation 2de3f5ca-53a8-465d-90c2-41cd7ee63c22 · outbound

This paper cites For model-based O2O RL, [31] explores regions with high uncertainty and returns in learned model.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL For model-based O2O RL, [31] explores regions with high uncertainty and returns in learned model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.673069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.182233Z digest=sha256:f5182f31df74653208f03f9134bfdbc7f64fa8b2fafbbdbccfbbe173610888fe

Observation 292accf8-fd76-46a6-88f4-c3b9448ede9b · outbound

This paper cites Moreover, [4] find that LayerNorm is favourable for efficient online RL with offline data.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Moreover, [4] find that LayerNorm is favourable for efficient online RL with offline data

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.658434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.187659Z digest=sha256:935391d51176038b46e356ff6b9282154c536a719bb52bfc080d35757e320d05

Pith citing papers

No inbound Pith citation observations are available.