Pith. sign in

Paper Citation Record · LEDGER

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL

As of 12 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2412.18855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18855 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:30:49.187659Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact3
  • verified fuzzy22
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3be00b87-80c9-4e0d-83fb-77693bf1b357 · outbound

This paper cites Constrained policy optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Constrained policy optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.935676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.935676Z digest=sha256:9a16b979fe6c5204987d1c4ab68d932d064ad2739f0f5fa9b62c6afddb7bfde7

Observation b8e8779d-b6d3-40c8-8977-23a777fb2748 · outbound

This paper cites Reinforcement learning: Theory and algorithms.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reinforcement learning: Theory and algorithms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.087909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:48.941106Z digest=sha256:a43965c54692b8b83236c7f53b1bc4ac435b2b7e96ac472670771704cf30f58d

Observation c1eb6107-27ea-47da-b206-13c9cf8efe86 · outbound

This paper cites Reincarnating reinforcement learning: Reusing prior computation to accelerate progress.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reincarnating reinforcement learning: Reusing prior computation to accelerate progress

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.071758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:48.946054Z digest=sha256:f37d44dfdada82426ec6c95010f06d7dbf392ec71c19070ea71702657e24f90b

Observation aeba181c-5585-4ce2-88a1-953ce5a8e638 · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Efficient Online Reinforcement Learning with Offline Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.950682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.950682Z digest=sha256:b4c89599ac55d01d42e512b6cf598d149b1709d723a5880256a15bcee7592341

Observation f57d225d-8fdb-48d5-9a7b-b1e140eff677 · outbound

This paper cites Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.955499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.955499Z digest=sha256:265a6742dc917bab2d4cb6e9d63c458436ff98ccc6050a7a5f4263ec377a228a

Observation 931c0e32-c4c3-48b9-a8da-f1bc7699846c · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Decision transformer: Reinforcement learning via sequence modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.960378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.960378Z digest=sha256:6283a7d6ead359e9ef8fa46202eecb48398d5d8967226ab97d59499c8d1cb55c

Observation 8d97be98-45b7-48a3-8e26-6d41cd6f76c7 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Safe Exploration in Continuous Action Spaces

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.965286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.965286Z digest=sha256:038961e291f1b9d4e987448022ca287c841ec552a5acb831234aae46742f58a0

Observation 16a0bb77-79d9-41f5-bf6e-e2a850b4f462 · outbound

This paper cites Uncertainty-aware model-based offline reinforcement learning for automated driving.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uncertainty-aware model-based offline reinforcement learning for automated driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.047352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:48.969948Z digest=sha256:c794b06e1ca1f278f8f34f0932b340849d70019cdd98792a7403da7efb87364f

Observation 1be9be94-cf64-49a2-a916-7783e452b959 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.974180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.974180Z digest=sha256:cbf3475dc072688ed7107ec3547d9994fe88bc4bc7d77e5702556835cb47a193

Observation 3b822ef4-6f86-4a48-95c5-9a5f7ddfd70f · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A minimalist approach to offline reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.978973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.978973Z digest=sha256:df8989336a13f3a6ff9815e07b12f6993d40426c07bcfc9ee41c621afa4f6e80

Observation 26167c32-88bf-4702-b62e-2c88e1012874 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Addressing function approximation error in actor-critic methods

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.983533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.983533Z digest=sha256:3c8fcdcd64bb35041f71131eada082466117449ede5267202c57239b011ed982

Observation 3e884ace-2ae7-4635-b915-da21ba897fbc · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Off-policy deep reinforcement learning without exploration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.987827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.987827Z digest=sha256:b43002d58c518fa207f040ae42e47913c3dcde94e268c8720bbd730fa436ca36

Observation 94ccdcad-0bf9-47a6-ab8c-1e72494dfeba · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.991946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.991946Z digest=sha256:f8d83e753d85971cdac6d9ee4e79e34cb864f72b342d947a9d6a19c24cb8b956

Observation 1078a651-dda5-442a-99a1-292e83de04f2 · outbound

This paper cites Open and real-world human-ai coordination by heterogeneous training with communication.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Open and real-world human-ai coordination by heterogeneous training with communication

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.001318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:48.997192Z digest=sha256:4af06eff036f69a1263010908e65ed5ace9f65390db070e576a7eb382282d7a1

Observation 791cd18e-9715-4197-8547-41090e436c30 · outbound

This paper cites A simple unified uncertainty-guided framework for offline-to-online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A simple unified uncertainty-guided framework for offline-to-online reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.001554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.001554Z digest=sha256:d20c255a0de7e654aa2a89813492f4ecdcedc3fb916c11bf845c095edf638bcc

Observation a2a59874-9d4a-42a2-8e17-a87128268627 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reinforcement learning with deep energy-based policies

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.005983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.005983Z digest=sha256:50ad45d0f50f5c9649394b08348bf77bca136e2133c43c236c10fac83a0b4f4b

Observation 680d33f4-22a6-4f31-8ff5-940f5489d238 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.010070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.010070Z digest=sha256:9c5fc83c90340b6bbf2d5d8b9d34739e0288c8918840b38ac177f42696e6757d

Observation f919f741-e582-43b1-9b2d-b1242342896c · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.014463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.014463Z digest=sha256:f71b7553ff7356ec750844a6e47e83fb9a931536d8c64144231c69677678e634

Observation 516b2635-e169-418f-987d-e79741e2376e · outbound

This paper cites Uncertainty-driven pessimistic q-ensemble for offline-to- online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uncertainty-driven pessimistic q-ensemble for offline-to- online reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.969216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.019462Z digest=sha256:9e114c3c1d3a90346b68e8ede2a0f51bc52228f42b1ad698a2ed34d953bd8ef1

Observation 64af4ca2-64c0-40f6-9c16-df69ed7322cc · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline reinforcement learning with implicit q-learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.023652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.023652Z digest=sha256:95b60d185a52cf10261b4e1728b35cb532160416a6cb2dbf0649a786b0a6d9b0

Observation 6746a8bb-fd32-4958-a3e1-2deb673aa3ab · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Conservative q-learning for offline reinforcement learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.027727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.027727Z digest=sha256:33f5f58ca440c56fcb6e6273fa373310f2ff48b51369b2eb176caa89ac47f3ae

Observation 17fb3d7f-9744-436b-8fe3-685d92aaf3ba · outbound

This paper cites Batch policy learning under constraints.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Batch policy learning under constraints

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.935505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.031806Z digest=sha256:8352194018b4bff8acd17cde4799c8e4cc7a987e2038b411621f21cd0f507863

Observation c00cb2d8-d4c6-46e5-b035-6c9dc9e02144 · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.919106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.035910Z digest=sha256:a74475a14c03a0cf2cfedf487ca5f6407f62b08779fbe6cb2e71a308cab9f18d

Observation 391ffc41-f4b9-468d-a5cc-564f956456eb · outbound

This paper cites Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.039985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.039985Z digest=sha256:93fa317c49a133fcded63eb4b1a701f57ca05de4828c02da1bf50f1310f733b9

Observation 7b857664-28ff-4841-a166-1eed0a315c88 · outbound

This paper cites PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.044417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.044417Z digest=sha256:a2204db25b12cf5e483f564ac5259b5b1013efdf76fdcd5cf18eeef122e6f345

Observation 82fd9920-06d9-4eac-9227-6fc5976b813a · outbound

This paper cites Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.048977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.048977Z digest=sha256:a43aa55631c8ad9c0b5f426a475eb592c0e3477f9809efafd86052b8405617af

Observation 53a18ba4-5f14-4d3c-8f35-9dd8a1d5391c · outbound

This paper cites Mildly conservative q-learning for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Mildly conservative q-learning for offline reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.903853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.053729Z digest=sha256:b002b76bc34f4354a152d3fc02006d396fe1f3da95213ae273747d989ef86d26

Observation 9dfaeb79-4d21-4164-9f6e-3c6368a1aa84 · outbound

This paper cites MOORe: Model-based Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL MOORe: Model-based Offline-to-Online Reinforcement Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:30:49.411742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.057905Z digest=sha256:a4365ae876ad60c919fdff8cd119a32fba48c419b06a4e66284cb42142f69625

Observation fbd3e492-d978-437d-938e-9749de752dea · outbound

This paper cites Supported trust region optimization for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Supported trust region optimization for offline reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.887854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.062641Z digest=sha256:6f60984a988a123376d5312122e7d67b7baa4622130bd033f889698147ff99da

Observation 1281f0c8-a005-4ff5-b632-a1173f4816a4 · outbound

This paper cites Fine-tuning offline policies with optimistic action selection.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Fine-tuning offline policies with optimistic action selection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.872845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.066875Z digest=sha256:db9b31ad558a5df676cf7b2f18256a2ba8539b408ee44f533b2d21f3bcfb16b7

Observation 5d037d44-40f3-4ef0-96c3-9441caeb1c8b · outbound

This paper cites Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:30:49.391005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.071160Z digest=sha256:6fcebb1c1cd2944f36da8716d40b476c627aa03f2a7575e5913784e009804ee9

Observation 87cd898a-545a-4798-8d34-cfe6b3d63c7c · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.075601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.075601Z digest=sha256:366cd5dc5c7ba825d25787da88dfa3f28a91a31d284a608cc9cfaceba7523ee3

Observation 2e2f4430-3cb6-4f79-ba2e-f0e4e1694785 · outbound

This paper cites Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.080017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.080017Z digest=sha256:f163f462198097661ccd7060b94e0641e99f326a785d26393d17e78537f48532

Observation 77df57cc-3353-487f-a5cf-429697734043 · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Bridging offline reinforcement learning and imitation learning: A tale of pessimism

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.858375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.084656Z digest=sha256:ca2f229240cbd3ef36024b2ddcb3b70e444decf310922de7bbdea9cea1392929

Observation 16c7240c-2801-434c-9ca7-525d0eecf8fa · outbound

This paper cites Trust region policy optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Trust region policy optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.089965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.089965Z digest=sha256:49cb338ed977403564caa39672fe8aa1e63b78c791a171163cf49cbe2ebe91e0

Observation 3f1115b8-aee4-4409-a4d4-0693a3fdc85d · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.094279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.094279Z digest=sha256:c4f2b2bd1dd8186497a6ae47a584851d75564447970130c26d5279448614028e

Observation 3b2e1149-fdc4-4d38-8b3e-976e37b5566f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.098890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.098890Z digest=sha256:898f396c66c1ba29310791fa25f16ad54402e8fe73fada41a91a6ace792246c6

Observation 88530470-1a24-44bc-9604-bad92fa4e204 · outbound

This paper cites Sutton and AndrewG.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Sutton and AndrewG

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.831044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.103172Z digest=sha256:d88223588e06b4f87237f08a2e7170e15102e66c54bd5942c93a91a7c08ba726

Observation 601c7bc8-c6a0-41ef-9765-2c427eebd7dc · outbound

This paper cites Lever- aging factored action spaces for efficient offline reinforcement learning in healthcare.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Lever- aging factored action spaces for efficient offline reinforcement learning in healthcare

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.815625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.107419Z digest=sha256:823f6783a6882d57ccd36dd864fb883b482ce3510439b5571505600c23209878

Observation a778597d-eb97-48f1-a963-2c861bc24650 · outbound

This paper cites CORL: Research-oriented deep offline reinforcement learning library.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL CORL: Research-oriented deep offline reinforcement learning library

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.801226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.111884Z digest=sha256:81348a55c7fdbf6315d1e67dcfcf13bab739f3ef10cab1576df92a6a3a94e44c

Observation 01c0eaf4-3bea-4485-811b-50926caddbbc · outbound

This paper cites Reward Constrained Policy Optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reward Constrained Policy Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.116005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.116005Z digest=sha256:b27d1453c480231cb3d781dd4b4d6e9b5f2e277c7346be70f40df3469eb01922

Observation d2f83805-5139-4b0d-a893-8b0679a025e8 · outbound

This paper cites Jump-start reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Jump-start reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.786819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.120558Z digest=sha256:e033fa20ed7358ca4cfd400b230e24d40ce0321be912ac5f3fbb9b780a82647f

Observation d77c9109-c886-4784-b58a-52e7e42c26cf · outbound

This paper cites Train once, get a family: State-adaptive balances for offline-to- online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Train once, get a family: State-adaptive balances for offline-to- online reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.772852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.125050Z digest=sha256:cd78f199f46cb8355c7a5ef51996bebf47aadba6028df7c3244d3c69a116e861

Observation 7cafb6e8-c05c-48d4-ba66-1963ea3253e7 · outbound

This paper cites Supported policy optimization for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Supported policy optimization for offline reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.758387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.129229Z digest=sha256:31eb68beadf96432d71e515acb4ade1c7e4ef965bdba95a801244708d76635fa

Observation d7ec5233-1756-45db-b3f8-afee46f24507 · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.133453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.133453Z digest=sha256:28e959e8df34c4c4b3c7fefb794edf34f13d2ba9090e97b62a41438841f31401

Observation 7759f3da-00c4-4d39-b472-4b77da252b4d · outbound

This paper cites A policy-guided imitation approach for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A policy-guided imitation approach for offline reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.733631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.137747Z digest=sha256:88620ffe87d49be3c9eefb3483ceb72b2fbd1f03113eca2fa7257ae116d6d82e

Observation 20d6d736-a044-495b-8099-7d0881f00018 · outbound

This paper cites Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.142024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.142024Z digest=sha256:dea029edadaea6c3252d095c20de3f0acc44052c0d16ce4104310a19d494403c

Observation 7bafb258-e4a9-481d-8b7f-b1ff5d9a276c · outbound

This paper cites Actor-critic alignment for offline-to-online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Actor-critic alignment for offline-to-online reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.717995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.146683Z digest=sha256:5d73f2b0a274873c60d0e420e3a18c7f29bf8c711da7725ac2a0ecc80b719659

Observation 4058ab69-5d44-424e-977b-1bd44d22b625 · outbound

This paper cites Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:30:49.280625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.151220Z digest=sha256:17aea06237445cc695d32673ad9f8d6277206de9a26c7698a2e3fb857a382a65

Observation f05126e9-b6dd-412d-9435-c2f8ec460bd6 · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.155903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.155903Z digest=sha256:5dd7781fbbfba7d69cce4751a31f0156db9265b4d54d086b4ced034fb23dd97b

Observation 36075a83-cc30-4203-aa73-eade42be6aea · outbound

This paper cites ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.160669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.160669Z digest=sha256:4b63cea41fd205b59701ace40e4cab97d57121f6c92c37be324d5421da211459

Observation f1c5647e-6306-4562-ab2d-7b441ae543d8 · outbound

This paper cites Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.165225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.165225Z digest=sha256:5c5d8cc55750bc06a9cac1f202c8e677f5a8e1bc0be3c0f2f0e4599c5cd2e38e

Observation 652d11f2-dc55-47c1-92c3-2f48f7e03ee6 · outbound

This paper cites Online decision transformer.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Online decision transformer

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:30:49.702843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.170538Z digest=sha256:51e8c3359afcc880e05e71e573ddef8013dd005f1372ba5cce55912bdf633ed2

Observation 6a610778-6fdc-45ac-94b5-8d919abe8ac1 · outbound

This paper cites However, our O2SAC still outperforms it with less computational cost during online fine-tuning and less requirements for offline policy.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL However, our O2SAC still outperforms it with less computational cost during online fine-tuning and less requirements for offline policy

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.688112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.175429Z digest=sha256:71589a43f73c568cf2c1b1ba9cd491d1b5742e899e44190864e995cf05ed0f74

Observation 2de3f5ca-53a8-465d-90c2-41cd7ee63c22 · outbound

This paper cites For model-based O2O RL, [31] explores regions with high uncertainty and returns in learned model.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL For model-based O2O RL, [31] explores regions with high uncertainty and returns in learned model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.673069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.182233Z digest=sha256:6dce6a77962a75b7dab622834b7a4c3a9ebcb92edf62f58922f52bba8424fffa

Observation 292accf8-fd76-46a6-88f4-c3b9448ede9b · outbound

This paper cites Moreover, [4] find that LayerNorm is favourable for efficient online RL with offline data.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Moreover, [4] find that LayerNorm is favourable for efficient online RL with offline data

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.658434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:30:49.187659Z digest=sha256:d701b165ad9eb748f75bd0cc3cf4625ac332be4edd91d03db45a1742914e29a0

Pith citing papers

No inbound Pith citation observations are available.