Pith. sign in

Paper Citation Record · LEDGER

Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2211.11802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.11802 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:45.273426Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:54:05.802664Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 563b4579-5f45-48bf-9108-6c2930b0c1e2 · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.273426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.273426Z digest=sha256:1fb05910c531110635d935af83c6f28c431d230d00e51ef84a3aebb228cab906

Observation fbd0d19b-7554-49b0-bd54-5a6b6af55f83 · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:37:13.800212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T03:36:09.272019Z digest=sha256:bcf4ea3b6c9d754f824643a289526c061a7739bd31371da7161fcc4b7672f2b6

Observation 89260a7d-7d73-4d9d-ad97-af1777d7af74 · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T02:53:14.061010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:53:14.061010Z digest=sha256:adbf7fe2cac2ae32bc51f8413668c5194f921d60c27cfa619ff078989d5ab6ab

Observation 20abb1f2-ffa7-4130-a0a0-f815389ee71b · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:08.177721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:fae9e94ba53f1146c2f5d0ef3b0146ece4e5b08b3cfe2e860a1b6ada29ccb822

Observation 9e4d69d8-4e4d-4576-9fc7-7b3f5de7e2c1 · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.804432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:bf0bfd6aebeaaf70f1492a82a6e23ba274377737d131ce36f8d4bc5cad716bfd

Observation e8b156a5-d68f-4183-94d1-4b025ca79bcd · inbound

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation cites this paper.

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T13:15:10.238096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:15:10.238096Z digest=sha256:87ddf43d16cfa1edae518a54320924d5da5590872a8a9482967757f5c9ceed68