Pith. sign in

Paper Citation Record · LEDGER

Conservative Q-Learning for Offline Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2006.04779.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.04779 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:45:34.852244Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

537
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 798f5a8b-a8f1-46a2-afda-7e554d96674b · inbound

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation cites this paper.

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation Conservative Q-Learning for Offline Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:51:55.904631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T08:51:55.826747Z digest=sha256:32e946e3cba21f5aba1bd1084a69ddaec3f9f9512d31844cc884b21e9eeeabbd

Observation 114f4217-f920-4a89-86b6-529d4ba6016a · inbound

Offline Reinforcement Learning with Implicit Q-Learning cites this paper.

Offline Reinforcement Learning with Implicit Q-Learning Conservative Q-Learning for Offline Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:47:05.661694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:47:05.621624Z digest=sha256:baa2b1d107710a3c5899518d18cd696fef30b260db9c9f5525a15d8c4954a756

Observation 918ed758-8c20-49ea-bc32-31db1ae7d66e · inbound

DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions cites this paper.

DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions Conservative Q-Learning for Offline Reinforcement Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:56:26.127974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T13:55:16.513839Z digest=sha256:fb178f99f3bd3dafd3bdacc7004a7c1ede03816d654c3b595793708334ee9fb2

Observation e8e4e2d1-bdee-42ff-b814-21d6f5458937 · inbound

Value Flows cites this paper.

Value Flows Conservative Q-Learning for Offline Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:31.063875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:31.063875Z digest=sha256:6c1710db3ccbfebc7647226a216542da08bbac106eed5a28725ee6695624c494

Observation 593a8806-a6e6-448e-b68c-492efa911c48 · inbound

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training cites this paper.

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Conservative Q-Learning for Offline Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:48:51.003494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T01:46:21.744857Z digest=sha256:3d4ed53620d45f99c9ce23dcc95fc2415b3671aa0939b7856ba488da5f7b9465

Observation 18072b28-e076-4023-97dd-5c4a920c8b40 · inbound

The hidden risks of temporal resampling in clinical reinforcement learning cites this paper.

The hidden risks of temporal resampling in clinical reinforcement learning Conservative Q-Learning for Offline Reinforcement Learning

Reference 62

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T07:07:29.616897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:05:16.379996Z digest=sha256:12872d4257bfe50344371f2420e644aef5bb9646c2512b891167e166b2be0730

Observation df39f09e-07b6-4ef8-8e75-5515c23b8209 · inbound

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation cites this paper.

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation Conservative Q-Learning for Offline Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:49:54.486151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:49:03.333757Z digest=sha256:e68d8924a4634863028183f2f27c795d6dc28d0e5ef24512c84df5d8968b5743

Observation 771a535d-9a5a-41fb-b930-f2cd098e19ff · inbound

JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing cites this paper.

JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing Conservative Q-Learning for Offline Reinforcement Learning

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:10:52.115427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:40:01.560145Z digest=sha256:c75d5f82925fdff434d8bee31a5ebdaf382caa1c65cc6e6c9ee2d86adcbf233f

Observation 13352aff-ad43-46fa-9169-02517c493e7d · inbound

JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing cites this paper.

JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing Conservative Q-Learning for Offline Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T16:46:55.806545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:46:55.806545Z digest=sha256:b32cbcc354723467b67ae4faa6eff7a2398de526dfcc6a2f6e8bb48b1e6b18c0

Observation 9aeb8997-d922-4e50-aa27-2de53b38dbee · inbound

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture cites this paper.

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture Conservative Q-Learning for Offline Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:29:06.408108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T13:59:40.638085Z digest=sha256:c4d336c7373f3b7073052f5772e835870a22fb16ba7dcee987515533b55ffcbb

Observation f1446329-990b-4a11-bbe1-7ec26ce075f7 · inbound

An adaptive variance estimator for relative sparsity cites this paper.

An adaptive variance estimator for relative sparsity Conservative Q-Learning for Offline Reinforcement Learning

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.591920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T19:35:46.097113Z digest=sha256:49284adc21de0107d8e86ec9dffcf91b423c1c1d34e8c3a6b34fdb28a21c3823

Observation 54b2b2da-a21f-40ea-bfd8-cebf495def48 · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Conservative Q-Learning for Offline Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:08.274944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:775900802de6d25ee7612b780a185966147258dfe1ac36094adf8b3c83a125f7

Observation 6ff6e96e-8a17-48c8-9446-e719ce29eaaa · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Conservative Q-Learning for Offline Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.817616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:93fceba1dd77247e92330bc7542c4caac7707ac01270083c21e4f456faecfb47

Observation 363088f6-84cb-4d4c-a516-d248c8a40dcf · inbound

Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation cites this paper.

Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation Conservative Q-Learning for Offline Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:52:46.236699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T20:49:15.724896Z digest=sha256:f51922ebf62b8e31637ff39aab1bb5b0f10a97cfc28966388713182d57d5ac20

Observation 43d4de9f-b2d2-4358-b1d2-285c45147a1c · inbound

ISEP: Implicit Support Expansion for Offline Reinforcement Learning via Stochastic Policy Optimization cites this paper.

ISEP: Implicit Support Expansion for Offline Reinforcement Learning via Stochastic Policy Optimization Conservative Q-Learning for Offline Reinforcement Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.443271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T12:07:21.037980Z digest=sha256:da124e8ec9837f51aaf22b82f391bd8c57a18a34f1b1fc81ada0f35362f97c1e

Observation 91af34ea-1812-483d-9e44-00cd4875c82b · inbound

Abstraction for Offline Goal-Conditioned Reinforcement Learning cites this paper.

Abstraction for Offline Goal-Conditioned Reinforcement Learning Conservative Q-Learning for Offline Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:16.575399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T07:46:20.289421Z digest=sha256:57acbc37dd3abb67ceb9c46afb9030cf480fb02b9090956ea4be23ae86965448

Observation 1c1e04c5-22b8-4520-ac8b-0185d71fed66 · inbound

Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization cites this paper.

Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization Conservative Q-Learning for Offline Reinforcement Learning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:09:37.550272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T14:48:19.683101Z digest=sha256:5afbf850a8b23c2f2a5fdcdf332750e1fd396ece86d5f5186db6d738c0d74faa

Observation 96215305-918f-42b4-9869-48c085c7f94f · inbound

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience cites this paper.

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience Conservative Q-Learning for Offline Reinforcement Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:58.742568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T01:57:04.293058Z digest=sha256:54e6ea7a4c9e7b9139962bb00f544dfa734b8285508f2a99d62e646d04dd0a9f

Observation d5013757-8dfe-4512-a595-49fb8945f0c8 · inbound

From Bootstrapping to Sequence Modeling: A Unified Generative Framework for Personalized Landing-Page Modeling cites this paper.

From Bootstrapping to Sequence Modeling: A Unified Generative Framework for Personalized Landing-Page Modeling Conservative Q-Learning for Offline Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:55:51.405173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T02:56:52.303152Z digest=sha256:4a6e709e3c74c6375dc004e71250a1e79f6de344deaeab7e160df7bab12d5bea

Observation fb1a7742-8efb-4993-a66b-a3fba610e1ad · inbound

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models cites this paper.

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models Conservative Q-Learning for Offline Reinforcement Learning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:21.012695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T06:46:02.255137Z digest=sha256:53fe48c5bfb99a4765378b39c0ce539f4ee4a800641bcd8224a550ea4475ffab

Observation 09c7d9f4-c115-4b28-88b4-ff7706ff3381 · inbound

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies cites this paper.

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies Conservative Q-Learning for Offline Reinforcement Learning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:58:05.713079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T11:50:49.215046Z digest=sha256:ad4ef43d3adea18da021933f5a03b44812291ff53911ab09513f8e59e00254b6

Observation 43654418-5787-417f-b0a6-d550dd7cc28f · inbound

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies cites this paper.

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies Conservative Q-Learning for Offline Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T08:26:59.428524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:26:59.428524Z digest=sha256:7673fbedaf86883d7e6355ecb30a6473469b5f01f0e22860d95154e3456f4742

Observation b74f09b0-d2d0-424b-980b-f5c4a16560f2 · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Conservative Q-Learning for Offline Reinforcement Learning

Reference 172

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:13.959321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:13.959321Z digest=sha256:58625aa30ffcab2e013d04613fa82619f90de517c510d4ca54e63502b96b0bec

Observation ad055bab-729a-4f67-bdfa-70a49a77e3a4 · inbound

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search cites this paper.

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search Conservative Q-Learning for Offline Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T01:24:20.678068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:24:20.678068Z digest=sha256:24a5e6811322f3fc3772e4c0443c56fc1adc8e5b949d21a2069b67fc8db858e1

Observation 536637ce-d99c-4c05-b02d-56f7c329c78d · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Conservative Q-Learning for Offline Reinforcement Learning

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:34.852244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:34.852244Z digest=sha256:6f53f387a6bf70a33e71ce61393e2310e6cdd82b883d43597a69ca1cd8aa0772