Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2006.04779.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:45:34.852244Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
537
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 798f5a8b-a8f1-46a2-afda-7e554d96674b · inbound
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation Conservative Q-Learning for Offline Reinforcement Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 114f4217-f920-4a89-86b6-529d4ba6016a · inbound
Offline Reinforcement Learning with Implicit Q-Learning Conservative Q-Learning for Offline Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 918ed758-8c20-49ea-bc32-31db1ae7d66e · inbound
DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions Conservative Q-Learning for Offline Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e8e4e2d1-bdee-42ff-b814-21d6f5458937 · inbound
Value Flows Conservative Q-Learning for Offline Reinforcement Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 593a8806-a6e6-448e-b68c-492efa911c48 · inbound
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Conservative Q-Learning for Offline Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 18072b28-e076-4023-97dd-5c4a920c8b40 · inbound
The hidden risks of temporal resampling in clinical reinforcement learning Conservative Q-Learning for Offline Reinforcement Learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df39f09e-07b6-4ef8-8e75-5515c23b8209 · inbound
Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation Conservative Q-Learning for Offline Reinforcement Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 771a535d-9a5a-41fb-b930-f2cd098e19ff · inbound
JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing Conservative Q-Learning for Offline Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13352aff-ad43-46fa-9169-02517c493e7d · inbound
JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing Conservative Q-Learning for Offline Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aeb8997-d922-4e50-aa27-2de53b38dbee · inbound
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture Conservative Q-Learning for Offline Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f1446329-990b-4a11-bbe1-7ec26ce075f7 · inbound
An adaptive variance estimator for relative sparsity Conservative Q-Learning for Offline Reinforcement Learning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54b2b2da-a21f-40ea-bfd8-cebf495def48 · inbound
RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Conservative Q-Learning for Offline Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ff6e96e-8a17-48c8-9446-e719ce29eaaa · inbound
RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Conservative Q-Learning for Offline Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 363088f6-84cb-4d4c-a516-d248c8a40dcf · inbound
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation Conservative Q-Learning for Offline Reinforcement Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 43d4de9f-b2d2-4358-b1d2-285c45147a1c · inbound
ISEP: Implicit Support Expansion for Offline Reinforcement Learning via Stochastic Policy Optimization Conservative Q-Learning for Offline Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 91af34ea-1812-483d-9e44-00cd4875c82b · inbound
Abstraction for Offline Goal-Conditioned Reinforcement Learning Conservative Q-Learning for Offline Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c1e04c5-22b8-4520-ac8b-0185d71fed66 · inbound
Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization Conservative Q-Learning for Offline Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96215305-918f-42b4-9869-48c085c7f94f · inbound
Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience Conservative Q-Learning for Offline Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d5013757-8dfe-4512-a595-49fb8945f0c8 · inbound
From Bootstrapping to Sequence Modeling: A Unified Generative Framework for Personalized Landing-Page Modeling Conservative Q-Learning for Offline Reinforcement Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb1a7742-8efb-4993-a66b-a3fba610e1ad · inbound
Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models Conservative Q-Learning for Offline Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 09c7d9f4-c115-4b28-88b4-ff7706ff3381 · inbound
Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies Conservative Q-Learning for Offline Reinforcement Learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 43654418-5787-417f-b0a6-d550dd7cc28f · inbound
Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies Conservative Q-Learning for Offline Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b74f09b0-d2d0-424b-980b-f5c4a16560f2 · inbound
Reinforcement Learning: From Algorithms To Foundation Models Conservative Q-Learning for Offline Reinforcement Learning
Reference 172
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad055bab-729a-4f67-bdfa-70a49a77e3a4 · inbound
Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search Conservative Q-Learning for Offline Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536637ce-d99c-4c05-b02d-56f7c329c78d · inbound
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Conservative Q-Learning for Offline Reinforcement Learning
Reference 126
Source-reported events for the cited work
Unavailable: canonical work link unavailable.