Pith. sign in

Paper Citation Record · LEDGER

Dual RL: Unification and New Methods for Reinforcement and Imitation Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2302.08560.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.08560 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:16:41.004427Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T22:15:49.901050Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 258ea22a-5c21-487d-aaff-13b1f6eecba0 · inbound

TRAM: Test-Time Risk Adaptation with Mixture of Agents cites this paper.

TRAM: Test-Time Risk Adaptation with Mixture of Agents Dual RL: Unification and New Methods for Reinforcement and Imitation Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:15:49.904188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T22:15:03.665956Z digest=sha256:ec37ee0453ced8fa6e9cae52da2a24588f921dc8500d24c21b2ad57e004e7337

Observation 04308499-b75f-4399-8e8c-76d18a719431 · inbound

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL cites this paper.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Dual RL: Unification and New Methods for Reinforcement and Imitation Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:41.004427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:41.004427Z digest=sha256:c529ef4556e1508b22fa7f1d5b609d6b69be2f72116beead86e3369380ade2f9

Observation 2b55bd79-908e-40f9-bdd9-072f03d2af0a · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Dual RL: Unification and New Methods for Reinforcement and Imitation Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:46.615368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:32525fcc865378b8207fbd1caa9329895845e74ebb7217d9cf414c4a5eda9859

Observation 70989f6e-e878-4c1b-a5ec-6c1d25947041 · inbound

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning cites this paper.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Dual RL: Unification and New Methods for Reinforcement and Imitation Learning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:47.200049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:f123943c393fb98d25161491acbff6ab52c0e62911db2a46d6cc0a117ec32ab1

Observation d84b81c6-1e4a-497c-b012-1b16eb4ff24b · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Dual RL: Unification and New Methods for Reinforcement and Imitation Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:00.642393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:96b44e642d2d35b6bd60e881771112f0f62f2b5cda9c9a5731df6399b75d286f

Observation abc6a6f2-c3ef-4997-ac10-30503af28278 · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Dual RL: Unification and New Methods for Reinforcement and Imitation Learning

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:58.145612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:53cbfe77ed3070945e3acb9bbf186417117217a4b03109e59398b47ce42fe8ef