Pith. sign in

Paper Citation Record · LEDGER

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL

As of 23 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.22401.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22401 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:16:41.636528Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04d6714c-a09f-401c-8705-c4a706c8fd77 · outbound

This paper cites A policy π : S 7→∆(A) specifies an action selection rule, where π(a|s) specifies the probability of taking action a in state s for each (s, a) ∈ S × A.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL A policy π : S 7→∆(A) specifies an action selection rule, where π(a|s) specifies the probability of taking action a in state s for each (s, a) ∈ S × A

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:16:42.380764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:16:41.636528Z digest=sha256:78404f29234e2731f886ced0a8cf4f61567efd9630ac62343cdeba2dad324f21

Observation 9dbf8921-89fd-426a-bcc4-b369363ac640 · outbound

This paper cites (40) Here, H(·) is the entropy function satisfying 0 ⩽ H(p) ⩽ log |A|, ∀p ∈ ∆(A).

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL (40) Here, H(·) is the entropy function satisfying 0 ⩽ H(p) ⩽ log |A|, ∀p ∈ ∆(A)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:16:42.753911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:16:41.516439Z digest=sha256:c8e651b465c8c1da4bad5cd81f5fe54e4c1e99828d9425573fb261785f661ba7

Observation 7418e6c8-5dc5-4ab3-98e0-36583c608810 · outbound

This paper cites Imitation Learning via Off-Policy Distribution Matching.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Imitation Learning via Off-Policy Distribution Matching

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:40.577492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:40.577492Z digest=sha256:503dad92e25c5d44ba0de280adea784ae057d1dd0adb5f50bba840fe6d5179f5

Observation 25f56446-c572-4bca-9973-944118a22325 · outbound

This paper cites Assumption 5 (Policy class).

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Assumption 5 (Policy class)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:16:42.508330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:16:41.589826Z digest=sha256:3bd00ede30286426faa482d3723d5041f02d20ae550e1efc718ebb0601cd51f7

Observation a90d7736-e7ad-4fdf-8ab1-90be08d0b6d5 · outbound

This paper cites GPT-4 Technical Report.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL GPT-4 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:40.863640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:40.863640Z digest=sha256:f4ee5c45aea616779181ec2993cde2f6d1f298cb765e1e35d09db6f6548086c9

Observation eb877a13-5209-4cd4-9e20-482f5adfc989 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Gemini: A Family of Highly Capable Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:41.087989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:41.087989Z digest=sha256:3b2707bf0e5029d62c1e8521900fe75f438cbc3e8baa240882bb12e191359720

Observation 6e8bf332-c7ab-4cc7-80f6-30b07b7d1439 · outbound

This paper cites Primal-Dual $\pi$ Learning: Sample Complexity and Sublinear Run Time for Ergodic Markov Decision Problems.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Primal-Dual $\pi$ Learning: Sample Complexity and Sublinear Run Time for Ergodic Markov Decision Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:41.136263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:41.136263Z digest=sha256:897e4f213725ad78bfe60a5e76bf45d0d4347ac62bd72932e01a3e89869aeca4

Observation 96767539-3980-4239-b77a-43d95f4b0436 · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:41.164165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:41.164165Z digest=sha256:63423ba148f9e1c6fc55982b6da1966768883fa129de461d25a526604d8a435d

Observation 4f1f6c9b-6aee-4768-be1b-f7710a86285f · outbound

This paper cites Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:16:41.806555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:16:41.207446Z digest=sha256:aa77b87176176c02c769ba6a3829dcb6ff4b67ac689f7df8949b0a5662fc061e

Observation 10127970-7d4e-49af-b5ff-daac192413ef · outbound

This paper cites Zanette, D.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Zanette, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:16:42.966818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:16:41.279713Z digest=sha256:3b525a51a022c412b5fd79b6149286e21af2b2b1bc3ba3825142082ccb427bb1

Observation 9fac93e6-af53-45fd-9ed7-fd92bd07a6d4 · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:41.348358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:41.348358Z digest=sha256:6176e8b0ca4144580e7e58fbcc411e66059798041be1dd019a71839b3a4f1e59

Observation 83d01094-2d25-4e16-962b-af20816a1a09 · outbound

This paper cites GEC: A Unified Framework for Interactive Decision Making in MDP, POMDP, and Beyond.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL GEC: A Unified Framework for Interactive Decision Making in MDP, POMDP, and Beyond

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:41.423136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:41.423136Z digest=sha256:f32000c0f825c4a5b01f2535f46ca582590efb10298bde59be5e7358a982ea1b

Observation 5b3b281a-35fc-4323-9584-723b350d38b8 · outbound

This paper cites Lemma 1 (Freedman’s inequality, Lemma D.2 in Liu et al.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Lemma 1 (Freedman’s inequality, Lemma D.2 in Liu et al

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:16:42.879115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:16:41.475713Z digest=sha256:644bd4cc5819f249e462e7d206dc83166b40f177ae0e02def296a693bfb1a0b7

Observation 02be0b09-c698-4362-8ffc-b6ac32f2e126 · outbound

This paper cites (113) This gives the desired result.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL (113) This gives the desired result

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:16:42.625708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:16:41.557045Z digest=sha256:1247f60bf918d467fda1415257e86c7d15d068274bdfb9cf2fa6d8b6346cc479

Observation fa730b0c-dc34-43b0-b04b-20de32495e77 · outbound

This paper cites Actor-Critics Can Achieve Optimal Sample Efficiency.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Actor-Critics Can Achieve Optimal Sample Efficiency

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:41.051490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:41.051490Z digest=sha256:995589b400aefb0288e2b14f09920623d6948f1ceb0822f2bbc00c43abe5e06f

Observation 6295c50d-8ffc-42ef-bed6-8beab192319b · outbound

This paper cites Spectral Decomposition Representation for Reinforcement Learning.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Spectral Decomposition Representation for Reinforcement Learning

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:40.943404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:40.943404Z digest=sha256:37b5f3cc6276e37849468ea5f0f135467829f6dcbf459e1401ce262d94da36a6

Observation f1f62133-3519-4288-a00b-dea1468a5493 · outbound

This paper cites Moulin, G.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Moulin, G

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:40.674415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:40.674415Z digest=sha256:00c2ad2d80dd0d1db3f735fad9c132d81d943211dcf0f2c184c10fc0724616ec

Observation 4a0cc76b-5241-46e8-9451-e38e9786a534 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:40.137931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:40.137931Z digest=sha256:dc9cdf6c53d06fddab8053f87cb06b444a43f83a94086e85a34fbb0aecd07dd7

Observation 04308499-b75f-4399-8e8c-76d18a719431 · outbound

This paper cites Dual RL: Unification and New Methods for Reinforcement and Imitation Learning.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Dual RL: Unification and New Methods for Reinforcement and Imitation Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:41.004427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:41.004427Z digest=sha256:c9fcf464fabd0c065f3933283e8cdd8919e6ffd02509762923d4fe659aa5ed59

Observation 98ac1154-60b2-4318-8992-0d33fe16af9f · outbound

This paper cites AlgaeDICE: Policy Gradient from Arbitrary Experience.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL AlgaeDICE: Policy Gradient from Arbitrary Experience

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:40.782168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:40.782168Z digest=sha256:c6ca756f21ec647155b718b5002acfc4c93ae810ca9eef48ac143c0172933b52

Observation b92eef3e-3625-48b6-9da3-4c3413137018 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:40.421569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:40.421569Z digest=sha256:868e27e70c425c9f2fffbb7135deee3d55b653143942f224aa5af66e7cfae92e

Observation 514554d3-7a2a-4049-bfd0-f334b90c48a4 · outbound

This paper cites The Statistical Complexity of Interactive Decision Making.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL The Statistical Complexity of Interactive Decision Making

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:40.263471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:40.263471Z digest=sha256:cc6d30bfef28587d21a63a99986a188012721d5d7f678c49eab183aa3421f908

Observation 349dbb18-227f-4524-b75a-bd9eba90e993 · outbound

This paper cites Posterior sampling for reinforcement learning: worst-case regret bounds.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Posterior sampling for reinforcement learning: worst-case regret bounds

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:16:42.253497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:16:40.034564Z digest=sha256:06243146b024f4fc4c2a17948f182aade77fa1d8a016910ea2656dce31e714f4

Observation 662386d8-9299-4b4a-abb8-3cd756b6a977 · outbound

This paper cites On the Power of Multitask Representation Learning in Linear MDP.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL On the Power of Multitask Representation Learning in Linear MDP

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:40.625859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:40.625859Z digest=sha256:7b8ffd5cd9ea8c3cbe55f383f7a9cec288f2624c0976d6ae80f3246428a5775b

Observation 0d4a63e2-91b3-4398-b508-2a2dfbd53315 · outbound

This paper cites Reinforcement Learning via Fenchel-Rockafellar Duality.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Reinforcement Learning via Fenchel-Rockafellar Duality

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:40.717812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:40.717812Z digest=sha256:1e54ea78556956a433b947411ce3bdb5ed925be8c6094af69f2736001e1273e8

Pith citing papers

No inbound Pith citation observations are available.