Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:16:41.636528Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.22401.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:16:41.636528Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 04d6714c-a09f-401c-8705-c4a706c8fd77 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL A policy π : S 7→∆(A) specifies an action selection rule, where π(a|s) specifies the probability of taking action a in state s for each (s, a) ∈ S × A
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9dbf8921-89fd-426a-bcc4-b369363ac640 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL (40) Here, H(·) is the entropy function satisfying 0 ⩽ H(p) ⩽ log |A|, ∀p ∈ ∆(A)
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7418e6c8-5dc5-4ab3-98e0-36583c608810 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Imitation Learning via Off-Policy Distribution Matching
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f56446-c572-4bca-9973-944118a22325 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Assumption 5 (Policy class)
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a90d7736-e7ad-4fdf-8ab1-90be08d0b6d5 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL GPT-4 Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb877a13-5209-4cd4-9e20-482f5adfc989 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Gemini: A Family of Highly Capable Multimodal Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e8bf332-c7ab-4cc7-80f6-30b07b7d1439 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Primal-Dual $\pi$ Learning: Sample Complexity and Sublinear Run Time for Ergodic Markov Decision Problems
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96767539-3980-4239-b77a-43d95f4b0436 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f1f6c9b-6aee-4768-be1b-f7710a86285f · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 10127970-7d4e-49af-b5ff-daac192413ef · outbound
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9fac93e6-af53-45fd-9ed7-fd92bd07a6d4 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83d01094-2d25-4e16-962b-af20816a1a09 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL GEC: A Unified Framework for Interactive Decision Making in MDP, POMDP, and Beyond
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b3b281a-35fc-4323-9584-723b350d38b8 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Lemma 1 (Freedman’s inequality, Lemma D.2 in Liu et al
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 02be0b09-c698-4362-8ffc-b6ac32f2e126 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL (113) This gives the desired result
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fa730b0c-dc34-43b0-b04b-20de32495e77 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Actor-Critics Can Achieve Optimal Sample Efficiency
Reference 1998
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6295c50d-8ffc-42ef-bed6-8beab192319b · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Spectral Decomposition Representation for Reinforcement Learning
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f62133-3519-4288-a00b-dea1468a5493 · outbound
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a0cc76b-5241-46e8-9451-e38e9786a534 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04308499-b75f-4399-8e8c-76d18a719431 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98ac1154-60b2-4318-8992-0d33fe16af9f · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL AlgaeDICE: Policy Gradient from Arbitrary Experience
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b92eef3e-3625-48b6-9da3-4c3413137018 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 514554d3-7a2a-4049-bfd0-f334b90c48a4 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL The Statistical Complexity of Interactive Decision Making
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349dbb18-227f-4524-b75a-bd9eba90e993 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Posterior sampling for reinforcement learning: worst-case regret bounds
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 662386d8-9299-4b4a-abb8-3cd756b6a977 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL On the Power of Multitask Representation Learning in Linear MDP
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d4a63e2-91b3-4398-b508-2a2dfbd53315 · outbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Reinforcement Learning via Fenchel-Rockafellar Duality
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.