Pith. sign in

Paper Citation Record · LEDGER

Extreme Q-Learning: MaxEnt RL without Entropy

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2301.02328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.02328 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:34:23.962457Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:36.943655Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 575ff729-ed03-4652-b4e6-6f605c574e73 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Extreme Q-Learning: MaxEnt RL without Entropy

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:48:36.476653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:5a3148f607b507ad81094cf8434094d90c6d77821071afacc231211b87292073

Observation 987adb72-3301-44da-a7f6-33ac305439e6 · inbound

Are Expressive Models Truly Necessary for Offline RL? cites this paper.

Are Expressive Models Truly Necessary for Offline RL? Extreme Q-Learning: MaxEnt RL without Entropy

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:12:40.424073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:12:40.424073Z digest=sha256:645e0fe94a96cc59f1b9a527a2cacf0112bd2064bd1c68da969929bae4f72392

Observation 94ccdcad-0bf9-47a6-ab8c-1e72494dfeba · inbound

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL cites this paper.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.991946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.991946Z digest=sha256:8f7cbcca1fbb3ac2a2059e6dcbdbf019c4f59fc95af8c46179a3fd9f2e5ae20d

Observation 8b590690-0025-41e2-9905-07b1a5960926 · inbound

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint cites this paper.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Extreme Q-Learning: MaxEnt RL without Entropy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.228255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.228255Z digest=sha256:c0efc7c8b16c48e550b99357018954e8a927b71faffadd214b54ee2e60f9bb62

Observation c83f43dd-5490-40c7-aa7f-f65ffa84cb46 · inbound

TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion cites this paper.

TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion Extreme Q-Learning: MaxEnt RL without Entropy

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:23.962457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:23.962457Z digest=sha256:258a27adc767b510a145246ce060f0cf855056631cf422d3a244f8c9811e8a53

Observation 202fe6ac-af03-4756-8236-36e4ef3b38c3 · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.632354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.632354Z digest=sha256:9f1a2a0c8a486f296e5d9ae55c4374c025081f7ba0bc20a11a5b5a22186d3e91

Observation 14dd046c-95d9-4c78-9362-9a3b7e400315 · inbound

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood cites this paper.

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood Extreme Q-Learning: MaxEnt RL without Entropy

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:24.560434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:22:24.560434Z digest=sha256:dbfc0fd0cfe7df24672d21c37653d85bef73fdb0c155dceb8e982d2f683f7ee2

Observation 6b556c63-0a75-4c9d-a391-14da58e48501 · inbound

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion cites this paper.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Extreme Q-Learning: MaxEnt RL without Entropy

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.548506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.548506Z digest=sha256:b199c0bfc111b3f332d5909309c9ca20af2f74059b00d1639aba961289cd1c9b

Observation 09ecb643-05b9-49bf-94ae-8708d9a5997a · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Extreme Q-Learning: MaxEnt RL without Entropy

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:46.587307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:fa6ef8990df975c162c933467d427cb785a611e6d0e67332c003f93084ab4fa1

Observation d11e5360-8f6a-457d-9fe9-64a546f422c7 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.657389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:1f1a67d844c84d191167e51a2cd3826891c3325385bc9416e8f1176b019de705

Observation bf7a61dd-b81f-4142-a925-aa35fc35e55f · inbound

Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer cites this paper.

Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer Extreme Q-Learning: MaxEnt RL without Entropy

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:10.148459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T13:56:01.365382Z digest=sha256:8b9d8879c23d404b4c216c2d92dbb86845cbde93335bba4a47b5cb15d06ce9c9

Observation 3d758c3f-83aa-4a98-86e0-7347ec0464ca · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:45.853436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:4ab79b0152f100a21a7ae77b41697fa1da28483c86f236333a8f487079dc5e30

Observation 0729d11e-2144-4596-9322-bf9bc9c14bbb · inbound

Spectral Souping: A Unified Framework for Online Preference Alignment cites this paper.

Spectral Souping: A Unified Framework for Online Preference Alignment Extreme Q-Learning: MaxEnt RL without Entropy

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.924661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T07:54:56.356555Z digest=sha256:0a2e1490517375631c235deb2c2a436fab086cc6f22e1267b12401d3a04669ed

Observation e8b078c2-2a53-450e-8180-6e537c3574f3 · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.945140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:713326a78d78028eaa87f32c29d75f268da92fe5403b7246efa05c8a10796de8

Observation b8ac3f31-467e-42f2-b8c9-b082063529b4 · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Extreme Q-Learning: MaxEnt RL without Entropy

Reference 269

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:22.550304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:22.550304Z digest=sha256:3073114a9d25768c9909b69a48965bfa718fb1ba8df289f6af127f728ce4b88d

Observation 20012c6e-ff60-40d2-8d45-2a8968370b20 · inbound

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning cites this paper.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.191923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.191923Z digest=sha256:9a320200abe530009eee71730b04a187ab2e5117fb42f5ade2e89bc7abc023ee

Observation 804d8ce9-a6bc-43be-a021-0c169128cbf5 · inbound

Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL cites this paper.

Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:30.190639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:30.190639Z digest=sha256:6887feba7b175b6520dac8661825ca3c8a6fca363d6cb9f8042280108e993d5b

Observation 4decfe0d-e354-4fdd-beec-7f24902c33de · inbound

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control cites this paper.

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control Extreme Q-Learning: MaxEnt RL without Entropy

Reference 244

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:46.611586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:46.611586Z digest=sha256:c16b575821c4123d802c4b3451d9b530a096ba6e11ab362fb15c59b704b825e4