Pith. sign in

Paper Citation Record · LEDGER

Extreme Q-Learning: MaxEnt RL without Entropy

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2301.02328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.02328 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:45.632354Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:36.943655Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 575ff729-ed03-4652-b4e6-6f605c574e73 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Extreme Q-Learning: MaxEnt RL without Entropy

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:48:36.476653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:9ceb0de61cb55f8357053fd318c0894e1574ab7801a15df472a3fab4f4f68b29

Observation 202fe6ac-af03-4756-8236-36e4ef3b38c3 · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.632354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.632354Z digest=sha256:7f0ad93ec2c2ec6a607ca781d39f19d7737b90e8e0fbc603cdfad6f2796353c4

Observation 14dd046c-95d9-4c78-9362-9a3b7e400315 · inbound

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood cites this paper.

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood Extreme Q-Learning: MaxEnt RL without Entropy

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:24.560434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:22:24.560434Z digest=sha256:094dee5e9061c69d58aacfda56ddd5e0ffae830ce30676d6969b2e2c426a2016

Observation 6b556c63-0a75-4c9d-a391-14da58e48501 · inbound

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion cites this paper.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Extreme Q-Learning: MaxEnt RL without Entropy

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.548506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.548506Z digest=sha256:91cffae66b0f16711fcfbda32adcf3877f8ef91efe49cee3ef4375906afb962c

Observation 09ecb643-05b9-49bf-94ae-8708d9a5997a · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Extreme Q-Learning: MaxEnt RL without Entropy

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:46.587307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:ad36ed11d140ce8090d46d1d435bd6800dded48de21b8db1cd0848e997f38008

Observation d11e5360-8f6a-457d-9fe9-64a546f422c7 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.657389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:52b41292efeded0f7bf0ed65b6648f5abda64dcc4204263aa064017dfc15a9a6

Observation bf7a61dd-b81f-4142-a925-aa35fc35e55f · inbound

Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer cites this paper.

Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer Extreme Q-Learning: MaxEnt RL without Entropy

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:10.148459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T13:56:01.365382Z digest=sha256:0f5c8b161b54b5ba57ddcc8d3f72d116b815249ddd017c2d86e453df68670036

Observation 3d758c3f-83aa-4a98-86e0-7347ec0464ca · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:45.853436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:ae399573098bf66d77815d87a16edc874374f0bddd2d4a4b541d6550407df924

Observation 0729d11e-2144-4596-9322-bf9bc9c14bbb · inbound

Spectral Souping: A Unified Framework for Online Preference Alignment cites this paper.

Spectral Souping: A Unified Framework for Online Preference Alignment Extreme Q-Learning: MaxEnt RL without Entropy

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.924661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:54:56.356555Z digest=sha256:82ea0c16aaaf5542d6e75557f6b89aafd538182832617cd464ffde9fccb419a7

Observation e8b078c2-2a53-450e-8180-6e537c3574f3 · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.945140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:14f42cfccee19e11285e3b6414c91f0a1b36ddc14e1f578780affffcca37a08a

Observation b8ac3f31-467e-42f2-b8c9-b082063529b4 · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Extreme Q-Learning: MaxEnt RL without Entropy

Reference 269

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:22.550304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:22.550304Z digest=sha256:7997d182eafa9bf9698099475a0da1184d21188a985451eb0625339c77fd6499