Pith. sign in

Paper Citation Record · LEDGER

Combining policy gradient and Q-learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:1611.01626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1611.01626 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:22:18.826187Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-18T14:02:39.654453Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e3af8771-7b49-4e74-b918-5fd3f8416270 · inbound

Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor cites this paper.

Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor Combining policy gradient and Q-learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:48:10.753812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T01:48:10.689906Z digest=sha256:3ac83f8a5f14ce17f5cb8f151ce5ffbcccc720fa76e55984c8cfaf60525f839a

Observation 7961c1ad-3399-4b34-90e3-82227103d096 · inbound

When Maximum Entropy Misleads Policy Optimization cites this paper.

When Maximum Entropy Misleads Policy Optimization Combining policy gradient and Q-learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:18.826187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:18.826187Z digest=sha256:0f4780616f3afe323d21f5e69f68da2cb81b0ca11be322f1abf1fb0d368a3346

Observation 328fca76-6e97-4472-bf4c-4bad2dbc6546 · inbound

On the Robustness of Derivative-free Methods for Linear Quadratic Regulator cites this paper.

On the Robustness of Derivative-free Methods for Linear Quadratic Regulator Combining policy gradient and Q-learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:55.429439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:55.429439Z digest=sha256:b0d9ff9f4e4feef0719ef5b3f2652eea9652e6a3acbc2de55ac5211dbf93942a

Observation f28ccca0-cc70-4a9f-a57d-33a551ae3286 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF Combining policy gradient and Q-learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:02:39.658388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:93ff317c33535235d8cdced38294daeb4af413ce22192a6d38143c3efe3392bf

Observation 283a1d5e-9550-4ee2-9753-70901f363af4 · inbound

Soft $Q(\lambda)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces cites this paper.

Soft $Q(\lambda)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces Combining policy gradient and Q-learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:30:31.963465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:26:42.812540Z digest=sha256:356df286eda5172cdc698c6e8852fea93da5261331881e13c56bfb9488c053dc

Observation 6a21aeb6-634e-42df-ba9d-42299475379b · inbound

Soft $Q(\lambda)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces cites this paper.

Soft $Q(\lambda)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces Combining policy gradient and Q-learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-02T16:26:13.114236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:26:13.114236Z digest=sha256:66a6fac60b63704f7a5a894c2d1a58fcc6beee2649180c5992d39d868512aefc