Pith. sign in

Paper Citation Record · LEDGER

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners

As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2607.13274.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13274 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:44:33.680003Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 76181836-1d2f-4406-9fe1-d9a89c40dbb7 · outbound

This paper cites Mastering the game of Go without human knowledge.Nature, 550:354–359, 2017.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Mastering the game of Go without human knowledge.Nature, 550:354–359, 2017

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:26.741095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:26.741095Z digest=sha256:963d17b3e1f022ee79f36e120e1116c0391588cbc3f4147beffc4f3671d9703f

Observation 59852375-7144-44d1-940a-20658f9d4847 · outbound

This paper cites an unresolved cited work.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:26.833885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:26.833885Z digest=sha256:a37b035ce1c4c8ea26b7f6a588474de3bfc5c03dfc848c3b7ddd5abe75bb0477

Observation e89e5b8a-40e7-454e-8148-cf1139b16432 · outbound

This paper cites Wurman, Samuel Barrett, Kenta Kawamoto, James MacGlashan, Kaushik Sub- ramanian, Thomas J.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Wurman, Samuel Barrett, Kenta Kawamoto, James MacGlashan, Kaushik Sub- ramanian, Thomas J

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:26.993836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:26.993836Z digest=sha256:abcf04c74087b8640dfd9e8f09f8b41a3cbd413b37010f8888965057bf469c3c

Observation b4607d67-9f1c-405d-9864-2c677e2bb66a · outbound

This paper cites Magnetic control of tokamak plasmas through deep reinforcement learning.Nature, 602:414–419, 2022.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Magnetic control of tokamak plasmas through deep reinforcement learning.Nature, 602:414–419, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:27.075781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:27.075781Z digest=sha256:9ec67c2cdf61ba0fd45eef56aebef8f64adaedd8663f98f26c9ea0c65447227f

Observation a1b5f11a-da62-4b10-b9cc-8799fa724a69 · outbound

This paper cites Dense reinforcement learning for safety validation of autonomous vehicles.Nature, pages 620–627, 2023.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Dense reinforcement learning for safety validation of autonomous vehicles.Nature, pages 620–627, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:27.097072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:27.097072Z digest=sha256:e26d004dd908355a4ef64503426d3352e8593700f48960027d4fedeb32742e91

Observation bd4bbc9d-398a-4ac5-9dea-54f8ae4f78b0 · outbound

This paper cites A survey of reinforcement learning for software engineering, 2025.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners A survey of reinforcement learning for software engineering, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:27.195390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:27.195390Z digest=sha256:b43bee7f4eac1a8d8f656a9daf761aa369138b43704cf1c07409a9faee55a93b

Observation be714932-901a-4db4-9703-3cd568a0519d · outbound

This paper cites Rainbow: Combining im- provements in deep reinforcement learning.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Rainbow: Combining im- provements in deep reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:27.356890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:27.356890Z digest=sha256:f82b7f9b0da68c4920b9abc1cd4afad5891f1a4033d8c649c8157f88fb64021d

Observation b4303a44-34aa-4d46-9466-6608155991e0 · outbound

This paper cites Noisy networks for exploration.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Noisy networks for exploration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:27.435167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:27.435167Z digest=sha256:6535541b050cf44b75fe641376837d4c1a02261f5e1f3cb776895dcc794134e3

Observation a944e7ce-44de-48dd-8e4d-677ab55be01b · outbound

This paper cites A distributional perspective on rein- forcement learning.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners A distributional perspective on rein- forcement learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:27.574779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:27.574779Z digest=sha256:9cec6f7b71f878613365eb90bac52ff94b606f2d2d939748db76176dfd1ebd4e

Observation 22a7e2b2-29e6-4412-81de-81d1231e7b29 · outbound

This paper cites Prioritized Experience Replay.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Prioritized Experience Replay

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:27.738949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:27.738949Z digest=sha256:0f7b5c9570f6a728693644afb665b72ff0306fb557645d5221965fc959236456

Observation 6b2a0f1c-0755-4013-b223-82f82ac2067a · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.Journal of artificial intelligence research, 47:253–279, 2013.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners The arcade learning environment: An evaluation platform for general agents.Journal of artificial intelligence research, 47:253–279, 2013

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:27.913325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:27.913325Z digest=sha256:ea058e2480f4579100cc384b76f50a803b74cef40c0ae95f9cff3523b6c6555c

Observation fc7f2cef-3918-4a40-8a1c-d78b0d02dca5 · outbound

This paper cites Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:28.035495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:28.035495Z digest=sha256:6577f5d84fdad37f2f07230cc294d1c037e95d30d292c35b19745c9de1ddab3b

Observation 8c522549-04fd-4eef-a2b6-25f6d23fa443 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:28.128592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:28.128592Z digest=sha256:b16652d5d68869b0aa2ab73399a918200b2ba1f53933a187b52c54721d4b1efd

Observation 3622c9b4-b431-4b34-b57d-4232e433564d · outbound

This paper cites Mujoco: A physics engine for model-based control.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Mujoco: A physics engine for model-based control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:28.296584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:28.296584Z digest=sha256:1f27a4359c0dd3607bad8378e16a1e1864d23a523f673ec710d8ad4aba1edd17

Observation cb76570f-75c5-4c89-a37f-a6884f7e255b · outbound

This paper cites DeepMind Control Suite.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners DeepMind Control Suite

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:28.406736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:28.406736Z digest=sha256:22a97578aa89d637a3e1420309d6f60d38eb4e2284210e834fd37583692d01d1

Observation 2d45eb03-2ba0-431c-8b8b-2cbc5e397cea · outbound

This paper cites Maximum a posteriori policy optimisation.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Maximum a posteriori policy optimisation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:28.573897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:28.573897Z digest=sha256:597fcd71206ac21f0d139d80fe167d516e4a2d1a019eeb08c3421af814d9970a

Observation 48312fc9-f76d-4fdd-bcf4-8928084190ba · outbound

This paper cites Greedy actor-critic: A new conditional cross-entropy method for policy im- provement.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Greedy actor-critic: A new conditional cross-entropy method for policy im- provement

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:28.703913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:28.703913Z digest=sha256:3a2729e78bece5c5c74eb32137591d9f92e1d1af746c8f19143a32f24cd5f448

Observation 2a8c6eac-f866-4e6a-b9bd-1ec76db9abd2 · outbound

This paper cites Offline reinforcement learn- ing with tsallis regularization.Transactions on Machine Learning Research, 2024.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Offline reinforcement learn- ing with tsallis regularization.Transactions on Machine Learning Research, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:28.820042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:28.820042Z digest=sha256:91b33ba13bba8d5d7f83eecfef6ed3e69916a2f38bfd7e287cab22ce19f3e154

Observation 381f5ffd-61bc-44c1-87d7-054c39b80524 · outbound

This paper cites Trust region policy optimization.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Trust region policy optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:28.938223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:28.938223Z digest=sha256:6c3e29292ca343e1e94182eef689a697ba94247cbccfdd10680ede869948588b

Observation 67699977-a8ed-4fc4-8b97-8c5271e0a354 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:29.079070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:29.079070Z digest=sha256:adc572e8aee033a7e333da486a524b468de1b27999c8eab438c5dff71d8d7a67

Observation fa2c6676-148b-487a-9451-0cb1bfcf320b · outbound

This paper cites Lillicrap, Jonathan J.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Lillicrap, Jonathan J

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:29.216492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:29.216492Z digest=sha256:e7ca9a8b3a3db83d5fe73b953157b45cea92f98ec57a902764ba7a8709f42568

Observation df7e68d9-646a-47fb-839f-9fbe7ba5a81f · outbound

This paper cites Al-Sakkari, Ahmed Ragab, Mohamed Ali, Hanane Dagdougui, and Daria C.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Al-Sakkari, Ahmed Ragab, Mohamed Ali, Hanane Dagdougui, and Daria C

Reference 22

Resolution
verified exact
doi, observed 2026-08-02T05:49:43.992833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-02T05:44:29.334558Z digest=sha256:0d0b9b45feb3d57dba7f37f8e8db162d68b30f389b23f862b4577102d0addf70

Observation f64d9506-0b65-41d6-b332-2b56207d0c84 · outbound

This paper cites Williams.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Williams

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:29.442793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:29.442793Z digest=sha256:aec40bb720cf324dfbd031c739ebddf333758e082dcb1711e809fb5403e35518

Observation b752b392-1d7b-4d64-96c2-a63a5491fb2a · outbound

This paper cites Rupam Mahmood, and Martha White.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Rupam Mahmood, and Martha White

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:29.576781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:29.576781Z digest=sha256:d2716931832fb9d925d7356dd78b0410f099782ea9f9dd687cf13bfc32d2a6f0

Observation 58c5a498-9f74-4ba4-bb0b-37901f2ac260 · outbound

This paper cites Mirror descent and nonlinear projected subgradient methods for convex optimization.Operations Research Letters, 31(3):167–175, 2003.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Mirror descent and nonlinear projected subgradient methods for convex optimization.Operations Research Letters, 31(3):167–175, 2003

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:29.671054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:29.671054Z digest=sha256:74117ebff5ea654a55e02fc52e59d411770394a60b919e5681c487973a241093

Observation 8d621a94-4a8a-452b-9391-e1ab9efdc825 · outbound

This paper cites Investigating the utility of mirror descent in off-policy actor-critic.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Investigating the utility of mirror descent in off-policy actor-critic

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:29.764860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:29.764860Z digest=sha256:cfc205e9b9193bbeca6ebf5fc1cb8cbd7c290984d72a9fd90589ec8baa509235

Observation 04eecc7d-6ba6-4fde-ab13-d21ffb221b69 · outbound

This paper cites Leverage the average: an analysis of regularization in rl.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Leverage the average: an analysis of regularization in rl

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:29.907200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:29.907200Z digest=sha256:eb9a620c00cb488845d3c0017ba86f2b9b2ba6673d2fb2084764431f661139e2

Observation 67ac4cf5-be0f-41ff-afe9-1634ad231491 · outbound

This paper cites Mirror descent policy optimization.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Mirror descent policy optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:29.989923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:29.989923Z digest=sha256:1aa094d583927c6dd354379369e49abddd8f8c01d8bb5dce619f1de5fb6b766b

Observation 1e891a81-6264-431b-bb8d-46f584c05af3 · outbound

This paper cites Machado, Pablo Samuel Castro, and Nicolas Le Roux.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Machado, Pablo Samuel Castro, and Nicolas Le Roux

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.108251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.108251Z digest=sha256:6752182505c43ff6efa716b21a670d769cf5a0eed5d51628e206901a0885172e

Observation 6c47e1de-26fa-4673-b753-b9c88acda69e · outbound

This paper cites Approximately optimal approximate reinforcement learn- ing.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Approximately optimal approximate reinforcement learn- ing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.220660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.220660Z digest=sha256:54e9ddb15d416d8fd04bb40123fc3d8932a7d478ad973eaf086766a85d238a9e

Observation 17bad94f-2570-41c3-838c-93adbc75832e · outbound

This paper cites Revisiting mixture policies in entropy-regularized actor-critic.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Revisiting mixture policies in entropy-regularized actor-critic

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.305447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.305447Z digest=sha256:e71574cbceec62534ea1f0f5596e1fed2a73d2cb6aba76497c241ef60acc2785

Observation 85a58543-8e10-4b50-9dc5-6e709fc0af7f · outbound

This paper cites Improving stochastic policy gradi- ents in continuous control with deep reinforcement learning using the beta distribution.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Improving stochastic policy gradi- ents in continuous control with deep reinforcement learning using the beta distribution

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.397569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.397569Z digest=sha256:a1c46848459222e2cd2cf65f12ccb0d708aa689057bd28098bc9fd4d9057b704

Observation 0b2f2a20-a5d4-4ea3-9e12-233a9e2a24ad · outbound

This paper cites Student-t policy in reinforcement learning to acquire global optimum of robot control.Applied Intelligence, 49(12):4335–4347, 2019.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Student-t policy in reinforcement learning to acquire global optimum of robot control.Applied Intelligence, 49(12):4335–4347, 2019

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.553143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.553143Z digest=sha256:b6ab858efa9a4518b34f0ddefbf7270a97d3f9c5374ea479263db225d6e32cab

Observation cd677f8a-2668-4017-b1d8-b920c087939f · outbound

This paper cites q-exponential policy optimization.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners q-exponential policy optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.661370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.661370Z digest=sha256:51c66759f2cd3b40ba8025a03b663d04230b1d79e33263749ab48635018cb18f

Observation dd697e9d-b983-4464-8662-6d3bab226f82 · outbound

This paper cites Policy Representation via Diffusion Probability Model for Reinforcement Learning.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Policy Representation via Diffusion Probability Model for Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.791324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.791324Z digest=sha256:4873be595cece5677d0bdc498c8365c1fa39a3864a8fe7ff94898c2822db1541

Observation 131fc597-58d5-4fb6-a993-89847e6b741a · outbound

This paper cites Learning a Diffusion Model Policy from Rewards via Q-Score Matching.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.881499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.881499Z digest=sha256:2f6528c42188568c8a77a96648c5299c8cf3aa9fac6d46ef587001ce33b28f9c

Observation ebf78c45-fe66-4109-8457-dbd017bb2122 · outbound

This paper cites Diffusion-based reinforcement learning via q-weighted variational policy opti- mization.Advances in Neural Information Processing Systems, 37:53945–53968, 2024.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Diffusion-based reinforcement learning via q-weighted variational policy opti- mization.Advances in Neural Information Processing Systems, 37:53945–53968, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:31.053462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:31.053462Z digest=sha256:3ecd11616efaf3677dae859a47e1dc8c9a0d5bedeafca1a7a6f1ee7bc54b4dbb

Observation 73ccc44a-1e9e-4a18-bb03-ac97bfc23d00 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Addressing function approximation error in actor-critic methods

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:31.165487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:31.165487Z digest=sha256:c51ce57859d40ea2446fdbe2bc325259c74a94fcc48b4842f71c667d588dd167

Observation 7117d35f-2783-43b5-9836-874abf3dd5ca · outbound

This paper cites Sampling from energy- based policies using diffusion.Reinforcement Learning Journal, 6:2291–2307, 2025.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Sampling from energy- based policies using diffusion.Reinforcement Learning Journal, 6:2291–2307, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:31.253418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:31.253418Z digest=sha256:36b48d39ecb3231f8533a2e43384b8b63f09a8ebe6f0f01ef06500adb66059a7

Observation 36910a57-7546-4778-8251-6b6599474f86 · outbound

This paper cites The primacy bias in deep reinforcement learning.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners The primacy bias in deep reinforcement learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:31.339763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:31.339763Z digest=sha256:c06da09a913e5ee27f5ceb0b621e3470992a5a4355cb6b239177c0d0a1069d19

Observation f7108f00-a3e7-4df8-bfb1-436926ae75dd · outbound

This paper cites Addressing the plasticity- stability dilemma in reinforcement learning.arXiv preprint arXiv:2512.01034, 2025.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Addressing the plasticity- stability dilemma in reinforcement learning.arXiv preprint arXiv:2512.01034, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:31.453906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:31.453906Z digest=sha256:ef3473314ac1f2034e9a411d78730f9d2953cfe6d14390a9111d8765657737f8

Observation 842181b4-1d58-46be-9e82-a2bf4e6ed9a1 · outbound

This paper cites Sample-efficient reinforcement learning by breaking the replay ratio barrier.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Sample-efficient reinforcement learning by breaking the replay ratio barrier

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:31.578005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:31.578005Z digest=sha256:a3478dc1bdd93a17f22906f2c1e8d6a9042534d865676dad8da6d61ff60e8179

Observation 41ce20f4-2251-4e64-8f52-c79e484116cb · outbound

This paper cites Mad-td: Model-augmented data stabilizes high update ratio rl.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Mad-td: Model-augmented data stabilizes high update ratio rl

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:31.692429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:31.692429Z digest=sha256:e431d15c8bfdc0c6148ab57433a40e9e2d55f9a0c287bedb39cf7a9459b86ec3

Observation 818b023d-dd62-452f-98d8-fb451027ef8b · outbound

This paper cites Stochastic approximation with two time scales.Systems & Control Letters, 29(5):291–294, 1997.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Stochastic approximation with two time scales.Systems & Control Letters, 29(5):291–294, 1997

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:31.868372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:31.868372Z digest=sha256:3bef9d38fed15004639be9938e0b9e34678f6a560302acfb2978fe501c684b68

Observation 391ac6bd-0c48-4198-b938-9b9795d6df0e · outbound

This paper cites Actor-critic algorithms.Advances in neural information processing systems, 12, 1999.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Actor-critic algorithms.Advances in neural information processing systems, 12, 1999

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:32.067957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:32.067957Z digest=sha256:41a82c11179c56bfb35d1ce611426f2a34eafad2b6fc96b72d28061e32ea5596

Observation 71942004-296b-4e93-b1df-62a99b4bff43 · outbound

This paper cites Sutton, Mohammad Ghavamzadeh, and Mark Lee.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Sutton, Mohammad Ghavamzadeh, and Mark Lee

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:32.204583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:32.204583Z digest=sha256:5274b57ed797ffc5bcbdab12fd183c19fa5df301b56ab52c982deec4e538a5b1

Observation fe2f7147-e909-46e6-9b9a-7582c394ff28 · outbound

This paper cites Likelihood ratio gradient estimation for stochastic systems.Communications of the ACM, 33(10):75–84, 1990.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Likelihood ratio gradient estimation for stochastic systems.Communications of the ACM, 33(10):75–84, 1990

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:32.415091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:32.415091Z digest=sha256:eb4c5468774262ea53791a7ad6e7ae89763bdb3f5b1b59a52604291cd0e128f6

Observation f0394865-a082-433e-8b6f-e4fe874cabba · outbound

This paper cites Auto-Encoding Variational Bayes.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Auto-Encoding Variational Bayes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:32.555808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:32.555808Z digest=sha256:29189da679b5713e04da18a1b1b27af1495610c81c99edda00ab27f62e8f4478

Observation b96b1259-a825-477f-8179-6e7d8912837f · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Categorical Reparameterization with Gumbel-Softmax

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:32.660438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:32.660438Z digest=sha256:b80f2b4f646dca0eab441f6483abb63f37efd327998c380e976ca90e0d97ceb5

Observation f7eea0fb-51a1-4c6b-b58b-f27c934d7922 · outbound

This paper cites Variance reduction properties of the reparameterization trick.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Variance reduction properties of the reparameterization trick

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:32.780673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:32.780673Z digest=sha256:a5acfa8c0d4d0067da44b445d91032f7341f16822f211d61f04239f761524283

Observation 45bf848a-18c9-400a-8924-47603f8b1aa6 · outbound

This paper cites Pipps: Flexible model-based policy search robust to the curse of chaos.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Pipps: Flexible model-based policy search robust to the curse of chaos

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:32.920408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:32.920408Z digest=sha256:c41f3ee68c54945ee47c6950db63853b1a6a5e8e6ba8bc7df92f12551e5a0f57

Observation 0329d157-f071-4f2d-9fc5-0fdf011a2e18 · outbound

This paper cites Pid controllers: theory, design, and tuning.The international society of measurement and control, 1995.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Pid controllers: theory, design, and tuning.The international society of measurement and control, 1995

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:33.087488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:33.087488Z digest=sha256:2860b9775d19efd9f8c71cb7a5786478ce31f8fc00b586c28c8ee453b569e36d

Observation 9a40193e-c5ad-48bd-8820-0d92e43bef11 · outbound

This paper cites Fat-to-thin policy optimization: Offline rl with sparse policies.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Fat-to-thin policy optimization: Offline rl with sparse policies

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:33.268132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:33.268132Z digest=sha256:39ad1d0d67c4c98c7b21f429bbec07c4de21c1ab5219e2b0aa704762c585be2e

Observation c3b3a3fa-98cb-47df-98ab-0bc9931f7baa · outbound

This paper cites an unresolved cited work.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Unresolved cited work

Reference 54

Resolution
malformed identifier
no resolver link, observed 2026-08-02T05:44:33.396253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:33.396253Z digest=sha256:c3118a01af48787467be907281bce4ba8aa12d00561ba142b7c672a9b1788744

Observation 8bb9d238-ec1c-4719-8711-5b07d498252d · outbound

This paper cites an unresolved cited work.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:33.521780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:33.521780Z digest=sha256:766d690f3c135d83d4c22b2addcfb9e86dee560a9e7ad7f6379679b9aebba87b

Observation 35dea39d-dfa5-42d1-b459-eb75807c0839 · outbound

This paper cites an unresolved cited work.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:33.680003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:33.680003Z digest=sha256:15889df0771740dd5377c9d5d79d275d325a0d94dc3e2c7f4dc22068059b2951

Pith citing papers

No inbound Pith citation observations are available.