Pith. sign in

Paper Citation Record · LEDGER

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic

As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2605.09157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09157 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:33:08.990277Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact13
  • verified fuzzy39
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 06d884c3-8bf8-4ebd-ae18-fcea3c060862 · outbound

This paper cites Reinforcement learning: Theory and algorithms.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Reinforcement learning: Theory and algorithms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.475200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:e48837243fdd38e7ca85bc4a324ad80fc2e308aeed7a86058efb37ed6e1a4e89

Observation f1c6a39a-bc3c-4ee0-bfeb-cab54e2c18a7 · outbound

This paper cites Understanding the impact of entropy on policy optimization.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Understanding the impact of entropy on policy optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.471699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:05bcdd1270045c91dd646a02c468de49ca330658d831578de67dc1b5a40ad477

Observation b32971d6-eae0-44ce-a01b-38de8aa42cb1 · outbound

This paper cites Maximum Entropy Reinforcement Learning with Mixture Policies.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Maximum Entropy Reinforcement Learning with Mixture Policies

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:22.891727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:cfd2d1f32a7ffd62659bfbbd65563c91ee9de90911e081112fd68f27139705f2

Observation e2be3b7e-52b7-4277-91c9-591d28336bc4 · outbound

This paper cites On the sample complexity and metastability of heavy-tailed policy search in continuous control.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic On the sample complexity and metastability of heavy-tailed policy search in continuous control

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.457205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:b1d6c221c116f621865b2bc0eadd73ff2661dea3e74ad9c1dc03fd2ace90a74c

Observation 05b4147b-b880-403a-80eb-33d724805114 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:11:22.773603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:5b2685d0fe92d1d26abf74ebdbd75ada8f0d0a3ce0955da04f7ec49af4a78c25

Observation 2fc38d5b-9d69-4c65-9198-9ee53fdcec2c · outbound

This paper cites JAX : composable transformations of P ython+ N um P y programs.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic JAX : composable transformations of P ython+ N um P y programs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.468041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:aee22211039386b6eca3bc8bf1062f6f75b16d2ebcd5b67ac1d10be13203e08f

Observation 5f3b3339-1883-4c3c-bca2-1c56f4043eb4 · outbound

This paper cites OpenAI Gym.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic OpenAI Gym

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:11:22.836602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:5d0246d4bb476740cc991e5a2e3290c053b44685d144b809e086a1e835ac0f82

Observation 7db74fb0-6519-4492-b292-b43139624b99 · outbound

This paper cites On upper and lower bounds for the variance of a function of a random variable.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic On upper and lower bounds for the variance of a function of a random variable

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.442004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:05874013fec1ad53b7e6d9940948e87e5bc6617a333cbfdc8cb35c605485c5b8

Observation 3ba1d749-aa97-4144-b0ec-7bc13c1d5770 · outbound

This paper cites Myosuite: A contact-rich simulation suite for musculoskeletal motor control.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Myosuite: A contact-rich simulation suite for musculoskeletal motor control

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.333236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:565c4754e84f40d2ee8ba0705e690a8ec1dd41a55a9aba8f6c5feee23b58dd6e

Observation 118e43a9-53e5-4c8f-b83e-3e05fb59c4eb · outbound

This paper cites Specializing versatile skill libraries using local mixture of experts.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Specializing versatile skill libraries using local mixture of experts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.420430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:201de28b006a816ea2d43dae51f9665c0d76b7a32485876599d068241643e50d

Observation 909acd8f-cbe1-45ec-9780-dd1435c3acb8 · outbound

This paper cites Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.423647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:0d2df230abd1934e3ab71db1a1cba35f8b69361981da9bef4e89de89087875c2

Observation a66c1beb-ea0d-428b-b567-aa9ab1150f10 · outbound

This paper cites Hierarchical relative entropy policy search.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Hierarchical relative entropy policy search

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.438300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:3c60115d30729fc1380056284be6c736eae570f0a6d476ce8566cb50c40e991f

Observation b1c7b8da-de2e-4be9-bd8a-ffea9a47fbc9 · outbound

This paper cites Model-free reinforcement learning with continuous action in practice.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Model-free reinforcement learning with continuous action in practice

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.412879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:67f5ad07848d03622a4c9cae4675c2b5a694d20f0a1edbfefb4739b6c4003789

Observation b32533d6-f5cf-4a08-b6c8-c9c783185888 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Addressing function approximation error in actor-critic methods

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.416738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:754af9c931090a0e7dfe52abb0c7b11ac5f0585c1adb79507e53885c6174f2ce

Observation 6cb18d3e-4abf-4b81-998f-bf71e8ebb0c2 · outbound

This paper cites Uncertainty in deep learning.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Uncertainty in deep learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.434060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:42d6fd84678b1b102c50048f08200c2515fb5345229f548bccdfeee8a60af472

Observation 62e38cab-91c7-4fe7-bfc4-ec742dbdce88 · outbound

This paper cites Acquiring diverse robot skills via maximum entropy deep reinforcement learning.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Acquiring diverse robot skills via maximum entropy deep reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.463877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:2ff19494ac436436855e17313f604b9a001ceca2136a2dac69bd5bf1954e66d8

Observation ee7d3d50-1cf3-4368-8805-dc3180aa5f10 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Reinforcement learning with deep energy-based policies

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.341082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:407245703e314747d6e1d957a28d8edb20d2c0645eb00ce528cb25ba8c7d3de1

Observation 6651af1a-3c40-41d8-9399-0b80f53dfc78 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.445703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:ed86e45aab02d9b937a39e72f73574932e41a2680ed2a25931241f28e63f868a

Observation ec6809bb-1d1e-429c-b991-6a88a7570380 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:48:10.852058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:23b98d20a5cdbabfd135f904e5f8bb13e6322672027d215b01fa4af04cfec323

Observation ab3fee38-09ea-4d14-8349-e644b5aee0e6 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Soft Actor-Critic Algorithms and Applications

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:f036a1bce5b4b5994168d615fcc1d914232cfa45cda68312a227528747c1c4f9

Observation 0cb0d849-adb8-470f-9f09-e0b366946746 · outbound

This paper cites Learning latent dynamics for planning from pixels.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Learning latent dynamics for planning from pixels

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.348406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:8bd8c6c0e6940941f7c4aad6898c6cce84ed76b726d199d0fcf30b75bee9c358

Observation 24c601ab-44b5-42d1-bb89-1caa69715d2a · outbound

This paper cites Learning continuous control policies by stochastic value gradients.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Learning continuous control policies by stochastic value gradients

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.449355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:7176cbc8b3920d5391f0cc5990f72bdc435f65694948530a81f05f7d73241feb

Observation 9bdc9302-e4c8-40e7-962e-0a641dce73a8 · outbound

This paper cites Off-policy Maximum Entropy Reinforcement Learning : Soft Actor-Critic with Advantage Weighted Mixture Policy(SAC-AWMP).

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Off-policy Maximum Entropy Reinforcement Learning : Soft Actor-Critic with Advantage Weighted Mixture Policy(SAC-AWMP)

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:22.842950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:9877002e64dc819df53351cfa332d0fef6160c9d0d5939d7a2872e9986335370

Observation 05787a78-2e17-40e7-8e28-bac7c0d6a8cd · outbound

This paper cites Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:22.869984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:275fc77a929217d7537b9e85a17179ad9a3f22ed9c6f0d110b0edf4632b2140c

Observation 5bd5f78e-1ac0-4fc7-9760-2cdbb4358992 · outbound

This paper cites Categorical reparameterization with gumbel-softmax.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Categorical reparameterization with gumbel-softmax

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.408638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:2abcc7599ec2de6566c4d9281877824020aeab5fab37e1cd2bcdc4343c6c0527

Observation 8204c685-b8ff-4d7a-8205-18a6fe6554bb · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Adam: A Method for Stochastic Optimization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:11:22.822486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:9c2b23e0a0478c3dece8439cbe636ebd3ef36e8b2cba581689be7d2d5df1d6d2

Observation edc0f208-e1a9-40ad-885c-6cfe51ab770f · outbound

This paper cites Student-t policy in reinforcement learning to acquire global optimum of robot control.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Student-t policy in reinforcement learning to acquire global optimum of robot control

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.344630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:e507baedf14f270599460806e3e73e0dcbc64126ea70c898ec5e147ad809059e

Observation d19de21d-04e0-434b-8da2-c5d9069155be · outbound

This paper cites Model-free policy learning with reward gradients.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Model-free policy learning with reward gradients

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.403437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:f0e506bfad00bed855c671e9167d4ab110b1afb3e723e9abe037a7757f4f4e4b

Observation edbf07ce-4e8c-4b45-81e6-b909cb9ab6a8 · outbound

This paper cites Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.360619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:36b84ac97b07beac623ae7f9cf2e01868eba2b275b98de596c683f0274fe6ef4

Observation 401bcb7b-f6a2-438f-a555-c3c2390f38ee · outbound

This paper cites Continuous control with deep reinforcement learning.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Continuous control with deep reinforcement learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:11:22.882243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:b656c811abd85935ae519806a5d8e27943a724b2046086cf308612d82f0c2f22

Observation 5efc6f9f-c731-4be9-bd12-3b99f976e4c2 · outbound

This paper cites The concrete distribution: A continuous relaxation of discrete random variables.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic The concrete distribution: A continuous relaxation of discrete random variables

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.395773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:2981f3ffdc531819d4048d693ab8dc1aad2a1b39269650340966b5728faa4cbf

Observation 778148b4-3481-47bf-a690-7047da57cfa5 · outbound

This paper cites Leveraging exploration in off-policy algorithms via normalizing flows.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Leveraging exploration in off-policy algorithms via normalizing flows

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.426968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:60b949e858e2942e8c405dea59ce56da998785949f7b054b5aee0023ebfcf50e

Observation 230c3560-1cd5-42af-9c9f-aaf452f73c92 · outbound

This paper cites S\ 2\ AC : Energy-based reinforcement learning with stein soft actor critic.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic S\ 2\ AC : Energy-based reinforcement learning with stein soft actor critic

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.388980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:417e534ada242f9fb05e18433c01b2735f9569025d928b6a0a2ae60750c49158

Observation 886fd1d4-3473-43bd-9540-cfb50801a3fa · outbound

This paper cites Reducing reparameterization gradient variance.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Reducing reparameterization gradient variance

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.382179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:00c40f8a4f026479db838971f9e8891c71eb64149534f316f5a9ef0a70c44d7f

Observation 8d6b7756-aefc-4310-af21-97711315d561 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Monte carlo gradient estimation in machine learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.352376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:8dd5586bd410dc399d0e7299384ee037437bcf9acfe4badb90721f16a1135ed0

Observation 6e72d158-8db5-491b-8aad-b2dafa102802 · outbound

This paper cites Robot skill adaptation via soft actor-critic gaussian mixture models.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Robot skill adaptation via soft actor-critic gaussian mixture models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.375590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:aca1a2fa40797c1a6695383f9220d440132ab60bdfb290ab5db32d864b820428

Observation a00a0810-6996-4cd5-a81f-d8d212df6b29 · outbound

This paper cites Greedy actor-critic: A new conditional cross-entropy method for policy improvement.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Greedy actor-critic: A new conditional cross-entropy method for policy improvement

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.385436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:53faad6e5822f97f4a384aedccd6b0d94a80d03e667797974133c27859e53ac9

Observation ac3acd16-06f0-43fe-a8b1-58c906e84da7 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Pytorch: An imperative style, high-performance deep learning library

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.430748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:5c17f503b8c11f6ac79bc14eb54ab42bbfefe788fdc186349a751e9f5204c48b

Observation 74b000b0-27a9-498e-bb44-2d6f490ea4e3 · outbound

This paper cites Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:22.860842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:ca61c3a4431f0c9e55b96cd3d7cba7d78ee964e48cbd6a267566f43b7633a4f8

Observation 94c67878-6e7c-49af-a7fb-a6104e4d653d · outbound

This paper cites Probabilistic Mixture-of-Experts for Efficient Deep Reinforcement Learning.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Probabilistic Mixture-of-Experts for Efficient Deep Reinforcement Learning

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:11:22.787704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:504c8d9e7092d80f52fed08e1df945745fe2dea69d7af3ec34055cc8767f982a

Observation 64d3372a-b612-4f17-b336-2fe671f8e3d3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Proximal Policy Optimization Algorithms

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:11:22.853183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:1913901e4e15ad66d871dedecff863a709acdf846276fbacfe7d64cb7d147771

Observation db6aecc0-cc08-491f-989d-461bb9b13083 · outbound

This paper cites Strength through diversity: Robust behavior learning via mixture policies.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Strength through diversity: Robust behavior learning via mixture policies

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.460617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:cf0ec90f4823b2f5c60a515fecefd4642ed4ed58be70ef695999bb88e5f9fcde

Observation 829b5396-cb1f-47cb-a712-8bb4668cee7a · outbound

This paper cites Reinforcement learning: An introduction.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Reinforcement learning: An introduction

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.453498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:6129b28c1aab68a395411b7093f6a2e3efd2ee5002f8a544004e9def970d9af6

Observation d2a5742c-9dbb-4f2d-851a-0a58ca80a436 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Policy gradient methods for reinforcement learning with function approximation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.368000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:fa7b4b7eff2ba25229d91760a8a63bd6105495b698cb60e2bde91dae28e54165

Observation 4feeb983-2734-402d-8b34-b4e00a139782 · outbound

This paper cites Implicit Policy for Reinforcement Learning.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Implicit Policy for Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T23:22:02.244370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:c771aa3714024b3721ac9e9c9461a6c25e2d4024c605cbaea0a52584412988d8

Observation e07a08fa-4270-4e26-89b5-a175398eb3a5 · outbound

This paper cites DeepMind Control Suite.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic DeepMind Control Suite

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:44:14.794093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:684f8afc775aafa90a279e90b85a0da00acb7d54040a58a61f9177462e484f2d

Observation b285df28-723d-41dc-9c03-e28c7ae86042 · outbound

This paper cites SciPy 1.0: fundamental algorithms for scientific computing in Python.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic SciPy 1.0: fundamental algorithms for scientific computing in Python

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.378562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:60657d7ee1ec06eb89769a71ca150e2d777dd8c039dea336a42573af8bc63d13

Observation fb737d3b-6080-41b3-b898-2656305d3810 · outbound

This paper cites Diffusion policies as an expressive policy class for offline reinforcement learning.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Diffusion policies as an expressive policy class for offline reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.364539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:e0b6887cbac2ad83af3b13922ad860d765869b5c819d889a8468ec101f5a2f48

Observation 47089ed9-a9d7-4e04-812d-60245b9a7c24 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.400109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:98a9fc671a736faf2ac4ba56dd1fcda3dc9e80d5ff655c3aa5fb69a1a2a7b378

Observation 7ceb34f8-9a68-479a-8a09-19a580605d33 · outbound

This paper cites Variance reduction properties of the reparameterization trick.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Variance reduction properties of the reparameterization trick

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.392327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:dfbcaafb51ac2a2d3d3dbd8ae7de294368c02dedf617416b4a445c82c2d4eb83

Observation 8be5421a-2b10-4700-88c2-94ef8f603f4d · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.337375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:e2dbafe2efd9b9a8a8b90ccf58cb1a93f337401cf6f12c840fe15288fadcb315

Observation 4f2c4491-6930-4690-9a65-3fa6f9570eb8 · outbound

This paper cites Latent state marginalization as a low-cost approach for improving exploration.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Latent state marginalization as a low-cost approach for improving exploration

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.355835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:83dd629b3b1a2679b48ad8327ef537f64a86cec322abf1d9a4d0707b4f22b0aa

Observation 570cccc1-1752-43e0-b5fa-78d2ef814852 · outbound

This paper cites Model-based reparameterization policy gradient methods: Theory and practical algorithms.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Model-based reparameterization policy gradient methods: Theory and practical algorithms

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T14:31:39.371561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:1e73c07d4679d24fca8c0529852bb9ef6c89abe0dd9154dc7f2e2af6a43552a8

Pith citing papers

No inbound Pith citation observations are available.