Pith. sign in

Paper Citation Record · LEDGER

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion

As of 23 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2606.31691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.31691 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T05:28:28.298787Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 577b2d4a-0e01-445a-8c35-2f5370713f3c · outbound

This paper cites Deepmimic: Example-guided deep reinforcement learning of physics-based char- acter skills.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Deepmimic: Example-guided deep reinforcement learning of physics-based char- acter skills

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.696535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:360763e2c90daf7e74c4e24f67075d117e241cb129015eaf3952defb9c7352c2

Observation c0851825-7d3c-4dc7-9249-86fdc1a84903 · outbound

This paper cites Towards robust motion control in multi- source uncertain scenarios by robust policy iteration,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Towards robust motion control in multi- source uncertain scenarios by robust policy iteration,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.665133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:0a452790686462bf89e7f15957dcc214c035ab8851110e77df08057a67c512d2

Observation 3bb1bd68-d373-4717-82e6-a2c714e9802e · outbound

This paper cites Isaac gym: High performance gpu-based physics simulation for robot learning,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Isaac gym: High performance gpu-based physics simulation for robot learning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.670441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:9e95bb6d17ef09e5640cb92a8d6f95b4d3c2e24145362d071c96e1f83f7351c4

Observation 85d90521-0a46-4d7d-bbbb-db3f667cb576 · outbound

This paper cites Learning to walk in minutes using massively parallel deep reinforcement learning.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Learning to walk in minutes using massively parallel deep reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.661458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:796764bceb0a2e296311b86b489ac752154a1b11e7e8f350df5d7a931be3b311

Observation beb051b3-c625-4441-a55e-36d2e95d0834 · outbound

This paper cites Parallelq- learning: Scaling off-policy reinforcement learning under massively parallel simulation,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Parallelq- learning: Scaling off-policy reinforcement learning under massively parallel simulation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.659633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:c8c75dec8ebdcb66e12dc6c1e681556c2d935ea39bc9315c85e598d9109c8e16

Observation 18ef4f72-b86d-457f-aae2-034b3470a634 · outbound

This paper cites Randomized ensembled double q-learning: Learning fast without a model,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Randomized ensembled double q-learning: Learning fast without a model,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.689497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:726cfcf08e046631ee803d3dd2f3764256e01b5cc7010c0eef50dbd40417ac24

Observation d6a03610-57c2-453d-a00d-3c9d83c53571 · outbound

This paper cites Understanding and preventing capacity loss in reinforcement learning,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Understanding and preventing capacity loss in reinforcement learning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.691225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:e3609543e2c68240922790c48cc03b329d14f150da1f032de52f5100d7f484ae

Observation 56a943d7-4848-4e59-bf97-2187a60888ca · outbound

This paper cites The primacy bias in deep reinforcement learning,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion The primacy bias in deep reinforcement learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.692973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:26bd8947101c692e0c78ea69730ce4ce7c5577a4bc5694f48e217a5665b18dd6

Observation 199f3d32-e9c7-47d1-9c84-61554d50c5a1 · outbound

This paper cites Deep reinforcement learning with plasticity injection,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Deep reinforcement learning with plasticity injection,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.687595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:45537175fd39d93cdcb67400f63ddce80c7c4395d71cf5a1c7d3c16ead100000

Observation 6df9832f-17f5-4571-b5d5-08501d113738 · outbound

This paper cites Human-level control through deep reinforcement learning.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Human-level control through deep reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.666905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:74b19925cb5c1b1420d57e5c048b8fd7956aca950a75379aae28cda97c0c6d20

Observation 9425a985-eb4f-4405-abcf-d5f72282856b · outbound

This paper cites Continuous control with deep reinforcement learning,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Continuous control with deep reinforcement learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.663237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:dba5dd8417d4aef7d5516b678164c3ce7a5dd1d79b2512ebc2ab3d7c38ad66ca

Observation 7038c482-182f-4e7e-95a5-3658d64bb3d4 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Off-policy deep reinforcement learning without exploration

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.683973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:96eca3f569e9212d20a1fe9cbea5ad117aabba664d7a113106d6103b22976a46

Observation 67368d14-222b-4f99-80af-5d10b4fa672b · outbound

This paper cites Addressing function approxi- mation error in actor-critic methods.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Addressing function approxi- mation error in actor-critic methods

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.682164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:bdda5f0f1ad460dc18d2ad0dfca4636b2ec8e0b4eb73197da50b4cec8a585b37

Observation 6a6b1675-58f3-4067-98cf-fdb1d26e33a2 · outbound

This paper cites Smoothed action value functions for learning gaussian policies,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Smoothed action value functions for learning gaussian policies,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.668718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:3598c6a147b61d0013d6c0dd9f606fd6179e1eb3fe81dda5595a313c8808594b

Observation e5b99d32-f6ad-4d9a-98c4-4dea8920deb0 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Stabilizing off-policy q-learning via bootstrapping error reduction,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.678667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:88f27312189aaf252ac2dda616e0445183ec033da19f0f3fc1a83f6c19f5b248

Observation 082f1e61-89ec-4554-ae27-ec28c0b8afc8 · outbound

This paper cites Demonstrating MuJoCo playground,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Demonstrating MuJoCo playground,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.680348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:fe642d68716fa0d77b363183df41888db395e196313878b060cec43f64ae177c

Observation b9387db1-df13-429d-b33a-ff9f31a8295d · outbound

This paper cites Humanoid- Bench: Simulated humanoid benchmark for whole-body locomotion and manipulation,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Humanoid- Bench: Simulated humanoid benchmark for whole-body locomotion and manipulation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.685806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:68a31a9c9f9e65656ba0d316082177313848435f76e4fffe748545b3699e1ff7

Observation 49629b77-0c2b-4599-b51f-aede4b038750 · outbound

This paper cites Proximal Policy Optimization Algorithms.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Proximal Policy Optimization Algorithms

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:25:42.282822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:bffdbabf26fbba1f3d280e9fe78bf299e84ffb47ce0e46893142118629bcfa30

Observation ec7a3f50-326f-420d-a0ad-1a081c8816ec · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.694711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:c711254441b03cb9ca795bac2d360361cda30b009ba3f33fbeaa0b913180c4e0

Observation fa85eaec-3c3b-4835-b975-a38a3c06821e · outbound

This paper cites A distributional per- spective on reinforcement learning,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion A distributional per- spective on reinforcement learning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.698246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:7fa4f2ba84ed9d71a24001a903d80efa16410237ab66136427562631483d8e0b

Observation 777d2734-968a-4372-809d-a0cc91583f46 · outbound

This paper cites FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:42.276740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:d8e5074112760dc816fc18c8c9c5df1b408ff1204c68b2304f91c778abad7999

Observation 5a9aa663-3fb2-45b1-97e9-8c7f442239a9 · outbound

This paper cites Sferrazza, C., Huang, D.-M., Lin, X., Lee, Y ., and Abbeel, P.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Sferrazza, C., Huang, D.-M., Lin, X., Lee, Y ., and Abbeel, P

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:25:42.279278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:49e67dfe280b99ecb29653a47243f7efd4e33281dd2a04622d8e0b85e5228df3

Observation fee1a2df-bc03-4776-b843-a95095394f45 · outbound

This paper cites Understanding plasticity in neural networks,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Understanding plasticity in neural networks,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.699855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:76524295ca3a5f1a7c6b481de7f1942509d6acc2c799188770f87f20ef0ecb93

Observation f9cc123d-ef33-4701-ad96-21cbcefbcbea · outbound

This paper cites Dropout q-functions for doubly efficient reinforcement learning,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Dropout q-functions for doubly efficient reinforcement learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.701584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:e8fc4bb7b028d55fa6a03f0d39a9a96446cf92b2e5de4bcc8ffa7d463526131c

Observation 2c714ea5-f0d1-42bd-bd02-48764c69ca09 · outbound

This paper cites Distributional soft actor-critic with three refinements,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Distributional soft actor-critic with three refinements,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.673983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:4ec8da1a938b9cafa7381410343c63b0a09b62aedd27bf2c7f7514cccac6e18d

Observation 03469d2b-161f-4484-a2dc-8fc8a03152a6 · outbound

This paper cites Controlling overestimation bias with truncated mixture of continuous distributional quantile critics,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Controlling overestimation bias with truncated mixture of continuous distributional quantile critics,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.676853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:5cfc8866bd78aae15cd9c770279829e104fcd8623c161c2f9d5d54820b331c68

Observation 885ddcd0-b9ed-4882-9a81-0af74ef93f18 · outbound

This paper cites Improving Deep Reinforcement Learning by Reducing the Chain Effect of Value and Policy Churn.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Improving Deep Reinforcement Learning by Reducing the Chain Effect of Value and Policy Churn

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:25:42.280503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:f4a8543352474b2e1992d2e178e93a0a525d5a6a3bbb798a78183e8f69e5e3c8

Observation 94520889-b3d5-47ca-9d65-34385a5a08c7 · outbound

This paper cites Stop regressing: Training value functions via classification for scalable deep rl,.

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion Stop regressing: Training value functions via classification for scalable deep rl,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T21:42:56.672302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:28:28.298787Z digest=sha256:7507a0a0dc1dd84cbc972d1b474a0fb74fd098b71bd11ab8eec020d9fcb12a9f

Pith citing papers

No inbound Pith citation observations are available.