Pith. sign in

Paper Citation Record · LEDGER

A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2006.14171.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.14171 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:07:48.369414Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:39:41.546939Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3793f95a-11f8-46cf-86c2-36f4bc144295 · inbound

Dynamic Collaborative Material Distribution System for Intelligent Robots In Smart Manufacturing cites this paper.

Dynamic Collaborative Material Distribution System for Intelligent Robots In Smart Manufacturing A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:48.369414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:48.369414Z digest=sha256:7ddfed56848542e4102a71f5f0f29daa52fd39e85669fb45c34938be4eecfe4b

Observation 53be11a9-015f-4c00-9348-8634fb2cef0c · inbound

Data-Driven Policy Mapping for Safe RL-based Energy Management Systems cites this paper.

Data-Driven Policy Mapping for Safe RL-based Energy Management Systems A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:53.981879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:53.981879Z digest=sha256:4aa9fd50b69400bb7c23f10ca791f9a7e5c80f03b4c07bbecfaeb2156670f1ac

Observation 07e50d5f-b39a-42f3-8389-29f39ca7fc18 · inbound

Novel Multi-Agent Action Masked Deep Reinforcement Learning for General Industrial Assembly Lines Balancing Problems cites this paper.

Novel Multi-Agent Action Masked Deep Reinforcement Learning for General Industrial Assembly Lines Balancing Problems A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:12:00.444992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:12:00.444992Z digest=sha256:d1ff382f048dec695aeac2a497f13a1cdbc26d5a4fe3a1d981512974c3a7a091

Observation d0ad9eb3-73f8-444f-8359-c85db362407c · inbound

Learning to Assemble the Soma Cube with Legal-Action Masked DQN and Safe ZYZ Regrasp on a Doosan M0609 cites this paper.

Learning to Assemble the Soma Cube with Legal-Action Masked DQN and Safe ZYZ Regrasp on a Doosan M0609 A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T14:32:07.511541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:32:07.511541Z digest=sha256:a9ea4d9e1e957bbf479e6e02f3f888c824c05af1c588f23a9634b1e3410b9dd8

Observation fe0e8f53-bd8b-43a7-9692-8648a2b0ea01 · inbound

Towards Scalable O-RAN Resource Management: Graph-Augmented Proximal Policy Optimization cites this paper.

Towards Scalable O-RAN Resource Management: Graph-Augmented Proximal Policy Optimization A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:46:46.718720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:46:46.718720Z digest=sha256:744dda2beb2984d7b8ae2cf61a748d135adb021c36c3982e648151b6a9c2d14a

Observation b0b57996-e29b-451e-8ec2-88e49c4e411e · inbound

TARMM: Scaling Delay-Critical Edge AI Offloading in 5G O-RAN via Temporal Graph Mobility Management cites this paper.

TARMM: Scaling Delay-Critical Edge AI Offloading in 5G O-RAN via Temporal Graph Mobility Management A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:14.767914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T01:25:48.075081Z digest=sha256:781d227f439bf38602d10c0cda7ecea993d57f69367fa915a982741d38bbd6b6

Observation 0a71177f-d0ee-4528-9be5-39d1d5221667 · inbound

Your Loss is My Gain: Low Stake Attacks on Liquid Staking Pools cites this paper.

Your Loss is My Gain: Low Stake Attacks on Liquid Staking Pools A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:16:11.282290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T17:57:37.260492Z digest=sha256:cb662e399ce3d2384a339672d6aa661996500c4e8834839785684e4920b50637

Observation 4f5e053d-38c2-49ea-9cfe-528aecb8588a · inbound

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency cites this paper.

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:52:08.808054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T02:50:20.040271Z digest=sha256:05ce66a87d0058e49e9d13f9b1c4b9763ddbbebfbe275318c307f4736f61269f

Observation 6a89b502-5468-4592-97a4-a75c77dbc6b8 · inbound

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning cites this paper.

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:27:41.299409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T17:26:04.762451Z digest=sha256:21ec3226f12503c65bc85d9c58473eb3d68162fa8c5b101334a98a865c02b3b4

Observation 7f968414-1195-4db5-977e-e03a6be706bb · inbound

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning cites this paper.

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:15:04.608625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:08:47.084947Z digest=sha256:ba85f308c5726c1c0e4eb9d4cea26b9e4b97761add2449c4956effa4c1c30788

Observation 67e32243-3306-42fd-971d-72c4e23371c6 · inbound

AlphaTransit: Learning to Design City-scale Transit Routes cites this paper.

AlphaTransit: Learning to Design City-scale Transit Routes A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:24.317535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T11:55:20.454865Z digest=sha256:28cd8b75bbbc5b6e533c3602b732de081fafdda299ac14b11d9633ac74fc524c

Observation 7293490f-2098-479d-9857-e34000b2ab40 · inbound

Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets cites this paper.

Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:38.596865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T13:36:53.866075Z digest=sha256:3e50a11c8635ea013048c304aa327f9a5a4b291bd0fe69630477e011dd120990

Observation a3538820-553e-4d4c-ae3c-5905c9db9c56 · inbound

Deep RL for Fast Long-Horizon Operations Scheduling on NASA's Carruthers Geocorona Observatory Mission cites this paper.

Deep RL for Fast Long-Horizon Operations Scheduling on NASA's Carruthers Geocorona Observatory Mission A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:41.548227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T11:25:57.471803Z digest=sha256:d2ea7c2ecb6bdbf4e2692f8d1191b4972a799bf6a0f43f43f05975ec700a4b81

Observation 934b253b-b71e-4296-9722-fe10141e3034 · inbound

Optimal Reward Shaping: Autonomous Car Parking Case Study cites this paper.

Optimal Reward Shaping: Autonomous Car Parking Case Study A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-30T17:40:25.172202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T17:40:25.172202Z digest=sha256:99c879a194fa480b061a423af25434f42f3dbe21d88d91605dcf33e162ade58a