Pith. sign in

Paper Citation Record · LEDGER

A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2006.14171.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.14171 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:07:48.369414Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:39:41.546939Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3793f95a-11f8-46cf-86c2-36f4bc144295 · inbound

Dynamic Collaborative Material Distribution System for Intelligent Robots In Smart Manufacturing cites this paper.

Dynamic Collaborative Material Distribution System for Intelligent Robots In Smart Manufacturing A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:48.369414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:48.369414Z digest=sha256:7ddfed56848542e4102a71f5f0f29daa52fd39e85669fb45c34938be4eecfe4b

Observation 53be11a9-015f-4c00-9348-8634fb2cef0c · inbound

Data-Driven Policy Mapping for Safe RL-based Energy Management Systems cites this paper.

Data-Driven Policy Mapping for Safe RL-based Energy Management Systems A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:53.981879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:53.981879Z digest=sha256:4aa9fd50b69400bb7c23f10ca791f9a7e5c80f03b4c07bbecfaeb2156670f1ac

Observation 07e50d5f-b39a-42f3-8389-29f39ca7fc18 · inbound

Novel Multi-Agent Action Masked Deep Reinforcement Learning for General Industrial Assembly Lines Balancing Problems cites this paper.

Novel Multi-Agent Action Masked Deep Reinforcement Learning for General Industrial Assembly Lines Balancing Problems A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:12:00.444992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:12:00.444992Z digest=sha256:d1ff382f048dec695aeac2a497f13a1cdbc26d5a4fe3a1d981512974c3a7a091

Observation d0ad9eb3-73f8-444f-8359-c85db362407c · inbound

Learning to Assemble the Soma Cube with Legal-Action Masked DQN and Safe ZYZ Regrasp on a Doosan M0609 cites this paper.

Learning to Assemble the Soma Cube with Legal-Action Masked DQN and Safe ZYZ Regrasp on a Doosan M0609 A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T14:32:07.511541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:32:07.511541Z digest=sha256:a9ea4d9e1e957bbf479e6e02f3f888c824c05af1c588f23a9634b1e3410b9dd8

Observation fe0e8f53-bd8b-43a7-9692-8648a2b0ea01 · inbound

Towards Scalable O-RAN Resource Management: Graph-Augmented Proximal Policy Optimization cites this paper.

Towards Scalable O-RAN Resource Management: Graph-Augmented Proximal Policy Optimization A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:46:46.718720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:46:46.718720Z digest=sha256:744dda2beb2984d7b8ae2cf61a748d135adb021c36c3982e648151b6a9c2d14a

Observation b0b57996-e29b-451e-8ec2-88e49c4e411e · inbound

TARMM: Scaling Delay-Critical Edge AI Offloading in 5G O-RAN via Temporal Graph Mobility Management cites this paper.

TARMM: Scaling Delay-Critical Edge AI Offloading in 5G O-RAN via Temporal Graph Mobility Management A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:14.767914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T01:25:48.075081Z digest=sha256:63841e3544f05bf38bc908f2c0b9013de0852be0223b6c04a6737243fb5c6918

Observation 0a71177f-d0ee-4528-9be5-39d1d5221667 · inbound

Your Loss is My Gain: Low Stake Attacks on Liquid Staking Pools cites this paper.

Your Loss is My Gain: Low Stake Attacks on Liquid Staking Pools A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:16:11.282290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T17:57:37.260492Z digest=sha256:1ccda3db6dbe3eb6bb69cc52477fe222fa52f1130b83427a21546e9437a9d698

Observation 4f5e053d-38c2-49ea-9cfe-528aecb8588a · inbound

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency cites this paper.

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:52:08.808054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T02:50:20.040271Z digest=sha256:f168597a5e9c10e07d6497418f76367ba8a886fdf122a9e538c7f897047f9b0c

Observation 6a89b502-5468-4592-97a4-a75c77dbc6b8 · inbound

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning cites this paper.

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:27:41.299409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:26:04.762451Z digest=sha256:fa723ad419db430e9f9cc2529e4a5831cc992550f9ee2ec0b963d35b3e948b58

Observation 7f968414-1195-4db5-977e-e03a6be706bb · inbound

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning cites this paper.

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:15:04.608625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:08:47.084947Z digest=sha256:e22afb4fd9205a97918aea6e2b2fe2e684abdce8542b9d7059c946b5343ea1ac

Observation 67e32243-3306-42fd-971d-72c4e23371c6 · inbound

AlphaTransit: Learning to Design City-scale Transit Routes cites this paper.

AlphaTransit: Learning to Design City-scale Transit Routes A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:24.317535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T11:55:20.454865Z digest=sha256:d0207357f0c2dc7f528c126147a453210211d62250d3058c1ac998da3766a083

Observation 7293490f-2098-479d-9857-e34000b2ab40 · inbound

Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets cites this paper.

Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:38.596865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:36:53.866075Z digest=sha256:13281f046299c9f1434f06c0aa71c954aa9f53f66370b397bd368a84c53dea74

Observation a3538820-553e-4d4c-ae3c-5905c9db9c56 · inbound

Deep RL for Fast Long-Horizon Operations Scheduling on NASA's Carruthers Geocorona Observatory Mission cites this paper.

Deep RL for Fast Long-Horizon Operations Scheduling on NASA's Carruthers Geocorona Observatory Mission A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:41.548227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T11:25:57.471803Z digest=sha256:ad046a7575f065dd081a4691550d0f2711c66dbba081ffb1f9f6c6e161cfeca9

Observation 934b253b-b71e-4296-9722-fe10141e3034 · inbound

Optimal Reward Shaping: Autonomous Car Parking Case Study cites this paper.

Optimal Reward Shaping: Autonomous Car Parking Case Study A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-30T17:40:25.172202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T17:40:25.172202Z digest=sha256:99c879a194fa480b061a423af25434f42f3dbe21d88d91605dcf33e162ade58a