Pith. sign in

Paper Citation Record · LEDGER

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation

As of 12 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2501.12703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12703 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:56:51.117654Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 851083a3-8cf4-4036-a67c-4e29d48ebf7c · outbound

This paper cites Deep Reinforce- ment Learning for Robotic Manipulation—The State of the Art,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Deep Reinforce- ment Learning for Robotic Manipulation—The State of the Art,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.771394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:50.971230Z digest=sha256:4c48e8929bf1be16ca81e9464a5de14f9172712342367224d6e1eb08f69ff8a6

Observation fd00e4a1-9261-401e-8ac9-afcbafc6761a · outbound

This paper cites Trust Region Policy Optimization,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Trust Region Policy Optimization,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.756353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:50.994051Z digest=sha256:dd1e49ceb36a84feb78a84b1af3dde39fb7a060c7e765efe4f90eeaef06e22dd

Observation a2b5a92c-33fd-4cda-b83c-189239dfae9c · outbound

This paper cites Proximal Policy Optimization Algorithms.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Proximal Policy Optimization Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:50.989893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:50.989893Z digest=sha256:acb2bd7c162142f0dcb82446716f9e64de95125a4b8c0e60987dd4d60038056a

Observation 53dc0fd6-3ec5-445f-83a7-88984a35e637 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:51.004249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:51.004249Z digest=sha256:f0d45f7056ea824e42bfa86f713f3b59ffbee880b9dbc00305d73dc76c23719e

Observation d6c50502-6585-4964-a7e2-d8d07e4a37b2 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Adam: A Method for Stochastic Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:50.998963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:50.998963Z digest=sha256:b5b2edfe6b4b64721beaf2ee7d2998c9cc1e60c8fec1c3e8143663ff69edf028

Observation c80dd969-9d66-46f8-89a2-01d0cddfaf3c · outbound

This paper cites FIXAR: A Fixed-Point Deep Re- inforcement Learning Platform with Quantization-Aware Training and Adaptive Parallelism,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation FIXAR: A Fixed-Point Deep Re- inforcement Learning Platform with Quantization-Aware Training and Adaptive Parallelism,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.742467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.015420Z digest=sha256:b5b60157f43988f4406fee73441c08da179a21add93527bc17a0ae03ab18b147

Observation 370fcbf2-6990-4c08-9678-f851544824b3 · outbound

This paper cites QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:51.009947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:51.009947Z digest=sha256:b878529dea37e513190f9f81d6405c9e187fe743acd3a9e41a1e8573ea647e86

Observation f7fb7a99-274b-4858-96e1-152edea360ea · outbound

This paper cites EnvPool: A Highly Parallel Reinforcement Learning En- vironment Execution Engine,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation EnvPool: A Highly Parallel Reinforcement Learning En- vironment Execution Engine,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.714190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.025503Z digest=sha256:d0d8b768177dedabcfcc2796dc8bdcb4bd8e46f56aaf416f852ddb0af3553867

Observation 55aa3610-82b8-4d41-a3b8-acc07e81cc40 · outbound

This paper cites Accelerating Proximal Policy Optimization on CPU-FPGA Heterogeneous Platforms,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Accelerating Proximal Policy Optimization on CPU-FPGA Heterogeneous Platforms,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.728981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.020572Z digest=sha256:84a4f203cf5e60d65d3ad744ec38a09b0f2e011e690f88dd088a34f81c511a34

Observation cbdc85e0-98e7-4158-ada6-9c843614d871 · outbound

This paper cites GPU-Accelerated Robotic Simulation for Distributed Reinforcement Learning,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation GPU-Accelerated Robotic Simulation for Distributed Reinforcement Learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.682334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.034998Z digest=sha256:e2be69b87126806a04c817311f3342d89f436d501884bbc41471e787fbbe1ada

Observation f716cda3-6545-4b38-b74f-da7f535c7f6d · outbound

This paper cites Accelerating Reinforcement Learning through GPU Atari Emulation,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Accelerating Reinforcement Learning through GPU Atari Emulation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.698095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.030477Z digest=sha256:bb0710fbe4b16f93b9bb25c8b8cd631ddca81655ff3200dcfecae338a863909a

Observation b38a05ef-509b-443f-b3a2-95588fa65843 · outbound

This paper cites an unresolved cited work.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:56:51.652469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.049387Z digest=sha256:46f25fc16db9ef09e732827092fab7d3c3ab368313c011d073b7cbbd20f007fd

Observation 0b0eaba1-278a-4cbb-b6aa-c24862367dbc · outbound

This paper cites Note on a Method for Calculating Corrected Sums of Squares and Products,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Note on a Method for Calculating Corrected Sums of Squares and Products,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.667672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.039484Z digest=sha256:6fb35a43784cb34eb1ae39f75c6ece84f577bdb59a28f775e5225ce99c0d9dc5

Observation 96ab6666-9ee0-4b7c-9951-0a6590e6b8af · outbound

This paper cites (2018) Understanding Normalization of Advantage Function in PPO [Online].

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation (2018) Understanding Normalization of Advantage Function in PPO [Online]

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.606759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.063554Z digest=sha256:a21f0ae9eb1999acb6a0d8c9466b76daf8ba2ee570198cbd1e5b8bb5e55fc39a

Observation 4400bab7-7e5d-4998-bc50-f3338952fc91 · outbound

This paper cites an unresolved cited work.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:56:51.587589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.068212Z digest=sha256:7d7e99dcac561c93f0e4823b157da343b1ad620e9fe3e13ddf15a5a15f8948be

Observation 179541f6-90d1-4dbe-aa06-777520f4d671 · outbound

This paper cites an unresolved cited work.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:56:51.636925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.054097Z digest=sha256:5897f73265443550ad941dec723f1694a3fda72a1cdb009b4ce0839bcdbc0533

Observation 3251f2e3-1dcd-48ee-b6c9-159998dde902 · outbound

This paper cites Revisiting Deep Learning Paral- lelism: Fine-Grained Inference Engine Utilizing Online Arithmetic,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Revisiting Deep Learning Paral- lelism: Fine-Grained Inference Engine Utilizing Online Arithmetic,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.557903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.077437Z digest=sha256:fe8b89693990a58f2c198fa7f3fcd743f1199ace38bd47aa17aff3d6b138014e

Observation 701ae6d7-a28a-4dc8-81ed-d53969493647 · outbound

This paper cites Enabling Mixed-Timing NoCs for FPGAs: Reconfigurable Synthesizable Synchronization FIFOs,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Enabling Mixed-Timing NoCs for FPGAs: Reconfigurable Synthesizable Synchronization FIFOs,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.540618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.081737Z digest=sha256:5e381972a53154830db6d40f756628f22cc06e174f25bb18e51b6c9d7299de16

Observation 51f16659-a3a6-4dad-a98a-4e167a2b5d37 · outbound

This paper cites Reconfigurable Synthesizable Syn- chronization FIFOs,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Reconfigurable Synthesizable Syn- chronization FIFOs,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.525269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.085769Z digest=sha256:a1cbcba3cb66ebe748ec95040d375eaf5283f220e7736a308100369e89c1e63c

Observation 8aaf4c8c-1f35-48f8-a33e-c841b30cbe3d · outbound

This paper cites Safe Overclocking of Tightly Coupled CGRAs and Processor Arrays using Razor,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Safe Overclocking of Tightly Coupled CGRAs and Processor Arrays using Razor,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.572617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.072867Z digest=sha256:9d82d8356f74d8197b3dbc54c4084deeedcd70725e11d9c56dba112c318fb97a

Observation 2cf4b782-1809-47fa-846e-672adc3adb67 · outbound

This paper cites High-Throughput Synthesizable Synchronization FI- FOs for Mixed-Timing NoCs,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation High-Throughput Synthesizable Synchronization FI- FOs for Mixed-Timing NoCs,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.488365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.093537Z digest=sha256:67af95574968c2a9ce45816ac408720f0dcc611d48b3acc944ca246e0b01e295

Observation 67226e91-8ae9-4118-8765-1702d5debd09 · outbound

This paper cites Interleaved Architectures for High-Throughput Synthesizable Synchronization FIFOs,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Interleaved Architectures for High-Throughput Synthesizable Synchronization FIFOs,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.464824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.099058Z digest=sha256:619b0ce840a4c58614479ddca8b5b9203d23dd308a49aaa6084b080d6421c96f

Observation 6703a6f9-c379-4be3-9295-5f438e3b8517 · outbound

This paper cites Atalanta: A Bit is Worth a “Thousand.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Atalanta: A Bit is Worth a “Thousand

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.448593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.103856Z digest=sha256:cfbc63adbccf817708befe5ca2c6e87d3583c48ace1f03011f3681c10ae053ec

Observation a885b039-9f97-453a-8efe-d055e03434b6 · outbound

This paper cites Synthesizable Synchronization FIFOs Utilizing the Asynchronous Pulse-Based Handshake Protocol,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Synthesizable Synchronization FIFOs Utilizing the Asynchronous Pulse-Based Handshake Protocol,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.505668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.089604Z digest=sha256:a20128e88d8f3f0b7f8533ae49ef84c0b718beff757641c26c3e09ff3a26bb40

Observation 31363f23-6940-4a1b-9181-9ca4b59fa406 · outbound

This paper cites Boveda: Building an On-Chip Deep Learning Memory Hierarchy Brick by Brick,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Boveda: Building an On-Chip Deep Learning Memory Hierarchy Brick by Brick,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.415272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.112660Z digest=sha256:d5f1bcb37d3b7c8521c6ff8495d004ef24be552aa1d88339258432386e445705

Observation 7fc6d1aa-d000-4ca5-8945-078d9b7bdb80 · outbound

This paper cites Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.431929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.108343Z digest=sha256:fe28642c2fc0b02355f79b83d7f21dcc349b3b34d071b39249d9b23635b9a7c5

Observation 31dadffc-fdde-4cf5-8a32-e4db56796dcf · outbound

This paper cites Available: https://proceedings.mlsys.org/paper files/paper/ 2021/file/12a304a31e42dfefa21c82431e849124-Paper.pdf 9.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Available: https://proceedings.mlsys.org/paper files/paper/ 2021/file/12a304a31e42dfefa21c82431e849124-Paper.pdf 9

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.397499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.117654Z digest=sha256:3e07ea029d95ba6850253a5be7ba6b12689b468cbe61ecfb8909ff83b1c4d02f

Observation 54985f18-abc4-4ad2-8b2b-272fb6748f7b · outbound

This paper cites an unresolved cited work.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Unresolved cited work

Reference 485

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:56:51.621641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.059004Z digest=sha256:01de464da31016bc5166f99723e27d4d8088c82d6a15aa108e7a64f30d626c59

Observation 1e9d03f2-3603-451b-93ba-60e62f257be1 · outbound

This paper cites Available: https://www.jstor.org/stable/1266577.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Available: https://www.jstor.org/stable/1266577

Reference 1962

Resolution
verified exact
raw_fallback, observed 2026-08-10T16:56:51.310916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T16:56:51.044063Z digest=sha256:22383e8b992adc9dbd7fbb93e00adf61857fc0af8f8aaf96e023db4749de6a3d

Observation f6ad1393-60cf-4f1d-bc29-dc0df14e6709 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:50.980566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:50.980566Z digest=sha256:f7bcd692c9a45812ee7fbf8842839fe41f50a4a9f9999f35f581761bdde948aa

Pith citing papers

No inbound Pith citation observations are available.