Pith. sign in

Paper Citation Record · LEDGER

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation

As of 21 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2501.12703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12703 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:56:51.117654Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 851083a3-8cf4-4036-a67c-4e29d48ebf7c · outbound

This paper cites Deep Reinforce- ment Learning for Robotic Manipulation—The State of the Art,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Deep Reinforce- ment Learning for Robotic Manipulation—The State of the Art,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.771394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:50.971230Z digest=sha256:45dbea60c09e62db418c790a085e72c730a9f254c4c1f47bb37d5117ee3651f5

Observation fd00e4a1-9261-401e-8ac9-afcbafc6761a · outbound

This paper cites Trust Region Policy Optimization,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Trust Region Policy Optimization,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.756353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:50.994051Z digest=sha256:ccc09ef5cd7c9fbb31e6f8c3b3834503f18fe9e12dd6ef869f0f8f0e40716c84

Observation a2b5a92c-33fd-4cda-b83c-189239dfae9c · outbound

This paper cites Proximal Policy Optimization Algorithms.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Proximal Policy Optimization Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:50.989893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:50.989893Z digest=sha256:96015fa90dc27c61fa94b40afa4ab626c595908dc1b3c72139da9989300bbf71

Observation 53dc0fd6-3ec5-445f-83a7-88984a35e637 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:51.004249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:51.004249Z digest=sha256:24b8f6d46e25cc36f405ca024ab3a1d3343a4fd8f38f167d1b5fae1e044181bc

Observation d6c50502-6585-4964-a7e2-d8d07e4a37b2 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Adam: A Method for Stochastic Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:50.998963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:50.998963Z digest=sha256:fc464fcce0a48dda3d3e315e2ea0e7f56d1efe9db971f6b0695c5d7eba2f5859

Observation c80dd969-9d66-46f8-89a2-01d0cddfaf3c · outbound

This paper cites FIXAR: A Fixed-Point Deep Re- inforcement Learning Platform with Quantization-Aware Training and Adaptive Parallelism,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation FIXAR: A Fixed-Point Deep Re- inforcement Learning Platform with Quantization-Aware Training and Adaptive Parallelism,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.742467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.015420Z digest=sha256:0546813e1d2080318840b0f2517275fe90adde0e868faa83690d52c45c7922fa

Observation 370fcbf2-6990-4c08-9678-f851544824b3 · outbound

This paper cites QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:51.009947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:51.009947Z digest=sha256:a0d10e46f9466a6975347ea4d53e2c3413456a6799ed72eec0de92365cbbe32e

Observation f7fb7a99-274b-4858-96e1-152edea360ea · outbound

This paper cites EnvPool: A Highly Parallel Reinforcement Learning En- vironment Execution Engine,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation EnvPool: A Highly Parallel Reinforcement Learning En- vironment Execution Engine,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.714190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.025503Z digest=sha256:e57e9955dc5b840c7206193f7a5b791d9b40938f7d53d027a7933d8c909b68a2

Observation 55aa3610-82b8-4d41-a3b8-acc07e81cc40 · outbound

This paper cites Accelerating Proximal Policy Optimization on CPU-FPGA Heterogeneous Platforms,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Accelerating Proximal Policy Optimization on CPU-FPGA Heterogeneous Platforms,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.728981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.020572Z digest=sha256:771fc8d32d0c2fdf132de7287b85aa58e8486b03df0ecddaea095bb05753f890

Observation cbdc85e0-98e7-4158-ada6-9c843614d871 · outbound

This paper cites GPU-Accelerated Robotic Simulation for Distributed Reinforcement Learning,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation GPU-Accelerated Robotic Simulation for Distributed Reinforcement Learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.682334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.034998Z digest=sha256:ef554e2bb4c414a38c92b55cfe65921fdd34ba5bc349b4f36d181ac127044324

Observation f716cda3-6545-4b38-b74f-da7f535c7f6d · outbound

This paper cites Accelerating Reinforcement Learning through GPU Atari Emulation,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Accelerating Reinforcement Learning through GPU Atari Emulation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.698095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.030477Z digest=sha256:4ed95acb562a57ac850d92c6fca5eb7af9cffa959379c27d313b388051d9795a

Observation b38a05ef-509b-443f-b3a2-95588fa65843 · outbound

This paper cites an unresolved cited work.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:56:51.652469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.049387Z digest=sha256:80ffc18ed6159a4e229f35e27a82472452dd43781eab0db4a9c83e1c16bc280a

Observation 0b0eaba1-278a-4cbb-b6aa-c24862367dbc · outbound

This paper cites Note on a Method for Calculating Corrected Sums of Squares and Products,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Note on a Method for Calculating Corrected Sums of Squares and Products,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.667672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.039484Z digest=sha256:0bc75f35b48ba6cc8c848d52e3ec00578f217fe34847d7fbdfd498c77d2a5b90

Observation 96ab6666-9ee0-4b7c-9951-0a6590e6b8af · outbound

This paper cites (2018) Understanding Normalization of Advantage Function in PPO [Online].

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation (2018) Understanding Normalization of Advantage Function in PPO [Online]

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.606759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.063554Z digest=sha256:9bcac4980d48763b69ff32ba4cb24668b6bc7ab76172ee944b89a863004a5642

Observation 4400bab7-7e5d-4998-bc50-f3338952fc91 · outbound

This paper cites an unresolved cited work.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:56:51.587589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.068212Z digest=sha256:517cc812654196112dd7a49a8f3132b35f529afb008f4e540f4a2cfe6c8ec3c2

Observation 179541f6-90d1-4dbe-aa06-777520f4d671 · outbound

This paper cites an unresolved cited work.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:56:51.636925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.054097Z digest=sha256:42d858d76abd6ad0e3d94d4437009b5dbb1b0a0571a86b7844a0820b17d925a9

Observation 3251f2e3-1dcd-48ee-b6c9-159998dde902 · outbound

This paper cites Revisiting Deep Learning Paral- lelism: Fine-Grained Inference Engine Utilizing Online Arithmetic,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Revisiting Deep Learning Paral- lelism: Fine-Grained Inference Engine Utilizing Online Arithmetic,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.557903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.077437Z digest=sha256:d3d81b9bcbcec157461dd463a3afe1730cd482fe768966d6cb85afcf5d658595

Observation 701ae6d7-a28a-4dc8-81ed-d53969493647 · outbound

This paper cites Enabling Mixed-Timing NoCs for FPGAs: Reconfigurable Synthesizable Synchronization FIFOs,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Enabling Mixed-Timing NoCs for FPGAs: Reconfigurable Synthesizable Synchronization FIFOs,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.540618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.081737Z digest=sha256:7a7fac7a90f7f3d71d2d472711d5ed93c1c4ae07cd8884ea3b97506dfc103fa5

Observation 51f16659-a3a6-4dad-a98a-4e167a2b5d37 · outbound

This paper cites Reconfigurable Synthesizable Syn- chronization FIFOs,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Reconfigurable Synthesizable Syn- chronization FIFOs,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.525269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.085769Z digest=sha256:002b2125dad371fea10c0c55d106d96fe65a28ca940d0533c0b7a685a639d790

Observation 8aaf4c8c-1f35-48f8-a33e-c841b30cbe3d · outbound

This paper cites Safe Overclocking of Tightly Coupled CGRAs and Processor Arrays using Razor,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Safe Overclocking of Tightly Coupled CGRAs and Processor Arrays using Razor,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.572617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.072867Z digest=sha256:8fdc528fb0ab531a4a71edf4ba5bfa342500d510548a796ef8888f970c2558d8

Observation 2cf4b782-1809-47fa-846e-672adc3adb67 · outbound

This paper cites High-Throughput Synthesizable Synchronization FI- FOs for Mixed-Timing NoCs,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation High-Throughput Synthesizable Synchronization FI- FOs for Mixed-Timing NoCs,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.488365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.093537Z digest=sha256:f017286d3615178671082ed8366c46065d840e201580e91b66211f581745fca4

Observation 67226e91-8ae9-4118-8765-1702d5debd09 · outbound

This paper cites Interleaved Architectures for High-Throughput Synthesizable Synchronization FIFOs,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Interleaved Architectures for High-Throughput Synthesizable Synchronization FIFOs,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.464824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.099058Z digest=sha256:a9215456ddfbfa684569e54297f236ced6c0f042e752e8116b8f8a45fd1e8e46

Observation 6703a6f9-c379-4be3-9295-5f438e3b8517 · outbound

This paper cites Atalanta: A Bit is Worth a “Thousand.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Atalanta: A Bit is Worth a “Thousand

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.448593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.103856Z digest=sha256:6d0c5115434192d684e882356ac3fc4160ab37ab635d9c5eb0d0f6c26d764ecd

Observation a885b039-9f97-453a-8efe-d055e03434b6 · outbound

This paper cites Synthesizable Synchronization FIFOs Utilizing the Asynchronous Pulse-Based Handshake Protocol,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Synthesizable Synchronization FIFOs Utilizing the Asynchronous Pulse-Based Handshake Protocol,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.505668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.089604Z digest=sha256:3e56ea17051249bc9265547c495dd3395823da534f62a1c0c0351d738e2ca2d1

Observation 31363f23-6940-4a1b-9181-9ca4b59fa406 · outbound

This paper cites Boveda: Building an On-Chip Deep Learning Memory Hierarchy Brick by Brick,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Boveda: Building an On-Chip Deep Learning Memory Hierarchy Brick by Brick,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.415272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.112660Z digest=sha256:2e1dafbdba3fee21fb9dd0650dbc95119f942dcbe78d98cf2a9fb92fa879a6e1

Observation 7fc6d1aa-d000-4ca5-8945-078d9b7bdb80 · outbound

This paper cites Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models,.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.431929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.108343Z digest=sha256:197311a7b0104efb7cb313d24df89461748c6e3e9b4613107115f1e4d09f687e

Observation 31dadffc-fdde-4cf5-8a32-e4db56796dcf · outbound

This paper cites Available: https://proceedings.mlsys.org/paper files/paper/ 2021/file/12a304a31e42dfefa21c82431e849124-Paper.pdf 9.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Available: https://proceedings.mlsys.org/paper files/paper/ 2021/file/12a304a31e42dfefa21c82431e849124-Paper.pdf 9

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:56:51.397499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.117654Z digest=sha256:163333d246901bf4d47173637a010db0d6effada30d23bcfa53871b5b51696b3

Observation 54985f18-abc4-4ad2-8b2b-272fb6748f7b · outbound

This paper cites an unresolved cited work.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Unresolved cited work

Reference 485

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:56:51.621641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.059004Z digest=sha256:c8a8f5f04583169906180da921b99288dc51227417c157de835971d3b90ed9d1

Observation 1e9d03f2-3603-451b-93ba-60e62f257be1 · outbound

This paper cites Available: https://www.jstor.org/stable/1266577.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Available: https://www.jstor.org/stable/1266577

Reference 1962

Resolution
verified exact
raw_fallback, observed 2026-08-10T16:56:51.310916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T16:56:51.044063Z digest=sha256:177d4ad36c34fd38d1caca4f1cde18758f890c4ddbb788051e440e690fed0697

Observation f6ad1393-60cf-4f1d-bc29-dc0df14e6709 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:50.980566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:50.980566Z digest=sha256:e27def8870e053365ff6bdda553e3377bbc9941dd91c1b5a4e911bc0ef05fdc8

Pith citing papers

No inbound Pith citation observations are available.