Pith. sign in

REVIEW 2 cited by

QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.01055 v6 pith:QU4R2H37 submitted 2019-10-02 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords learningreinforcementactorqtextittrainingtimesblueachieving
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Deep reinforcement learning continues to show tremendous potential in achieving task-level autonomy, however, its computational and energy demands remain prohibitively high. In this paper, we tackle this problem by applying quantization to reinforcement learning. To that end, we introduce a novel Reinforcement Learning (RL) training paradigm, \textit{ActorQ}, to speed up actor-learner distributed RL training. \textit{ActorQ} leverages 8-bit quantized actors to speed up data collection without affecting learning convergence. Our quantized distributed RL training system, \textit{ActorQ}, demonstrates end-to-end speedups \blue{between 1.5 $\times$ and 5.41$\times$}, and faster convergence over full precision training on a range of tasks (Deepmind Control Suite) and different RL algorithms (D4PG, DQN). Furthermore, we compare the carbon emissions (Kgs of CO2) of \textit{ActorQ} versus standard reinforcement learning \blue{algorithms} on various tasks. Across various settings, we show that \textit{ActorQ} enables more environmentally friendly reinforcement learning by achieving \blue{carbon emission improvements between 1.9$\times$ and 3.76$\times$} compared to training RL-agents in full-precision. We believe that this is the first of many future works on enabling computationally energy-efficient and sustainable reinforcement learning. The source code is available here for the public to use: \url{https://github.com/harvard-edge/QuaRL}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation

    cs.AR 2025-01 reject novelty 5.0 of 10

    HEPPO-GAE is an FPGA design that pipelines GAE with 8-bit quantization and standardization, claiming 30% PPO speedup and 1.5x rewards, with speedups estimated rather than measured.

  2. Characterization and Mitigation of ADC Noise by Reference Tuning in RRAM-Based Compute-In-Memory

    cs.ET 2025-02 reject novelty 4.0 of 10

    A measured effective-bit noise model is used to simulate accuracy loss on four workloads; per-module or per-ADC reference tuning helps some tasks, but the drone result contradicts the claimed general effectiveness.

Pith tools