Pith. sign in

REVIEW 1 cited by

On Robust Reinforcement Learning with Lipschitz-Bounded Policy Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11432 v3 pith:RGHVBN5U submitted 2024-05-19 cs.LG cs.SYeess.SY

classification cs.LGcs.SYeess.SY
keywords lipschitznetworkspolicyperformancerobustcleanlayerlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a study of robust policy networks in deep reinforcement learning. We investigate the benefits of policy parameterizations that naturally satisfy constraints on their Lipschitz bound, analyzing their empirical performance and robustness on two representative problems: pendulum swing-up and Atari Pong. We illustrate that policy networks with smaller Lipschitz bounds are more robust to disturbances, random noise, and targeted adversarial attacks than unconstrained policies composed of vanilla multi-layer perceptrons or convolutional neural networks. However, the structure of the Lipschitz layer is important. We find that the widely-used method of spectral normalization is too conservative and severely impacts clean performance, whereas more expressive Lipschitz layers such as the recently-proposed Sandwich layer can achieve improved robustness without sacrificing clean performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Scalable Approach for Safe and Robust Learning via Lipschitz-Constrained Networks

    cs.LG 2025-06 reject novelty 3.0 of 10

    A loop-transformation convexification plus a randomized subspace sketch for Lipschitz-constrained training; the sketch's high-probability certificate is not mathematically justified.

Pith tools