Pith. sign in

Paper Citation Record · LEDGER

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization

As of 20 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2507.03372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03372 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:25:09.694462Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98c75044-eebe-4946-94c5-7ce03341af9d · outbound

This paper cites Reinforcement learning based recommender systems: A survey.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Reinforcement learning based recommender systems: A survey

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:15.221666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:06.596567Z digest=sha256:b949a96480ec75038853153bec38fed413845f0a317ffe64e474731ca644e3bd

Observation 9c0cddd7-2873-4880-b392-cd6d4553824d · outbound

This paper cites Safe learning in robotics: From learning-based control to safe reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Safe learning in robotics: From learning-based control to safe reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:15.105289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:06.686440Z digest=sha256:dedd0fd10aaacd3ad43a902976add1d4cdebdf27270c0f7e3280418d5f913459

Observation 6ea39787-c015-42f7-97d4-87ef7aa49182 · outbound

This paper cites Certifiable robustness to adversarial state uncertainty in deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Certifiable robustness to adversarial state uncertainty in deep reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.960200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:06.828541Z digest=sha256:68e8473b3195bf6ad81543fec38de2dcc6cd5983175e2ed8d6b939a0c976dc5a

Observation caff1573-0394-4bcb-b9fa-a300889a184c · outbound

This paper cites Maximum entropy RL (provably) solves some robust RL problems.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Maximum entropy RL (provably) solves some robust RL problems

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.759489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:06.954194Z digest=sha256:559ff34e5d144ca5bea6ae087272aa953c6eae5a30bb52665ded91598a5c1a5e

Observation 7d63bf7e-49bc-4d6c-b52e-e2b52d525f58 · outbound

This paper cites Online Robustness Training for Deep Reinforcement Learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Online Robustness Training for Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.041555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.041555Z digest=sha256:97ee34f5eb62d2399587925db51081e726be398e882db6b7fb2480027c524434

Observation 9943ad94-02b5-4333-bd8c-c41cc21ea8b6 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.152922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.152922Z digest=sha256:a3ac308740b819c3f1627456d620bf2db531e4494df1e72c143cc252128a8a0c

Observation dbc958f2-f026-4e4a-ba4b-5ff51183c325 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Addressing function approximation error in actor-critic methods

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.297412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.297412Z digest=sha256:69c433442f6c844b107e194c834b8eeecbe74b4f4f5d835ac232df9e3b44dd21

Observation 1cba9500-13f1-4507-9147-7826c7103d52 · outbound

This paper cites an unresolved cited work.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:25:14.501113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:07.392711Z digest=sha256:bf81753f154eed1eaa78f3430c0bdde0699d635634e3953fef84a936f47499bf

Observation 2dcf1263-1545-4cc2-b7d7-93d9f92998be · outbound

This paper cites Adversarial Attacks on Neural Network Policies.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Adversarial Attacks on Neural Network Policies

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.512173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.512173Z digest=sha256:dd49f7a01c26ca0c8594e84cec80d114462030242a9592a127377d4d6689c7af

Observation 9a6075f1-d532-4554-b8a6-80ae8c4bdf95 · outbound

This paper cites The 37 implementation details of proximal policy optimization.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization The 37 implementation details of proximal policy optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.618873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.618873Z digest=sha256:2595073c990f93358a325575e2409ea158fa14eba9a0d9e79ef0fbdfd5b7a763

Observation 53f44c71-7d3a-441d-b1aa-caa90f6f63e8 · outbound

This paper cites Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.252479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:07.772234Z digest=sha256:4fbd7e9e1bfac384fd0b70c9f201b03b74723802034ca473ee27d0b9036ae81d

Observation e464de02-4f51-4883-a397-cc3af2a83e52 · outbound

This paper cites Challenges and countermeasures for adversarial attacks on deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Challenges and countermeasures for adversarial attacks on deep reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.885692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.885692Z digest=sha256:d4628fe94cade97d7571235f745484b8e6277af0f768245f5a7417f7be9ba5d7

Observation a3cfaeaf-d03f-4b78-af8a-57a8459b7479 · outbound

This paper cites Learning quadrupedal locomotion over challenging terrain.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Learning quadrupedal locomotion over challenging terrain

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.996826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.996826Z digest=sha256:f23e64f425e31908374b7b1ff2f92cca3c403e6dc66c21944529257df6cf9e78

Observation 44e43070-404d-4db8-bdd2-7d4cb313cc98 · outbound

This paper cites Spatiotem- porally constrained action space attacks on deep reinforcement learning agents.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Spatiotem- porally constrained action space attacks on deep reinforcement learning agents

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.094582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.066905Z digest=sha256:0046b109dd4989137ce67d48233aa9dac926a0a2b38b1f28d3dcea72ebf94b95

Observation a10a8786-ad85-484e-b1cc-423d7da93dcb · outbound

This paper cites Efficient adversarial training without attacking: Worst-case-aware robust reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Efficient adversarial training without attacking: Worst-case-aware robust reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.974490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.139880Z digest=sha256:8b0d84533fe9dd2496a13c269b7207a7371c3fed9dad1d34346b427969959c19

Observation 0b0ba7e6-65fc-4a8d-b3a3-8a788f77ac76 · outbound

This paper cites Provably efficient black-box action poisoning attacks against reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Provably efficient black-box action poisoning attacks against reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.807962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.198726Z digest=sha256:13df82ed34793597de6a2d966d5d302ec2edb8e2aa8978b2ca92bd366f545294

Observation c7fdbf05-08ce-4686-8335-78b600678f61 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Towards deep learning models resistant to adversarial attacks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.649352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.277644Z digest=sha256:42ab73104a7ba03c5c024cf42ae4d499f5854d3f7f7c3b004cfaf0b946509fc6

Observation c0c52c15-98af-43dc-b174-6a9692e6d100 · outbound

This paper cites Human-level control through deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Human-level control through deep reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:08.372071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:08.372071Z digest=sha256:09b5e6aafc1d3ecd5d72611ff262bccaade4ff744b638535a7d55012896e9177

Observation 47840ea9-ebe9-4d31-a407-a54d4a0d794d · outbound

This paper cites Robust reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.511735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.460059Z digest=sha256:8eb71cdc0b9b03055fa2ca1a8fc7ec805077d4e013a673874cfcd18312d96170

Observation 87f4595a-ac59-4e40-88b9-c9c7c1bd5633 · outbound

This paper cites Assessing transferability from simulation to reality for reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Assessing transferability from simulation to reality for reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.350240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.546420Z digest=sha256:c2754ad87a2dd9230da1f58c6c99e614e6ee98cf07ccdbb6996ec60aa69a3fcf

Observation 0d15a0f3-b53f-4456-b67d-291feaba836a · outbound

This paper cites Robust deep reinforcement learning through adversarial loss.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust deep reinforcement learning through adversarial loss

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.050191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.594896Z digest=sha256:e70e9e7d8cd1c8fbbf9f290b40e727a15b2e263ea4588368220ce36cbccfe2c5

Observation e6659db3-eda5-44b5-b5d2-b16b34babfa5 · outbound

This paper cites Characterizing attacks on deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Characterizing attacks on deep reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.900816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.651218Z digest=sha256:bdcb34a873574f3efb8bc7a2f09b1cc2cdcb75e3637e598a841d58e2fd012d43

Observation 3815608d-00c1-4f3d-b6af-766565bfedc9 · outbound

This paper cites Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.720848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.722490Z digest=sha256:74fa0b16f0434e983b29aa17aacc09c3a9ae40d32c41ffb89c651a74fb1cdce5

Observation 94d03389-a69b-4bc9-b3ad-29fe3d273a85 · outbound

This paper cites The security of autonomous driving: Threats, defenses, and future directions.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization The security of autonomous driving: Threats, defenses, and future directions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.587133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.795215Z digest=sha256:1a4e0e63d77c231dae3b9dd2b6786f0a60ee5add90932ae405781d6368647530

Observation 3bbd293c-bd7a-49d5-ba9e-9413aa00170f · outbound

This paper cites Learning to walk in minutes using massively parallel deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Learning to walk in minutes using massively parallel deep reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:08.834445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:08.834445Z digest=sha256:a210bf0665e24bf682e93e1b0d3075c67f5b8ed9caafc63d56cc8b43587e7f03

Observation 3611f0eb-36b1-410b-b54d-49b501b96fe0 · outbound

This paper cites Improving robotic machining accuracy through experimental error investigation and modular compensation.The International Journal of Advanced Manufacturing Technology, 85:3–15, 2016.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Improving robotic machining accuracy through experimental error investigation and modular compensation.The International Journal of Advanced Manufacturing Technology, 85:3–15, 2016

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.445899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.885713Z digest=sha256:02713fc1b31cd89257e04844f48258025427f186bb841e5e18cb820969672b15

Observation a7d3b7ca-75f7-4bcf-8e53-e080d83928f5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:08.934071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:08.934071Z digest=sha256:eb07af22f4bda847785ac5b10bfa4e702fc906e10f3dd0f94f54e0079eebb1c3

Observation 44a3e993-a13a-4438-b3c0-8aaf2e46b642 · outbound

This paper cites Towards facilitating empathic conversations in online mental health support: A reinforcement learning approach.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Towards facilitating empathic conversations in online mental health support: A reinforcement learning approach

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.264004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.949626Z digest=sha256:ece7121535ca63f441da366bfca951a81315a2f3c3986b7375174ac5ee5c7165

Observation b967c140-f346-4417-b157-dbb834955c42 · outbound

This paper cites Certifiably robust policy learning against adversarial multi-agent communication.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Certifiably robust policy learning against adversarial multi-agent communication

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.092515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:08.990285Z digest=sha256:4ef2b019d5d76cae16a373091d62d333daf23ee3818855a150bf473b8e1b0744

Observation 33ec714b-5e8c-447e-b95f-3630aa23a6ec · outbound

This paper cites Certifiably robust policy learning against adversarial multi-agent communication.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Certifiably robust policy learning against adversarial multi-agent communication

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.940979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.047922Z digest=sha256:cff04b721752f9790fe3739f582f6bafc27f54e5b0284d3e6ef9f01655f6eb75

Observation ea130755-e901-474a-8744-0245ecee41df · outbound

This paper cites Who is the strongest enemy? towards optimal and efficient evasion attacks in deep RL.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Who is the strongest enemy? towards optimal and efficient evasion attacks in deep RL

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.706675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.087983Z digest=sha256:cecd1a14e7dbb498cfb14682782809abcdfd4a1276336e6f90284681f74690cd

Observation f76b6a1e-e35c-444f-9a36-5e76b5e90792 · outbound

This paper cites Reinforcement learning: An introduction.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Reinforcement learning: An introduction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:09.150195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:09.150195Z digest=sha256:f808d181227035cab4dbb6bba8171c8228cb1a4b31122e3e8eaf27c60229091f

Observation 2b867fe0-f00b-40ac-a803-ad3a36575da2 · outbound

This paper cites Action robust reinforcement learning and applications in continuous control.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Action robust reinforcement learning and applications in continuous control

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.495463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.207310Z digest=sha256:f62227ae6ed74134f7addd8d0d7ae74aef1e27b18c8a8a7ce1809c53ff081933

Observation 6323ea20-c761-4d13-9fbf-80fb45deebed · outbound

This paper cites Mujoco: A physics engine for model-based control.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Mujoco: A physics engine for model-based control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:09.241906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:09.241906Z digest=sha256:8a455cc08401f38ddb51a04b4eed5c622fdbbae90bc82102a774c2b83fe51aec

Observation 4086eb8d-ac23-4e96-9a29-ee8438a40774 · outbound

This paper cites Policy gradient method for robust reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Policy gradient method for robust reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.323799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.294486Z digest=sha256:7cb3a2484eff34bcc62e4b75b94fb2de0bb25e29e3b7ffe256f97322b00f86f1

Observation e99860d5-e676-412f-b7bc-4ca207a87a21 · outbound

This paper cites CROP: Certify- ing robust policies for reinforcement learning through functional smoothing.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization CROP: Certify- ing robust policies for reinforcement learning through functional smoothing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.035690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.363296Z digest=sha256:a94b91dd22486ca4328e3ae822fe0c29bb504b9736c0833e0dc00d08ff6eb705

Observation ca13e76a-64db-4e37-9bd1-b275932fc9a5 · outbound

This paper cites Robust deep reinforcement learning through bootstrapped opportunistic curriculum.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust deep reinforcement learning through bootstrapped opportunistic curriculum

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.903033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.388391Z digest=sha256:8b8164a93921bde7178c19a26cb1a54e7645b228216b3769c51d8b9da16c2b3f

Observation b5b0e446-6c9a-4989-a042-cd08b4221c80 · outbound

This paper cites Taac: Temporally abstract actor-critic for continuous control.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Taac: Temporally abstract actor-critic for continuous control

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.739145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.444717Z digest=sha256:c3f6f265eaaa08f25bb04e52bc97eaa26da1c08660f9e644a600c2318be1e5ef

Observation 1664edd7-da4d-4e27-bbde-ce91f91b6f18 · outbound

This paper cites Gradient surgery for multi-task learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Gradient surgery for multi-task learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:09.492261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:09.492261Z digest=sha256:412e51f88a35c798267573050e38a9f0b87aaf3a30f386af9d304f5605914d8d

Observation 92a427cb-3d4c-4772-be7c-b26a64e2dedb · outbound

This paper cites Robust reinforcement learning on state observations with learned optimal adversary.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust reinforcement learning on state observations with learned optimal adversary

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.561888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.519045Z digest=sha256:382ea631d8ce3c1e70088bee8dd6ffc6ff689c09fa1512a2f484c8206be8e7b5

Observation c07a37e2-7e1f-4c66-9f37-2339c35d2183 · outbound

This paper cites Robust deep reinforcement learning against adversarial perturbations on state observations.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust deep reinforcement learning against adversarial perturbations on state observations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.429921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.578035Z digest=sha256:4a858122867e7e7578b75efb0edd33c6ca3cdb9a97cf5d120746921117a3bb6f

Observation 9d2ed019-a327-4c39-9998-5c210206e090 · outbound

This paper cites Adaptive reward-poisoning attacks against reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Adaptive reward-poisoning attacks against reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.215720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.635734Z digest=sha256:8f2b221aecc91bc52c60792e498db426dd26edb7a82d9b8bda2519fe11b803b6

Observation dac71140-d879-4c66-acb0-e7b5eb92cdcb · outbound

This paper cites Sim-to-real transfer in deep reinforcement learning for robotics: a survey.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Sim-to-real transfer in deep reinforcement learning for robotics: a survey

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.069216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.661163Z digest=sha256:7de22c37a0475105008ddad00477f13266815ec664b0e0b3bdc2ef52ca0cb6d4

Observation cde022d4-24c3-4039-9dc4-a14562d2e012 · outbound

This paper cites Cadre: A cascade deep reinforcement learning framework for vision-based autonomous urban driving.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Cadre: A cascade deep reinforcement learning framework for vision-based autonomous urban driving

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:09.926141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:25:09.694462Z digest=sha256:b565cd4271a587b02a8139b3b5341bf8d2f8658aeb94243169b1ddcd4527412f

Pith citing papers

No inbound Pith citation observations are available.