Pith. sign in

Paper Citation Record · LEDGER

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2507.03372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03372 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:25:09.694462Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98c75044-eebe-4946-94c5-7ce03341af9d · outbound

This paper cites Reinforcement learning based recommender systems: A survey.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Reinforcement learning based recommender systems: A survey

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:15.221666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:06.596567Z digest=sha256:c1305f5d0409ffe702c52b1d70c503e503c850d46194d778907d05bb6b004b4a

Observation 9c0cddd7-2873-4880-b392-cd6d4553824d · outbound

This paper cites Safe learning in robotics: From learning-based control to safe reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Safe learning in robotics: From learning-based control to safe reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:15.105289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:06.686440Z digest=sha256:d1edcf4236d1d25e3d5275b5e6c1f58565cd14e543842981d2e84ef35adcb8c0

Observation 6ea39787-c015-42f7-97d4-87ef7aa49182 · outbound

This paper cites Certifiable robustness to adversarial state uncertainty in deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Certifiable robustness to adversarial state uncertainty in deep reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.960200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:06.828541Z digest=sha256:6a26f0d4165a7bf573a8c86bc919189a8f598c1e02c431d0f1c6e6a96c1a7d13

Observation caff1573-0394-4bcb-b9fa-a300889a184c · outbound

This paper cites Maximum entropy RL (provably) solves some robust RL problems.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Maximum entropy RL (provably) solves some robust RL problems

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.759489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:06.954194Z digest=sha256:6a1acc3c764c2562acdf171b43e5f5641506d36b4e9a0eed27107a1ab7138300

Observation 7d63bf7e-49bc-4d6c-b52e-e2b52d525f58 · outbound

This paper cites Online Robustness Training for Deep Reinforcement Learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Online Robustness Training for Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.041555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.041555Z digest=sha256:cc8f5352f77f0c1a9113011a1a927ade74691fd9486376f24a51b7bf7ffd0dfc

Observation 9943ad94-02b5-4333-bd8c-c41cc21ea8b6 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.152922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.152922Z digest=sha256:5253a3377653756665d36fbd3ad10d4f33df9296c597e9810d907d5ee9e5c277

Observation dbc958f2-f026-4e4a-ba4b-5ff51183c325 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Addressing function approximation error in actor-critic methods

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.297412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.297412Z digest=sha256:a3e8f04810665afae4ee9d9a9533e7693ca6ee64e67a7a6afb65c2f607df0497

Observation 1cba9500-13f1-4507-9147-7826c7103d52 · outbound

This paper cites an unresolved cited work.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:25:14.501113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:07.392711Z digest=sha256:f1f79638ab5a43c16bcf6583fc7b2edbcc52c67a247e8e7543fb25e9996dc5de

Observation 2dcf1263-1545-4cc2-b7d7-93d9f92998be · outbound

This paper cites Adversarial Attacks on Neural Network Policies.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Adversarial Attacks on Neural Network Policies

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.512173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.512173Z digest=sha256:1181d690d3e002f1f817b80de5685a62241b07f82a57ca2a22d7e3809aecbac3

Observation 9a6075f1-d532-4554-b8a6-80ae8c4bdf95 · outbound

This paper cites The 37 implementation details of proximal policy optimization.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization The 37 implementation details of proximal policy optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.618873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.618873Z digest=sha256:8e4003777731f0d4daf9531bf42fbf9a7c1b6cf9385de77bb2227555def23bbd

Observation 53f44c71-7d3a-441d-b1aa-caa90f6f63e8 · outbound

This paper cites Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.252479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:07.772234Z digest=sha256:7f200be97184bc62686a57ef9108b7914d4858daf4c212f1480afb9d25f37163

Observation e464de02-4f51-4883-a397-cc3af2a83e52 · outbound

This paper cites Challenges and countermeasures for adversarial attacks on deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Challenges and countermeasures for adversarial attacks on deep reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.885692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.885692Z digest=sha256:ff2f1b9a253505af41ab3315f8cf3075907db910810ce691abf921af6a48d9b3

Observation a3cfaeaf-d03f-4b78-af8a-57a8459b7479 · outbound

This paper cites Learning quadrupedal locomotion over challenging terrain.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Learning quadrupedal locomotion over challenging terrain

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.996826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.996826Z digest=sha256:fcab4e3d2a13c29ca28bf07ad66bda39d916460065097e6b8a4c5b4a5c337bcd

Observation 44e43070-404d-4db8-bdd2-7d4cb313cc98 · outbound

This paper cites Spatiotem- porally constrained action space attacks on deep reinforcement learning agents.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Spatiotem- porally constrained action space attacks on deep reinforcement learning agents

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.094582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.066905Z digest=sha256:d04afb9a67330513b3a51c55be6ca53c51fb4862654fc167f27b63b0dcdbe1c6

Observation a10a8786-ad85-484e-b1cc-423d7da93dcb · outbound

This paper cites Efficient adversarial training without attacking: Worst-case-aware robust reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Efficient adversarial training without attacking: Worst-case-aware robust reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.974490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.139880Z digest=sha256:9eb4c8cb2ab502aeee6eb0aeaf7c06624088281e9d6bae93ab364da4ecf5d9dc

Observation 0b0ba7e6-65fc-4a8d-b3a3-8a788f77ac76 · outbound

This paper cites Provably efficient black-box action poisoning attacks against reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Provably efficient black-box action poisoning attacks against reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.807962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.198726Z digest=sha256:7611c451ad4ec99fb0bb6b4841d313069feb1a976c4c6b7f71a188975036f0f8

Observation c7fdbf05-08ce-4686-8335-78b600678f61 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Towards deep learning models resistant to adversarial attacks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.649352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.277644Z digest=sha256:0be5408c9b08027e5be530a32f53595e05b59e1fc0a82faa6cff25a5b99e3f5e

Observation c0c52c15-98af-43dc-b174-6a9692e6d100 · outbound

This paper cites Human-level control through deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Human-level control through deep reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:08.372071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:08.372071Z digest=sha256:e44c714c6275c120146acff593fcedad850096a7f037b5c0d51990116fbd2283

Observation 47840ea9-ebe9-4d31-a407-a54d4a0d794d · outbound

This paper cites Robust reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.511735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.460059Z digest=sha256:e88b328e29b20e053354cd196960b8332277c857a31472ac2f3c37639fb5516a

Observation 87f4595a-ac59-4e40-88b9-c9c7c1bd5633 · outbound

This paper cites Assessing transferability from simulation to reality for reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Assessing transferability from simulation to reality for reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.350240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.546420Z digest=sha256:cc1c4ee7060512becd10a2fa23f7da7b78b36abfba897cbad05689c356b49a84

Observation 0d15a0f3-b53f-4456-b67d-291feaba836a · outbound

This paper cites Robust deep reinforcement learning through adversarial loss.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust deep reinforcement learning through adversarial loss

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.050191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.594896Z digest=sha256:3366afa6f43821c0a010fd286ed37537eb6cb216d16824c8afe724aa241c9902

Observation e6659db3-eda5-44b5-b5d2-b16b34babfa5 · outbound

This paper cites Characterizing attacks on deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Characterizing attacks on deep reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.900816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.651218Z digest=sha256:1ef51b081299218b260258dcba1465d6b99e891d1ed193ba7c764192aa2a5d78

Observation 3815608d-00c1-4f3d-b6af-766565bfedc9 · outbound

This paper cites Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.720848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.722490Z digest=sha256:724ddfc0606c82001f31b567a2dd783a1181bb120d6bc3656e9a63dce4015e6c

Observation 94d03389-a69b-4bc9-b3ad-29fe3d273a85 · outbound

This paper cites The security of autonomous driving: Threats, defenses, and future directions.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization The security of autonomous driving: Threats, defenses, and future directions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.587133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.795215Z digest=sha256:3d304c83cd0a239b243e3f05fa9bfe14649f335e383a380c2b88b5e5f9ba017b

Observation 3bbd293c-bd7a-49d5-ba9e-9413aa00170f · outbound

This paper cites Learning to walk in minutes using massively parallel deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Learning to walk in minutes using massively parallel deep reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:08.834445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:08.834445Z digest=sha256:b373d12f3c20721c1abc6cad5e8df70671db5a4a0a1a455d55829c2a4c9d5a02

Observation 3611f0eb-36b1-410b-b54d-49b501b96fe0 · outbound

This paper cites Improving robotic machining accuracy through experimental error investigation and modular compensation.The International Journal of Advanced Manufacturing Technology, 85:3–15, 2016.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Improving robotic machining accuracy through experimental error investigation and modular compensation.The International Journal of Advanced Manufacturing Technology, 85:3–15, 2016

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.445899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.885713Z digest=sha256:a130b5074414f215ad233318f8c74777d194b5f57ec6250212a99df9d2d3130c

Observation a7d3b7ca-75f7-4bcf-8e53-e080d83928f5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:08.934071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:08.934071Z digest=sha256:ed789bf9bdd4bc511facfba30142f36ab6b43dcdc401285bfccddfde79488fef

Observation 44a3e993-a13a-4438-b3c0-8aaf2e46b642 · outbound

This paper cites Towards facilitating empathic conversations in online mental health support: A reinforcement learning approach.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Towards facilitating empathic conversations in online mental health support: A reinforcement learning approach

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.264004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.949626Z digest=sha256:a3f75a704a678032d22eeed819d37d31031f9fec245bab1d12808c2e206b69ee

Observation b967c140-f346-4417-b157-dbb834955c42 · outbound

This paper cites Certifiably robust policy learning against adversarial multi-agent communication.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Certifiably robust policy learning against adversarial multi-agent communication

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.092515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:08.990285Z digest=sha256:84c5fe49501de1a2ee0417037843810944e059dc2e5daef98fd17bde0b021bcd

Observation 33ec714b-5e8c-447e-b95f-3630aa23a6ec · outbound

This paper cites Certifiably robust policy learning against adversarial multi-agent communication.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Certifiably robust policy learning against adversarial multi-agent communication

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.940979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.047922Z digest=sha256:1b058cd6677392443575a2ae14ccbb4b70a4126ad20b80a19a825bc6885e1e36

Observation ea130755-e901-474a-8744-0245ecee41df · outbound

This paper cites Who is the strongest enemy? towards optimal and efficient evasion attacks in deep RL.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Who is the strongest enemy? towards optimal and efficient evasion attacks in deep RL

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.706675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.087983Z digest=sha256:f1c09aaa2fe752c0b16491487cdeb8a28dfefdf13386909de500760c788ce366

Observation f76b6a1e-e35c-444f-9a36-5e76b5e90792 · outbound

This paper cites Reinforcement learning: An introduction.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Reinforcement learning: An introduction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:09.150195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:09.150195Z digest=sha256:fbc29820a0493a7ce56e1b1b340b147a89b9fe20bd502da570b8fe8f2946296d

Observation 2b867fe0-f00b-40ac-a803-ad3a36575da2 · outbound

This paper cites Action robust reinforcement learning and applications in continuous control.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Action robust reinforcement learning and applications in continuous control

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.495463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.207310Z digest=sha256:d884d0717b92805680f43168a1019b807a278c42da19099a5d96c7cb773754d7

Observation 6323ea20-c761-4d13-9fbf-80fb45deebed · outbound

This paper cites Mujoco: A physics engine for model-based control.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Mujoco: A physics engine for model-based control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:09.241906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:09.241906Z digest=sha256:989a2ddf6baf4efbb95c908733547425b12f1bff1ccc0b63e543dea2b5f2632e

Observation 4086eb8d-ac23-4e96-9a29-ee8438a40774 · outbound

This paper cites Policy gradient method for robust reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Policy gradient method for robust reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.323799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.294486Z digest=sha256:aaf1e749b2810fc79231eb98f01a9a903b49aa95c2e55e2a1b011e3377af26b7

Observation e99860d5-e676-412f-b7bc-4ca207a87a21 · outbound

This paper cites CROP: Certify- ing robust policies for reinforcement learning through functional smoothing.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization CROP: Certify- ing robust policies for reinforcement learning through functional smoothing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.035690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.363296Z digest=sha256:51949ee158fa872320d05d955a1bfb5e6fb7699f044067518b4d6d19e7e7d988

Observation ca13e76a-64db-4e37-9bd1-b275932fc9a5 · outbound

This paper cites Robust deep reinforcement learning through bootstrapped opportunistic curriculum.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust deep reinforcement learning through bootstrapped opportunistic curriculum

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.903033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.388391Z digest=sha256:a3db099cb8fd614f7c48180db86d0cef1e8a3b35af39490a22ffbcdc1feae59d

Observation b5b0e446-6c9a-4989-a042-cd08b4221c80 · outbound

This paper cites Taac: Temporally abstract actor-critic for continuous control.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Taac: Temporally abstract actor-critic for continuous control

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.739145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.444717Z digest=sha256:628f82cce2b67ddce1a2e6a51e0924bd0ec19056684d5428bbecc544632590f2

Observation 1664edd7-da4d-4e27-bbde-ce91f91b6f18 · outbound

This paper cites Gradient surgery for multi-task learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Gradient surgery for multi-task learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:09.492261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:09.492261Z digest=sha256:03fc6763a1146db895b30fb5e99888709c0b5f25c855e162e2da9a0b24d5957f

Observation 92a427cb-3d4c-4772-be7c-b26a64e2dedb · outbound

This paper cites Robust reinforcement learning on state observations with learned optimal adversary.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust reinforcement learning on state observations with learned optimal adversary

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.561888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.519045Z digest=sha256:f467f1db32f39c4be45aad98030f04f93756a420a05346a06ac442f0cd5f3e3e

Observation c07a37e2-7e1f-4c66-9f37-2339c35d2183 · outbound

This paper cites Robust deep reinforcement learning against adversarial perturbations on state observations.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust deep reinforcement learning against adversarial perturbations on state observations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.429921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.578035Z digest=sha256:2b08154ea5b718e39c85207e53c18e913adbc68bb78cfa36ed2045f8c220074a

Observation 9d2ed019-a327-4c39-9998-5c210206e090 · outbound

This paper cites Adaptive reward-poisoning attacks against reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Adaptive reward-poisoning attacks against reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.215720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.635734Z digest=sha256:3b8742772859ad24c1c4c28e9cae661a5765e40bc29761948eb6459b35cfb3b2

Observation dac71140-d879-4c66-acb0-e7b5eb92cdcb · outbound

This paper cites Sim-to-real transfer in deep reinforcement learning for robotics: a survey.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Sim-to-real transfer in deep reinforcement learning for robotics: a survey

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.069216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.661163Z digest=sha256:1514019ff7baf35e5113414b1e5da0de9b09d557c9ca0817f2424ec291c9d377

Observation cde022d4-24c3-4039-9dc4-a14562d2e012 · outbound

This paper cites Cadre: A cascade deep reinforcement learning framework for vision-based autonomous urban driving.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Cadre: A cascade deep reinforcement learning framework for vision-based autonomous urban driving

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:09.926141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:25:09.694462Z digest=sha256:9f1f5770d846c78d9a0866fb1ccec93862b09ba7ba6cc5515d124eb0f3a490d9

Pith citing papers

No inbound Pith citation observations are available.