Pith. sign in

Paper Citation Record · LEDGER

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2507.03372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03372 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:25:09.694462Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98c75044-eebe-4946-94c5-7ce03341af9d · outbound

This paper cites Reinforcement learning based recommender systems: A survey.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Reinforcement learning based recommender systems: A survey

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:15.221666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:06.596567Z digest=sha256:6023ebad37ab4abce27f3e3459e187a178aafac07779c2fa976016fcd848664a

Observation 9c0cddd7-2873-4880-b392-cd6d4553824d · outbound

This paper cites Safe learning in robotics: From learning-based control to safe reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Safe learning in robotics: From learning-based control to safe reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:15.105289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:06.686440Z digest=sha256:39b4adfdc5fc161213c1db80f7e865b916a57c998265281f915c970157eadc7c

Observation 6ea39787-c015-42f7-97d4-87ef7aa49182 · outbound

This paper cites Certifiable robustness to adversarial state uncertainty in deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Certifiable robustness to adversarial state uncertainty in deep reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.960200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:06.828541Z digest=sha256:f12f561a600cfb43ec8d021e846603b886d4aa77403253d9f0408f2525b986f6

Observation caff1573-0394-4bcb-b9fa-a300889a184c · outbound

This paper cites Maximum entropy RL (provably) solves some robust RL problems.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Maximum entropy RL (provably) solves some robust RL problems

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.759489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:06.954194Z digest=sha256:0303a46902fbd73a302da9945cc5afb5309825c5dacbeba3cabad32668cf9068

Observation 7d63bf7e-49bc-4d6c-b52e-e2b52d525f58 · outbound

This paper cites Online Robustness Training for Deep Reinforcement Learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Online Robustness Training for Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.041555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.041555Z digest=sha256:cc8f5352f77f0c1a9113011a1a927ade74691fd9486376f24a51b7bf7ffd0dfc

Observation 9943ad94-02b5-4333-bd8c-c41cc21ea8b6 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.152922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.152922Z digest=sha256:5253a3377653756665d36fbd3ad10d4f33df9296c597e9810d907d5ee9e5c277

Observation dbc958f2-f026-4e4a-ba4b-5ff51183c325 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Addressing function approximation error in actor-critic methods

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.297412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.297412Z digest=sha256:a3e8f04810665afae4ee9d9a9533e7693ca6ee64e67a7a6afb65c2f607df0497

Observation 1cba9500-13f1-4507-9147-7826c7103d52 · outbound

This paper cites an unresolved cited work.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:25:14.501113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:07.392711Z digest=sha256:79e175badd344cf3cca3e9a76863553fe002d9014b7f115d82069fde06c680ef

Observation 2dcf1263-1545-4cc2-b7d7-93d9f92998be · outbound

This paper cites Adversarial Attacks on Neural Network Policies.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Adversarial Attacks on Neural Network Policies

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.512173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.512173Z digest=sha256:1181d690d3e002f1f817b80de5685a62241b07f82a57ca2a22d7e3809aecbac3

Observation 9a6075f1-d532-4554-b8a6-80ae8c4bdf95 · outbound

This paper cites The 37 implementation details of proximal policy optimization.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization The 37 implementation details of proximal policy optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.618873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.618873Z digest=sha256:8e4003777731f0d4daf9531bf42fbf9a7c1b6cf9385de77bb2227555def23bbd

Observation 53f44c71-7d3a-441d-b1aa-caa90f6f63e8 · outbound

This paper cites Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.252479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:07.772234Z digest=sha256:26e2c3870a1bb22ba0b96ee6e75388e486d4d888294abf862921ef74b4383a37

Observation e464de02-4f51-4883-a397-cc3af2a83e52 · outbound

This paper cites Challenges and countermeasures for adversarial attacks on deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Challenges and countermeasures for adversarial attacks on deep reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.885692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.885692Z digest=sha256:ff2f1b9a253505af41ab3315f8cf3075907db910810ce691abf921af6a48d9b3

Observation a3cfaeaf-d03f-4b78-af8a-57a8459b7479 · outbound

This paper cites Learning quadrupedal locomotion over challenging terrain.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Learning quadrupedal locomotion over challenging terrain

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:07.996826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:07.996826Z digest=sha256:fcab4e3d2a13c29ca28bf07ad66bda39d916460065097e6b8a4c5b4a5c337bcd

Observation 44e43070-404d-4db8-bdd2-7d4cb313cc98 · outbound

This paper cites Spatiotem- porally constrained action space attacks on deep reinforcement learning agents.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Spatiotem- porally constrained action space attacks on deep reinforcement learning agents

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:14.094582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.066905Z digest=sha256:9fdc6e04955f752d8faf30cdd86a1c851acf525ba031ae4352b407831242bbcc

Observation a10a8786-ad85-484e-b1cc-423d7da93dcb · outbound

This paper cites Efficient adversarial training without attacking: Worst-case-aware robust reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Efficient adversarial training without attacking: Worst-case-aware robust reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.974490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.139880Z digest=sha256:b51280cafb505c1c9403f7adc0242cc57aef7ff42e0dfe4ac0e4ac77cd122db6

Observation 0b0ba7e6-65fc-4a8d-b3a3-8a788f77ac76 · outbound

This paper cites Provably efficient black-box action poisoning attacks against reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Provably efficient black-box action poisoning attacks against reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.807962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.198726Z digest=sha256:c83561e64859cffa9bbec8c9b933746774c5678df12e76c82cf6528238ae8c61

Observation c7fdbf05-08ce-4686-8335-78b600678f61 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Towards deep learning models resistant to adversarial attacks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.649352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.277644Z digest=sha256:f6354341eead9ea898c11f1d0a7aecd59c6f3fafc56005dd94eba0c5f778877b

Observation c0c52c15-98af-43dc-b174-6a9692e6d100 · outbound

This paper cites Human-level control through deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Human-level control through deep reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:08.372071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:08.372071Z digest=sha256:e44c714c6275c120146acff593fcedad850096a7f037b5c0d51990116fbd2283

Observation 47840ea9-ebe9-4d31-a407-a54d4a0d794d · outbound

This paper cites Robust reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.511735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.460059Z digest=sha256:8e8ebed2e3c5b8aaf7155164584bbc2ca2070b4d1e81d5fbed55c4775a00e854

Observation 87f4595a-ac59-4e40-88b9-c9c7c1bd5633 · outbound

This paper cites Assessing transferability from simulation to reality for reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Assessing transferability from simulation to reality for reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.350240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.546420Z digest=sha256:ae5d5762b1f0bd86c4784a6c76af01d0b62fd6dc0b262b7c8da25056cb6c41ea

Observation 0d15a0f3-b53f-4456-b67d-291feaba836a · outbound

This paper cites Robust deep reinforcement learning through adversarial loss.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust deep reinforcement learning through adversarial loss

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:13.050191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.594896Z digest=sha256:c76065dafe023427d525228a4532bdef768835fe4922719bc93838879cbc1d19

Observation e6659db3-eda5-44b5-b5d2-b16b34babfa5 · outbound

This paper cites Characterizing attacks on deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Characterizing attacks on deep reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.900816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.651218Z digest=sha256:e72c3ea8c827a86059955e805139c14907ae138fd411a3685b526e722c451eff

Observation 3815608d-00c1-4f3d-b6af-766565bfedc9 · outbound

This paper cites Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.720848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.722490Z digest=sha256:6a00f099562872e4aa38647f5400d3c2f1a199ceaeb712a8f8dcc3dcdfbf4cf0

Observation 94d03389-a69b-4bc9-b3ad-29fe3d273a85 · outbound

This paper cites The security of autonomous driving: Threats, defenses, and future directions.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization The security of autonomous driving: Threats, defenses, and future directions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.587133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.795215Z digest=sha256:c18aabea008ce4e73528cb37ccd54d330bb018192f09a18806d9dabc72e5cc7c

Observation 3bbd293c-bd7a-49d5-ba9e-9413aa00170f · outbound

This paper cites Learning to walk in minutes using massively parallel deep reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Learning to walk in minutes using massively parallel deep reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:08.834445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:08.834445Z digest=sha256:b373d12f3c20721c1abc6cad5e8df70671db5a4a0a1a455d55829c2a4c9d5a02

Observation 3611f0eb-36b1-410b-b54d-49b501b96fe0 · outbound

This paper cites Improving robotic machining accuracy through experimental error investigation and modular compensation.The International Journal of Advanced Manufacturing Technology, 85:3–15, 2016.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Improving robotic machining accuracy through experimental error investigation and modular compensation.The International Journal of Advanced Manufacturing Technology, 85:3–15, 2016

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.445899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.885713Z digest=sha256:5b5c3e9dadf613275ce7db94759fad8100a0280f6958e5a459f4cd2c2fa53986

Observation a7d3b7ca-75f7-4bcf-8e53-e080d83928f5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:08.934071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:08.934071Z digest=sha256:ed789bf9bdd4bc511facfba30142f36ab6b43dcdc401285bfccddfde79488fef

Observation 44a3e993-a13a-4438-b3c0-8aaf2e46b642 · outbound

This paper cites Towards facilitating empathic conversations in online mental health support: A reinforcement learning approach.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Towards facilitating empathic conversations in online mental health support: A reinforcement learning approach

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.264004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.949626Z digest=sha256:22a400fa62b88fbad53473a5991353f8079ba8a951e6feb8d09a4ee699d5200e

Observation b967c140-f346-4417-b157-dbb834955c42 · outbound

This paper cites Certifiably robust policy learning against adversarial multi-agent communication.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Certifiably robust policy learning against adversarial multi-agent communication

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:12.092515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:08.990285Z digest=sha256:88d242b99bce005c9fda21ee8236f66da44d7a229b41067248e0e49d0833e13e

Observation 33ec714b-5e8c-447e-b95f-3630aa23a6ec · outbound

This paper cites Certifiably robust policy learning against adversarial multi-agent communication.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Certifiably robust policy learning against adversarial multi-agent communication

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.940979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.047922Z digest=sha256:22d2c62d806a4948a6562c0ddf95e6b9ebb3e79aec433d86db7be61522ed09ad

Observation ea130755-e901-474a-8744-0245ecee41df · outbound

This paper cites Who is the strongest enemy? towards optimal and efficient evasion attacks in deep RL.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Who is the strongest enemy? towards optimal and efficient evasion attacks in deep RL

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.706675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.087983Z digest=sha256:b403388d0d2e3392051f3ac9e29c076d1e9692494d334f02a1a6888665fe87d7

Observation f76b6a1e-e35c-444f-9a36-5e76b5e90792 · outbound

This paper cites Reinforcement learning: An introduction.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Reinforcement learning: An introduction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:09.150195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:09.150195Z digest=sha256:fbc29820a0493a7ce56e1b1b340b147a89b9fe20bd502da570b8fe8f2946296d

Observation 2b867fe0-f00b-40ac-a803-ad3a36575da2 · outbound

This paper cites Action robust reinforcement learning and applications in continuous control.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Action robust reinforcement learning and applications in continuous control

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.495463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.207310Z digest=sha256:31988a2f6c067c61f11a641a8fee021464576f0c86b52d198a18134e3a407a21

Observation 6323ea20-c761-4d13-9fbf-80fb45deebed · outbound

This paper cites Mujoco: A physics engine for model-based control.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Mujoco: A physics engine for model-based control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:09.241906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:09.241906Z digest=sha256:989a2ddf6baf4efbb95c908733547425b12f1bff1ccc0b63e543dea2b5f2632e

Observation 4086eb8d-ac23-4e96-9a29-ee8438a40774 · outbound

This paper cites Policy gradient method for robust reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Policy gradient method for robust reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.323799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.294486Z digest=sha256:dcc66e6ceaba5112022c26e077c543db2427be4b28d14cd2f0e25315836ecc52

Observation e99860d5-e676-412f-b7bc-4ca207a87a21 · outbound

This paper cites CROP: Certify- ing robust policies for reinforcement learning through functional smoothing.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization CROP: Certify- ing robust policies for reinforcement learning through functional smoothing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:11.035690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.363296Z digest=sha256:1c0317d05cef16cf74eba88606d443c40cfe7402439fa7ef54ba83232618606d

Observation ca13e76a-64db-4e37-9bd1-b275932fc9a5 · outbound

This paper cites Robust deep reinforcement learning through bootstrapped opportunistic curriculum.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust deep reinforcement learning through bootstrapped opportunistic curriculum

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.903033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.388391Z digest=sha256:c7dd01d2421a27a674caaa38b82e552e516f6865ded01fc6d1ba488f4d3c3561

Observation b5b0e446-6c9a-4989-a042-cd08b4221c80 · outbound

This paper cites Taac: Temporally abstract actor-critic for continuous control.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Taac: Temporally abstract actor-critic for continuous control

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.739145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.444717Z digest=sha256:2845fc5e742435933d2cb5bd05a1bc19869d4f2315c8c48c1e72421f06dfe543

Observation 1664edd7-da4d-4e27-bbde-ce91f91b6f18 · outbound

This paper cites Gradient surgery for multi-task learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Gradient surgery for multi-task learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:09.492261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:09.492261Z digest=sha256:03fc6763a1146db895b30fb5e99888709c0b5f25c855e162e2da9a0b24d5957f

Observation 92a427cb-3d4c-4772-be7c-b26a64e2dedb · outbound

This paper cites Robust reinforcement learning on state observations with learned optimal adversary.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust reinforcement learning on state observations with learned optimal adversary

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.561888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.519045Z digest=sha256:894fca2d52fd8a743f08b255ffacd3a67ea71d563b4c2d372c4fd3d380ab62e2

Observation c07a37e2-7e1f-4c66-9f37-2339c35d2183 · outbound

This paper cites Robust deep reinforcement learning against adversarial perturbations on state observations.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Robust deep reinforcement learning against adversarial perturbations on state observations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.429921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.578035Z digest=sha256:cd6d5471cad1d882c6120da088755a495267264447927f2af576f9617ff878e5

Observation 9d2ed019-a327-4c39-9998-5c210206e090 · outbound

This paper cites Adaptive reward-poisoning attacks against reinforcement learning.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Adaptive reward-poisoning attacks against reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.215720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.635734Z digest=sha256:d7f16bbdc22e451aa74b5be8134d85894a3f293485f86d49e2501d92bba55601

Observation dac71140-d879-4c66-acb0-e7b5eb92cdcb · outbound

This paper cites Sim-to-real transfer in deep reinforcement learning for robotics: a survey.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Sim-to-real transfer in deep reinforcement learning for robotics: a survey

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:10.069216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.661163Z digest=sha256:957416d10a3eb8a4b0ffdc05b7fe1c8cca556958ddad8773f40ef52fd4c90991

Observation cde022d4-24c3-4039-9dc4-a14562d2e012 · outbound

This paper cites Cadre: A cascade deep reinforcement learning framework for vision-based autonomous urban driving.

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization Cadre: A cascade deep reinforcement learning framework for vision-based autonomous urban driving

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:09.926141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:25:09.694462Z digest=sha256:cc52718fcbe3a9b23a88e585eb4f937f4b3581382763161bb0774dd52b0ab4b4

Pith citing papers

No inbound Pith citation observations are available.