Pith. sign in

Paper Citation Record · LEDGER

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation

As of 19 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2506.16753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16753 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:30:12.341086Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 968649e3-b948-4873-9a9e-8a3a56b1e88e · outbound

This paper cites Then, we approximate the peak of the probability by a constant multiple of Dirac’s delta function asκworstδ(˜s⋆) and distribute the remaining probability equally as 1−κworst.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Then, we approximate the peak of the probability by a constant multiple of Dirac’s delta function asκworstδ(˜s⋆) and distribute the remaining probability equally as 1−κworst

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.707156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.271580Z digest=sha256:d02e5b50c240eb18d70c137edd656c701a4137ad945f6b5987bc00a00766f3d7

Observation 98c71c13-319a-4909-974a-cc15c25d1fbb · outbound

This paper cites (57) 5: end for 6: else 7: Do nothing (pass) 8: end if C.3.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation (57) 5: end for 6: else 7: Do nothing (pass) 8: end if C.3

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.687282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.278946Z digest=sha256:7983875f13857f22f4ae0f3f52e7d87decac39c62084652c0b81e7edaa5503ad

Observation 564ea699-d5ee-44ae-80e1-55fd41310cda · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.527935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.333756Z digest=sha256:a787d26a5d167a9e0810102471e7507520572d16dc6219a4452afbd58e210af9

Observation 4b18e3d9-5a3c-49d2-9e01-0ef81c64418f · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.677160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.282856Z digest=sha256:5b6f003ac68450b37ab06c645d41d550718c2064f091d49cc8178b2e6579174a

Observation 1fbe3810-9a7f-44b4-99f8-38038756669a · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Soft Actor-Critic Algorithms and Applications

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.198150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.198150Z digest=sha256:1bdb096b2ee72d95bebb6f2f72f04cf196d4ddfe05192a4476a2c023b2a82667

Observation af50499b-2f15-47c0-b5ac-33fc22bd01c9 · outbound

This paper cites Average episodic rewards (± standard deviation) for median- seed models of our proposed methods (V ALT-EPS, V ALT-SOFT) and other SAC baselines across four MuJoCo tasks.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Average episodic rewards (± standard deviation) for median- seed models of our proposed methods (V ALT-EPS, V ALT-SOFT) and other SAC baselines across four MuJoCo tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.583199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.315732Z digest=sha256:1669eafba70829ae2d16dafb91983abf60602517123af5e97648627066b3ec5b

Observation 407518af-8b6e-466c-b73f-f9596fcbffa8 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.572250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.319227Z digest=sha256:e33ff08866055769b68d15d99a840fde57d1d2c5e75b77a79dfd6323965ba3d0

Observation fd3dc850-0bf3-448c-b634-1be529023962 · outbound

This paper cites We denote the ablation setting as w/o PE, and regularization is applied in all settings, including the ablation.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation We denote the ablation setting as w/o PE, and regularization is applied in all settings, including the ablation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.549881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.325953Z digest=sha256:e4b1320652e3d92d4522757cea48dcb6df366ffcaa6f4f1283ec390cf5ce1fb1

Observation 3eed6b07-0e34-43c2-9df8-8d9f1575263d · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.746796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.255935Z digest=sha256:4a1809798fe1090cf950eb18ef791984f04b898362c5ed066306cd9511fbf4c8

Observation b9797397-b95e-4e70-9522-feab1d3950f9 · outbound

This paper cites We denote these settings asAdv*, where * indicates the adversary rate.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation We denote these settings asAdv*, where * indicates the adversary rate

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.539060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.329943Z digest=sha256:f6038a2d433f463b71a6bc59fd34858a48c23583a5ceffdc4f6af46683d6808d

Observation 124ec24f-4704-4ef5-8836-31de9203e49d · outbound

This paper cites B., Andrychowicz, M., Zaremba, W., and Abbeel, P.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation B., Andrychowicz, M., Zaremba, W., and Abbeel, P

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.796874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.225167Z digest=sha256:2eddd4a6a629899fae6a47acf2bbe9a158c3231325e6feedfb0585b5c7c211bd

Observation ef5b6cc1-b2c0-4482-bfd0-f62f49c8c9ca · outbound

This paper cites Robust Deep Reinforcement Learning Through Adversarial Attacks and Training : A Survey.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Robust Deep Reinforcement Learning Through Adversarial Attacks and Training : A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.228815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.228815Z digest=sha256:e4ff3c38226b7e45ac8469fadfe54772d3f66b8af6bc3a77dbb0b5ca2b13bde5

Observation 8f57e79b-8927-4729-a277-9d5fa1459075 · outbound

This paper cites L., Esfandiari, Y ., Lee, X.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation L., Esfandiari, Y ., Lee, X

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.787041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.236304Z digest=sha256:865c8837819b8ab6f7498a1fb9fe808037097dd21d8fe2ac732b2f71b72a25dc

Observation 57ed649e-24f4-40ed-849d-59373fce1789 · outbound

This paper cites Robust Reinforcement Learning using Adversarial Populations.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Robust Reinforcement Learning using Adversarial Populations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.244026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.244026Z digest=sha256:3edb11811f15f51ce94661c7ffedbcbe02561fe717bd22940110b70712106b56

Observation 799698de-26a9-413f-8b7b-89774f2ecdc9 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.756719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.251858Z digest=sha256:ca9d73655741bd67c4f664de5395dca92a3dcf2b1af5b6618e61cc6335ab5609

Observation b724273f-62e5-4a8e-bb22-43275a72e863 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.737090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.259790Z digest=sha256:4c1bbeadef14e0138f56cdbe8c357e41fb3d4629994890d55407ca0a3b813330

Observation 04b25261-974c-48a1-a00b-ddc3eb1deb40 · outbound

This paper cites Sincef(Q,st) is a monotonically increasing function for Q, then we can say: f(Q1,st)≤f(Q2 +ϵ,st) =ϵ +f(Q2,st) =∥Q1−Q2∥st,˜at +f(Q2,st) ↔f(Q1,st)−f(Q2,st)≤∥Q1−Q2∥st,˜at.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Sincef(Q,st) is a monotonically increasing function for Q, then we can say: f(Q1,st)≤f(Q2 +ϵ,st) =ϵ +f(Q2,st) =∥Q1−Q2∥st,˜at +f(Q2,st) ↔f(Q1,st)−f(Q2,st)≤∥Q1−Q2∥st,˜at

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.717295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.267806Z digest=sha256:729cc0456d86c1179c4a398d2209d088ba85a4eb8e8096f866e6023055e9c7d9

Observation d05cc4d2-c1d0-43be-a17e-5c27d94b8a90 · outbound

This paper cites While its effect is only slightly better in some tasks, we do not observe any disadvantages to using PER.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation While its effect is only slightly better in some tasks, we do not observe any disadvantages to using PER

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.667191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.286523Z digest=sha256:215bd6dea59caefeb271f928fd0758e446b4ba7efc2d9fecd2a09fc41e0e622c

Observation 1a90eaa9-ff76-45f2-b7c2-9475746be696 · outbound

This paper cites However, we observe that the agent’s learning became critically slow in HalfCheetah due to delays in updating the running statistics.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation However, we observe that the agent’s learning became critically slow in HalfCheetah due to delays in updating the running statistics

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.657131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.290245Z digest=sha256:34fe955087f4ce6521383e88380e304e6815abea0724f0907b8b402b6f15380b

Observation da42e9a6-1e9d-4a19-93ef-1f2aa25b2e64 · outbound

This paper cites EnvironmentEnv.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation EnvironmentEnv

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T19:30:12.647100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.294069Z digest=sha256:94d793487b5237421f00ced301dbf7bf56d86b5562c8eb0745196a826b8d03a0

Observation 837462b4-866c-46ff-ac59-7a08ef9cad52 · outbound

This paper cites The SAC (agent) component retains the same settings as the base SAC, while the PPO (adversary) component follows the adversary settings of ATLA-PPO (Zhang et al., 2021).

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation The SAC (agent) component retains the same settings as the base SAC, while the PPO (adversary) component follows the adversary settings of ATLA-PPO (Zhang et al., 2021)

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.636900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.297574Z digest=sha256:7f86c18b97413ae1d469f7457e8e8c6b2e572dc039c8583dd56b738556b1c35e

Observation 5b2d8083-6b5a-41c6-a14c-916e318863a9 · outbound

This paper cites The solid lines represent the average evaluation scores, and the shaded areas indicate standard deviations across different seeds.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation The solid lines represent the average evaluation scores, and the shaded areas indicate standard deviations across different seeds

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.615158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.304369Z digest=sha256:cc02c5fb1eec71e2e5032d536dd3d17364de6fe329f4cc536899c2d3dc457f26

Observation 0f58721f-51fc-4565-8b18-087b95c39f80 · outbound

This paper cites Compared to V ALT-EPS-SAC, V ALT-SOFT-SAC introduces more hyperparameters due to adversary policy training.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Compared to V ALT-EPS-SAC, V ALT-SOFT-SAC introduces more hyperparameters due to adversary policy training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.604239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.308042Z digest=sha256:b9815d7205e23462b3e61abc7a5d708451da982828b71832d1698b6e42c09144

Observation e8e3b7c4-c509-4e7a-8e9f-61d06d0d37b3 · outbound

This paper cites We perform multiple training runs to tune robust critic parameters (perturbation scale and regression term) around the benchmark’s attack scale.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation We perform multiple training runs to tune robust critic parameters (perturbation scale and regression term) around the benchmark’s attack scale

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.593830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.312017Z digest=sha256:8bc1f85e2e79ca03498b2748904dcb12fa65876646721831c7452c417dc83337

Observation 60900cd9-585e-4287-88b1-c3dd1dbebb09 · outbound

This paper cites We denote these settings as w/o PE and w/o PI.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation We denote these settings as w/o PE and w/o PI

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.561623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.322610Z digest=sha256:93bb446950f3760d25d008dca69d284a7f5e304636e1cb0f8657c8b17ad3fc5f

Observation 64f00d96-f918-42e2-972c-782d38a969c3 · outbound

This paper cites Adv1.0 assumes full adversarial influence (˜at∼π◦νsoft), while Adv0.0 uses the agent policy alone (at∼π).α = 4 const.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Adv1.0 assumes full adversarial influence (˜at∼π◦νsoft), while Adv0.0 uses the agent policy alone (at∼π).α = 4 const

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.517546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.337359Z digest=sha256:b95ad94acee6674d199bb77ac993c491e30ef6a83225ba2c1cbad1c084c741b0

Observation 8005e069-78af-4755-a9ee-f364aea26137 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.505561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.341086Z digest=sha256:77499762585c8457a07b21391b43ed32c60da11cf1fb88c5fe9f013067e5a7d7

Observation f5047eff-1cfa-4559-bcf1-72b3e6b1fe04 · outbound

This paper cites Adversarial Attacks on Neural Network Policies.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Adversarial Attacks on Neural Network Policies

Reference 1972

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.202401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.202401Z digest=sha256:d0c15c089e9e44afe123831e821df8bd913f0f18796926848bcd6f1a466639d1

Observation d93ac52b-5d6b-43b8-9e3a-4b6b75750782 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 1998

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.696834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.275342Z digest=sha256:8f478cc78f28db3e5a95de9c5fb9f86b9601534fbea94375bd67685db6ab0cf3

Observation 7832744d-ee2e-4b48-b3d2-59112542e08d · outbound

This paper cites Whatever Does Not Kill Deep Reinforcement Learning, Makes It Stronger.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Whatever Does Not Kill Deep Reinforcement Learning, Makes It Stronger

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.176065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.176065Z digest=sha256:7bb1117b4df2f250e416e4ab2834b21aee8976980d89c76347af78eb66aec345

Observation e35784f0-25fd-487b-81b7-c6de09dfa5ad · outbound

This paper cites Reinforcement Learning via Fenchel-Rockafellar Duality.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Reinforcement Learning via Fenchel-Rockafellar Duality

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.214036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.214036Z digest=sha256:0fc27f879e495c6d502e45d20983c83f1cd5131d3c0f57c2972b40b6a4ba4359

Observation 0c3720f9-ec5f-4dff-a113-3f2bd8c9dc66 · outbound

This paper cites OpenAI Gym.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation OpenAI Gym

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.185268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.185268Z digest=sha256:c15daa69da857b465e135643e0779aa14e859644097b0b5c6cd5d7378ae50758

Observation 7eb69763-072f-4380-b53b-c2f645dbbef4 · outbound

This paper cites Adversarial examples in the physical world.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Adversarial examples in the physical world

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.210101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.210101Z digest=sha256:debf3ce34d7ecd6f789f7a584d7e6a7f1a50e580394b3f35a8630432d2fa2ebc

Observation 62cd38ce-eda0-4063-b6a3-155a2d2b074c · outbound

This paper cites Delving into adversarial attacks on deep policies.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Delving into adversarial attacks on deep policies

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.206466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.206466Z digest=sha256:c7073c2a45b76ad90155e8cce1c85a51f60e56a5915c32ece58ba36d55e553db

Observation f38e2ea4-f007-4003-b5f1-923ab23c0f7b · outbound

This paper cites On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.193798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.193798Z digest=sha256:3c41438cccb8cc37466f320c4158a37cba2041768e97cab68e27f646e1bf0bd1

Observation dfdeffd3-b782-4701-961b-cca68da1e72d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.232612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.232612Z digest=sha256:4eebabbe05c35cd29a20ed3696c992716469846934b6ca7d76e372827f736661

Observation 9d590071-ed2f-48ad-abb6-a4218fe1bf0e · outbound

This paper cites Practical black-box attacks against 12 Off-Policy Actor-Critic for Observation Robustness: Virtual Alternative Training machine learning.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Practical black-box attacks against 12 Off-Policy Actor-Critic for Observation Robustness: Virtual Alternative Training machine learning

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.807109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.221604Z digest=sha256:dcc3e35e1b88cb783f98de16b2154b4321b86a421bacae65b322f769652f02f3

Observation eb18e4d5-08a8-4f09-94c2-5e69bd3e69b2 · outbound

This paper cites f-Divergence constrained policy improvement.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation f-Divergence constrained policy improvement

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.181202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.181202Z digest=sha256:63b8217e21e4f29b379096c42a869f40da5df0935f900c45995fe39371aae1da

Observation 49fcc47c-4f03-46a8-ad2b-4151366b8ad4 · outbound

This paper cites Soft Actor-Critic for Discrete Action Settings.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Soft Actor-Critic for Discrete Action Settings

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.189554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.189554Z digest=sha256:3ed7296c0b480967bc24e635023e485ae979134cc6e33be6cd7be67f6d8d261a

Observation c144b36e-3e1f-4f92-847c-0be805b0ea9f · outbound

This paper cites Mujoco: A physics engine for model-based control.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Mujoco: A physics engine for model-based control

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.776546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.240172Z digest=sha256:1892f873a3d01cd5612463b3f3173d70ed87091ebe03017ff53cc329b2091e31

Observation 79195780-b7e0-4635-84f8-b0844585f1c0 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.626030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.301038Z digest=sha256:de4bad93e3f6cd3813c1233df97afc1d9dfb1ba515e619283b15d2cb83c1ae53

Observation 782021d3-6eed-4c5e-a70c-c2180bce2b6f · outbound

This paper cites Furthermore, Liu et al.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Furthermore, Liu et al

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.727222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.263778Z digest=sha256:32f9819fae0d79c1af0afb32b428641e28fa026043c4f2e9bd9b6823c27a2cfe

Observation 64e57ba0-bf54-4121-942e-e2ca3ae7dcdb · outbound

This paper cites The limitations of deep learning in adversarial settings.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation The limitations of deep learning in adversarial settings

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.816986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.217835Z digest=sha256:ff497120f8c324b8360d13c7fac09a383f1288e852a63bee9ced9600954c84fa

Observation 7d1abe1c-cd0f-4f89-8225-c58d6d6f2082 · outbound

This paper cites In the following, we temporarily set aside strict notation and represent expressions like ” R a∈A·,da ” as ”P a∈A·” when discussing continuous action space.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation In the following, we temporarily set aside strict notation and represent expressions like ” R a∈A·,da ” as ”P a∈A·” when discussing continuous action space

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.766718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:30:12.248032Z digest=sha256:ab3be6ef1f0194d0b7b59cfd1dda18ad7e6a068213aaa30021a823a2676676cf

Pith citing papers

No inbound Pith citation observations are available.