Pith. sign in

Paper Citation Record · LEDGER

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation

As of 19 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2506.16753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16753 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:30:12.341086Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 968649e3-b948-4873-9a9e-8a3a56b1e88e · outbound

This paper cites Then, we approximate the peak of the probability by a constant multiple of Dirac’s delta function asκworstδ(˜s⋆) and distribute the remaining probability equally as 1−κworst.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Then, we approximate the peak of the probability by a constant multiple of Dirac’s delta function asκworstδ(˜s⋆) and distribute the remaining probability equally as 1−κworst

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.707156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.271580Z digest=sha256:8f3e0fc7b2c8b07c9e4ed46b72e5f5f0b16dbb6cd46f0020886f4ef27a3aa4a3

Observation 98c71c13-319a-4909-974a-cc15c25d1fbb · outbound

This paper cites (57) 5: end for 6: else 7: Do nothing (pass) 8: end if C.3.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation (57) 5: end for 6: else 7: Do nothing (pass) 8: end if C.3

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.687282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.278946Z digest=sha256:9ce28f4d2cc7ee1f2d0a3f6a5859fd1345216ea06520412c4eab2a90743d999f

Observation 564ea699-d5ee-44ae-80e1-55fd41310cda · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.527935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.333756Z digest=sha256:0639c9b3d30dc4e56822664ea30c8aa0160bb89b306f4cf42342467232e47b6c

Observation 4b18e3d9-5a3c-49d2-9e01-0ef81c64418f · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.677160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.282856Z digest=sha256:4eab30ba91401ada198c38deb319c2e312a8f5f98e5a0c93ca167b7ba1093e84

Observation 1fbe3810-9a7f-44b4-99f8-38038756669a · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Soft Actor-Critic Algorithms and Applications

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.198150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.198150Z digest=sha256:1bdb096b2ee72d95bebb6f2f72f04cf196d4ddfe05192a4476a2c023b2a82667

Observation af50499b-2f15-47c0-b5ac-33fc22bd01c9 · outbound

This paper cites Average episodic rewards (± standard deviation) for median- seed models of our proposed methods (V ALT-EPS, V ALT-SOFT) and other SAC baselines across four MuJoCo tasks.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Average episodic rewards (± standard deviation) for median- seed models of our proposed methods (V ALT-EPS, V ALT-SOFT) and other SAC baselines across four MuJoCo tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.583199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.315732Z digest=sha256:2edb50e1db5c44645ec6432d88336499901adead37153731fb150141588a541e

Observation 407518af-8b6e-466c-b73f-f9596fcbffa8 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.572250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.319227Z digest=sha256:a0a9a5b6b20f9772604eaa111e15701df2dba7e1028be47d4f187c8c20c13f76

Observation fd3dc850-0bf3-448c-b634-1be529023962 · outbound

This paper cites We denote the ablation setting as w/o PE, and regularization is applied in all settings, including the ablation.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation We denote the ablation setting as w/o PE, and regularization is applied in all settings, including the ablation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.549881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.325953Z digest=sha256:1c67376ba3b7c2d9b9c5b707cc842133fb81385cff42c01d267bac71ed1bd63a

Observation 3eed6b07-0e34-43c2-9df8-8d9f1575263d · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.746796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.255935Z digest=sha256:849b6958a4ad86d4af6754c6abc1fc1818e11b8d4b6c8718596c8794afa25888

Observation b9797397-b95e-4e70-9522-feab1d3950f9 · outbound

This paper cites We denote these settings asAdv*, where * indicates the adversary rate.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation We denote these settings asAdv*, where * indicates the adversary rate

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.539060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.329943Z digest=sha256:4a2cfebe80ab3eee2ad2d35346e96d231518b383f2f4a2236b2cc334d765d97a

Observation 124ec24f-4704-4ef5-8836-31de9203e49d · outbound

This paper cites B., Andrychowicz, M., Zaremba, W., and Abbeel, P.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation B., Andrychowicz, M., Zaremba, W., and Abbeel, P

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.796874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.225167Z digest=sha256:8152acdc5f89af83abf28879635dc461a7eee56dafb76b2b7ef2746f844cd278

Observation ef5b6cc1-b2c0-4482-bfd0-f62f49c8c9ca · outbound

This paper cites Robust Deep Reinforcement Learning Through Adversarial Attacks and Training : A Survey.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Robust Deep Reinforcement Learning Through Adversarial Attacks and Training : A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.228815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.228815Z digest=sha256:e4ff3c38226b7e45ac8469fadfe54772d3f66b8af6bc3a77dbb0b5ca2b13bde5

Observation 8f57e79b-8927-4729-a277-9d5fa1459075 · outbound

This paper cites L., Esfandiari, Y ., Lee, X.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation L., Esfandiari, Y ., Lee, X

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.787041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.236304Z digest=sha256:b58e57d403212db6324831cec3ae06d77e52239692ca90b5e1d8013ea1451c19

Observation 57ed649e-24f4-40ed-849d-59373fce1789 · outbound

This paper cites Robust Reinforcement Learning using Adversarial Populations.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Robust Reinforcement Learning using Adversarial Populations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.244026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.244026Z digest=sha256:693baa72bf39417421449bcc729d39fe8bd64b8b8ea929e23135f93bae0d48c1

Observation 799698de-26a9-413f-8b7b-89774f2ecdc9 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.756719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.251858Z digest=sha256:33bf3194b5c4e806529f3f789fac6d34a329ee14ee3d3e40894a325dd81119b7

Observation b724273f-62e5-4a8e-bb22-43275a72e863 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.737090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.259790Z digest=sha256:7c02aa1dd42f9e159d263ac46d0738e27a4369ea0b4239e60a9113347ea2b080

Observation 04b25261-974c-48a1-a00b-ddc3eb1deb40 · outbound

This paper cites Sincef(Q,st) is a monotonically increasing function for Q, then we can say: f(Q1,st)≤f(Q2 +ϵ,st) =ϵ +f(Q2,st) =∥Q1−Q2∥st,˜at +f(Q2,st) ↔f(Q1,st)−f(Q2,st)≤∥Q1−Q2∥st,˜at.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Sincef(Q,st) is a monotonically increasing function for Q, then we can say: f(Q1,st)≤f(Q2 +ϵ,st) =ϵ +f(Q2,st) =∥Q1−Q2∥st,˜at +f(Q2,st) ↔f(Q1,st)−f(Q2,st)≤∥Q1−Q2∥st,˜at

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.717295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.267806Z digest=sha256:87e2fa79400e8dd922e44f245cdf17c7b621d55aa68d086b41ec31c9017221b2

Observation d05cc4d2-c1d0-43be-a17e-5c27d94b8a90 · outbound

This paper cites While its effect is only slightly better in some tasks, we do not observe any disadvantages to using PER.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation While its effect is only slightly better in some tasks, we do not observe any disadvantages to using PER

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.667191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.286523Z digest=sha256:488ae3872a8f92bf8a1ad35ed8fcefb441b2958bffce64a9856dde7b8ea6af21

Observation 1a90eaa9-ff76-45f2-b7c2-9475746be696 · outbound

This paper cites However, we observe that the agent’s learning became critically slow in HalfCheetah due to delays in updating the running statistics.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation However, we observe that the agent’s learning became critically slow in HalfCheetah due to delays in updating the running statistics

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.657131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.290245Z digest=sha256:73df7a2fa3d8bd578ad07c610693d0019007425e822271d8c03c45f2efbf308c

Observation da42e9a6-1e9d-4a19-93ef-1f2aa25b2e64 · outbound

This paper cites EnvironmentEnv.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation EnvironmentEnv

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T19:30:12.647100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.294069Z digest=sha256:4741c7b7c49bb69511e7c692aff8f373310c1d23629c36c76ba30266c35db68f

Observation 837462b4-866c-46ff-ac59-7a08ef9cad52 · outbound

This paper cites The SAC (agent) component retains the same settings as the base SAC, while the PPO (adversary) component follows the adversary settings of ATLA-PPO (Zhang et al., 2021).

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation The SAC (agent) component retains the same settings as the base SAC, while the PPO (adversary) component follows the adversary settings of ATLA-PPO (Zhang et al., 2021)

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.636900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.297574Z digest=sha256:20c3531a646b8ba6a84e446aa09614ee3e4da02bafa08d7b556eff9cae9d5a2c

Observation 5b2d8083-6b5a-41c6-a14c-916e318863a9 · outbound

This paper cites The solid lines represent the average evaluation scores, and the shaded areas indicate standard deviations across different seeds.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation The solid lines represent the average evaluation scores, and the shaded areas indicate standard deviations across different seeds

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.615158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.304369Z digest=sha256:6b572d85d9bbcfbaf235d5e041bae3b815281028a5564696152f06fbccf46b24

Observation 0f58721f-51fc-4565-8b18-087b95c39f80 · outbound

This paper cites Compared to V ALT-EPS-SAC, V ALT-SOFT-SAC introduces more hyperparameters due to adversary policy training.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Compared to V ALT-EPS-SAC, V ALT-SOFT-SAC introduces more hyperparameters due to adversary policy training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.604239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.308042Z digest=sha256:0fca837d708dc5e0c21445ab2d80778d6e19065374da8801f7b2e549cf42de2e

Observation e8e3b7c4-c509-4e7a-8e9f-61d06d0d37b3 · outbound

This paper cites We perform multiple training runs to tune robust critic parameters (perturbation scale and regression term) around the benchmark’s attack scale.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation We perform multiple training runs to tune robust critic parameters (perturbation scale and regression term) around the benchmark’s attack scale

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.593830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.312017Z digest=sha256:39be212f38bc7b89dd56d1c2fc347dd863d7d70f4bd84b0fa149e96d616dcec5

Observation 60900cd9-585e-4287-88b1-c3dd1dbebb09 · outbound

This paper cites We denote these settings as w/o PE and w/o PI.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation We denote these settings as w/o PE and w/o PI

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.561623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.322610Z digest=sha256:35e90d385557838f44f9eb6a41487a9031818f43995021ddf2aad5b636e356e9

Observation 64f00d96-f918-42e2-972c-782d38a969c3 · outbound

This paper cites Adv1.0 assumes full adversarial influence (˜at∼π◦νsoft), while Adv0.0 uses the agent policy alone (at∼π).α = 4 const.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Adv1.0 assumes full adversarial influence (˜at∼π◦νsoft), while Adv0.0 uses the agent policy alone (at∼π).α = 4 const

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.517546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.337359Z digest=sha256:9c0b7bd4e6fca398ed9149e49cd22b1a9291a8835507b99a31c709191c7aa05f

Observation 8005e069-78af-4755-a9ee-f364aea26137 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.505561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.341086Z digest=sha256:8c0a0d5ec4fd4492ded006123bd1d27acff0334cce95c2c4f07e8868daf21a63

Observation f5047eff-1cfa-4559-bcf1-72b3e6b1fe04 · outbound

This paper cites Adversarial Attacks on Neural Network Policies.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Adversarial Attacks on Neural Network Policies

Reference 1972

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.202401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.202401Z digest=sha256:d0c15c089e9e44afe123831e821df8bd913f0f18796926848bcd6f1a466639d1

Observation d93ac52b-5d6b-43b8-9e3a-4b6b75750782 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 1998

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.696834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.275342Z digest=sha256:775f3a5c38f96fe84fae11ae504922ee32573209cb978cb18dbeb580419be930

Observation 7832744d-ee2e-4b48-b3d2-59112542e08d · outbound

This paper cites Whatever Does Not Kill Deep Reinforcement Learning, Makes It Stronger.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Whatever Does Not Kill Deep Reinforcement Learning, Makes It Stronger

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.176065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.176065Z digest=sha256:5b0887fec570d7c7ea996d84e855744f73e7572b15ba4e7234ccdc28354a807a

Observation e35784f0-25fd-487b-81b7-c6de09dfa5ad · outbound

This paper cites Reinforcement Learning via Fenchel-Rockafellar Duality.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Reinforcement Learning via Fenchel-Rockafellar Duality

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.214036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.214036Z digest=sha256:0fc27f879e495c6d502e45d20983c83f1cd5131d3c0f57c2972b40b6a4ba4359

Observation 0c3720f9-ec5f-4dff-a113-3f2bd8c9dc66 · outbound

This paper cites OpenAI Gym.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation OpenAI Gym

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.185268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.185268Z digest=sha256:c15daa69da857b465e135643e0779aa14e859644097b0b5c6cd5d7378ae50758

Observation 7eb69763-072f-4380-b53b-c2f645dbbef4 · outbound

This paper cites Adversarial examples in the physical world.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Adversarial examples in the physical world

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.210101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.210101Z digest=sha256:debf3ce34d7ecd6f789f7a584d7e6a7f1a50e580394b3f35a8630432d2fa2ebc

Observation 62cd38ce-eda0-4063-b6a3-155a2d2b074c · outbound

This paper cites Delving into adversarial attacks on deep policies.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Delving into adversarial attacks on deep policies

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.206466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.206466Z digest=sha256:c7073c2a45b76ad90155e8cce1c85a51f60e56a5915c32ece58ba36d55e553db

Observation f38e2ea4-f007-4003-b5f1-923ab23c0f7b · outbound

This paper cites On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.193798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.193798Z digest=sha256:3c41438cccb8cc37466f320c4158a37cba2041768e97cab68e27f646e1bf0bd1

Observation dfdeffd3-b782-4701-961b-cca68da1e72d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.232612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.232612Z digest=sha256:4eebabbe05c35cd29a20ed3696c992716469846934b6ca7d76e372827f736661

Observation 9d590071-ed2f-48ad-abb6-a4218fe1bf0e · outbound

This paper cites Practical black-box attacks against 12 Off-Policy Actor-Critic for Observation Robustness: Virtual Alternative Training machine learning.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Practical black-box attacks against 12 Off-Policy Actor-Critic for Observation Robustness: Virtual Alternative Training machine learning

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.807109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.221604Z digest=sha256:d95ad2b74bdd9ffe28c7a9536116877552b0fec4994ea33b0e97724690e01cdc

Observation eb18e4d5-08a8-4f09-94c2-5e69bd3e69b2 · outbound

This paper cites f-Divergence constrained policy improvement.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation f-Divergence constrained policy improvement

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.181202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.181202Z digest=sha256:63b8217e21e4f29b379096c42a869f40da5df0935f900c45995fe39371aae1da

Observation 49fcc47c-4f03-46a8-ad2b-4151366b8ad4 · outbound

This paper cites Soft Actor-Critic for Discrete Action Settings.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Soft Actor-Critic for Discrete Action Settings

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:12.189554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:30:12.189554Z digest=sha256:3ed7296c0b480967bc24e635023e485ae979134cc6e33be6cd7be67f6d8d261a

Observation c144b36e-3e1f-4f92-847c-0be805b0ea9f · outbound

This paper cites Mujoco: A physics engine for model-based control.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Mujoco: A physics engine for model-based control

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.776546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.240172Z digest=sha256:8505dba46859203daab720102faa4336f3b5264b50108144ac733580b64353d2

Observation 79195780-b7e0-4635-84f8-b0844585f1c0 · outbound

This paper cites an unresolved cited work.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:30:12.626030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.301038Z digest=sha256:f4975054e03ad25a2638846f23602315e6e2cf5c9e7db77b47ee20f0514f37f5

Observation 782021d3-6eed-4c5e-a70c-c2180bce2b6f · outbound

This paper cites Furthermore, Liu et al.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation Furthermore, Liu et al

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.727222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.263778Z digest=sha256:b212a62ec26a80070c471cf8ccf3b474dbedac13e4f3576a4df46dc3d409018c

Observation 64e57ba0-bf54-4121-942e-e2ca3ae7dcdb · outbound

This paper cites The limitations of deep learning in adversarial settings.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation The limitations of deep learning in adversarial settings

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.816986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.217835Z digest=sha256:5c114fe775bf95633f39b3cee95f9a798a70c27473da59795c6fdc0d5595163d

Observation 7d1abe1c-cd0f-4f89-8225-c58d6d6f2082 · outbound

This paper cites In the following, we temporarily set aside strict notation and represent expressions like ” R a∈A·,da ” as ”P a∈A·” when discussing continuous action space.

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation In the following, we temporarily set aside strict notation and represent expressions like ” R a∈A·,da ” as ”P a∈A·” when discussing continuous action space

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:30:12.766718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:30:12.248032Z digest=sha256:06e956a347e89f18e8fdcca7c6ef8011968edaf478514744c250227adefcfae4

Pith citing papers

No inbound Pith citation observations are available.