Pith. sign in

Paper Citation Record · LEDGER

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

As of 10 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2608.04317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04317 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:58:17.376083Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9e9ffb8-750e-4b83-8c2f-bd45f20644c3 · outbound

This paper cites https://www.atomicredteam.io/.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) https://www.atomicredteam.io/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.297825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.162099Z digest=sha256:66dd881d1c4a20f5a353820cd0784d38edaf11abbea55f8100d925f69f4c949d

Observation 55d89643-12c0-4fe2-b587-e12a11e90d67 · outbound

This paper cites https://aicyberchallenge.com/ overview/.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) https://aicyberchallenge.com/ overview/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.288443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.166233Z digest=sha256:11a7213a5f846a9c78b18678c34af0752fcd81db5f23861087af9243694a3199

Observation 01e142a8-f26c-45f9-9bbe-4eb79a5a53e1 · outbound

This paper cites https://attack.mitre.org/.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) https://attack.mitre.org/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.279178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.169561Z digest=sha256:a373a26f3dd6bbf0ece335215a8f7677c3b6496ca516018266f164a1f5ae6681

Observation 79f9f41b-f184-4565-abc2-5c6a8e94bdcd · outbound

This paper cites EnIGMA: Interactive tools substantially assist LM agents in finding security vulnerabilities.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) EnIGMA: Interactive tools substantially assist LM agents in finding security vulnerabilities

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.270066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.173238Z digest=sha256:10c4aac0f3e20db5baa71349d0dd16ba325984e5a6cb12f73ceb27226a3296dc

Observation 067df76b-02b2-4127-a1bc-ba247e54a6f4 · outbound

This paper cites Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.176790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.176790Z digest=sha256:9c13d62f0dd9227118aaff0b2851102ee9844fa61e079a9733b8e24ad47c6db6

Observation 141545ae-3ce0-4a1b-9c95-bb69a16af683 · outbound

This paper cites Ctibench: a benchmark for evaluating llms in cyber threat intelligence.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Ctibench: a benchmark for evaluating llms in cyber threat intelligence

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.255302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.180734Z digest=sha256:efbe8e2d153eb07ce430022e92d3d1b1a74b6d14091722afe1b0ed26963b7e97

Observation 489af948-f565-4c76-9a3a-e2ed8ea9bf0b · outbound

This paper cites Claude Mythos Preview red.anthropic.com — red.anthropic.com.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Claude Mythos Preview red.anthropic.com — red.anthropic.com

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.245978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.185307Z digest=sha256:22a81cd2139869041700acad69a058b5072b94a5002983946fdd94065cddbfa8

Observation a94a278f-1b97-4c27-a052-f8d65eea29d9 · outbound

This paper cites CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.189295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.189295Z digest=sha256:a1377b0487f00f8a04fed2c0faf08d3544054d4732f48969d9198fabe64a74f2

Observation 7cdc2bd7-d8d6-4bb1-8996-da6ea68649a7 · outbound

This paper cites Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.193944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.193944Z digest=sha256:2323f7845cfe3d40dff70d0ce5813792cc06dd0e06488939637338a8e0c06f7b

Observation 66b4479a-f346-4677-b5e0-5a1b4346697f · outbound

This paper cites Large language models are autonomous cyber defenders.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Large language models are autonomous cyber defenders

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.236818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.197620Z digest=sha256:f43de6d7c4ec1fe9a4ea7b2da91d98fc6fe02d3d59b3171f4d152ba5986c439c

Observation eafcbef2-bf1a-42ce-8839-962fd2b7ef32 · outbound

This paper cites Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.200744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.200744Z digest=sha256:7b9b4ef93d47449d0f8d40fd348fc0457170f816c055d1eea10fcec782ceef03

Observation f56c22bb-c346-46b6-a51d-8999397414a7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.203952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.203952Z digest=sha256:3b5b4f58c1810415121c2f1c85caaa0c83899925c5e99c7539ba927fa72d6ce5

Observation c44250ce-64e6-41b6-9e34-c9837de6b25f · outbound

This paper cites PentestGPT: Evaluating and harnessing large language models for automated penetration testing.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) PentestGPT: Evaluating and harnessing large language models for automated penetration testing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.222351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.207606Z digest=sha256:0f013bf50b45b1dd00917cfd8e34cfa60bdd760717c779c3ad95562d0cff2661

Observation 20e8ebcb-43a6-4c2f-9524-793b567730e6 · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.210586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.210586Z digest=sha256:0797afee9a9242db0a5acc6686f5f2dce4e0f47350b1af06acb621a4afec4b26

Observation 8d7de844-f4f5-43d7-aa20-28fb6659842a · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) LLM Agents can Autonomously Hack Websites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.213871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.213871Z digest=sha256:da5da7608082a3c78c86d72d0a4b3e0e505c2174febac5d95347adfa7fee807f

Observation 537a98f3-3e25-4ac9-9f28-20ec16beef8b · outbound

This paper cites Graphplanner: Graph- based agentic routing for LLMs.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Graphplanner: Graph- based agentic routing for LLMs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.213523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.217256Z digest=sha256:5dd870ed232dc1a9376d31ba05d80d66e57d248ad07d3f99570e5ac1f540104f

Observation 8d073036-155f-4786-b95f-11dcadf22c59 · outbound

This paper cites Redcode: Risky code execution and generation benchmark for code agents.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Redcode: Risky code execution and generation benchmark for code agents

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.203990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.220501Z digest=sha256:b7ab8b107f7130c8e51f64d47638709e7326b00ae722eaa6f52f0012d23ef57a

Observation 39a0136a-f926-47b7-a4fa-5a049b9ef4b5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.223565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.223565Z digest=sha256:829cc1f6e274f0dec6c7ab178fb292498e27ef06db008a1c6bb32cbda5182732

Observation 392eb2c5-267a-48d7-8cfd-b33f5b60f0b6 · outbound

This paper cites Getting pwn’d by ai: Penetration testing with large language models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Getting pwn’d by ai: Penetration testing with large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.193871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.227059Z digest=sha256:3a24473b34f1b2466512cf061c6a54290c1ee6e74a581023ce78bd43e7eae1c3

Observation c775929a-b989-4410-9ac8-1cc62c818250 · outbound

This paper cites Llms as hackers: Autonomous linux privilege escalation attacks.Empirical Software Engineering, 31(3):70, 2026.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Llms as hackers: Autonomous linux privilege escalation attacks.Empirical Software Engineering, 31(3):70, 2026

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.183206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.230065Z digest=sha256:421ce957edf85732270ae2ef043d373c42ec7c96541773c9478d4c43dba543ff

Observation a9a08bcc-f95d-4e36-9e3e-b8c4f572151a · outbound

This paper cites Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.173065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.233285Z digest=sha256:377b7b1b9f24d42b3b43eeb6522147d7c98d88dc9af776c9e0a109af39d4b864

Observation c10c72e4-b6db-44d3-b056-381995d46494 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Qwen2.5-Coder Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.236382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.236382Z digest=sha256:c9d825601a93669efd14eb78b39f21f974d160abe7bd18089163e5082c471bf0

Observation 84a5e156-aac6-4e3a-88ed-2c2efa7cef9b · outbound

This paper cites GPT-4o System Card.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.239797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.239797Z digest=sha256:7bc0c57016449afc8c623f2213b2709d4cc85ef5a3f1b1961a5d497fdcb84440

Observation 28054fda-b27f-4aaf-8f4b-81dd9ae96b37 · outbound

This paper cites Agentic ai for cyber defense: Llm-guided hierarchical multi-agent reinforcement learning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Agentic ai for cyber defense: Llm-guided hierarchical multi-agent reinforcement learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.160033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.243121Z digest=sha256:211baec5e4bc222fd76203b390823531f38766d9b849046175d59664a8f8064a

Observation 828afaeb-c042-4d3c-8910-8c58d6057b1c · outbound

This paper cites Search-r1: Training LLMs to reason and leverage search engines with reinforcement learning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Search-r1: Training LLMs to reason and leverage search engines with reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.148367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.246095Z digest=sha256:4f54ea429b64285fa4c9cd0817b236a4a2cd1702614a523643ec4ca59b02aae9

Observation 5157025d-7461-46e7-b441-5d66539fec3a · outbound

This paper cites Exploring the efficacy of multi-agent reinforcement learning for autonomous cyber defence: A cage challenge 4 perspective.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Exploring the efficacy of multi-agent reinforcement learning for autonomous cyber defence: A cage challenge 4 perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.249096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.249096Z digest=sha256:24cf20e2b7528766044787d814af6d0ea35a37dab3e47b80c94cea797ddc6c66

Observation e6e1b05a-f842-4dbb-b1e7-d144c3b9786f · outbound

This paper cites Automated cyber defense with generalizable graph-based reinforcement learning agents.arXiv preprint arXiv:2509.16151, 2025.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Automated cyber defense with generalizable graph-based reinforcement learning agents.arXiv preprint arXiv:2509.16151, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.252347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.252347Z digest=sha256:96505727c730fd0968ef7715fc220f5c4ebeca61e609017b91d44ab20f665dbf

Observation eea105c7-6006-490e-be8b-80e29d897860 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.255777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.255777Z digest=sha256:067647ada7b17ad625056c6725301c4fffce432fdacf50175008bc081bb6ce85

Observation 002fa239-7cad-40a3-9abc-bd5b8d3cfe68 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Gonzalez, Hao Zhang, and Ion Stoica

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.258881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.258881Z digest=sha256:5ecc5741d7b69fe6f0b8359975588a22e2d519ea998d55df2bb0fc41ce44ba0f

Observation e9181f4b-d3ac-47de-9116-2de7e18ab6a0 · outbound

This paper cites In-the-flow agentic system optimization for effective planning and tool use.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) In-the-flow agentic system optimization for effective planning and tool use

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.910451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.261962Z digest=sha256:c7618e1bdb2e0ec265a787500029b5d348f1d2bc23746c8407e087723bf9b571

Observation 122e0f90-d9be-4245-9c27-eee78120d9bd · outbound

This paper cites Code as policies: Language model programs for embodied control.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Code as policies: Language model programs for embodied control

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.264984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.264984Z digest=sha256:1d1a98fc4b70f733cb9197a50db1400ecf1f4f538fb4b70723c1086c6e9e5b55

Observation 74634654-66b4-4099-b9a5-9ac2f6dd52b0 · outbound

This paper cites Et-bert: A contextualized datagram representation with pre-training transformers for encrypted traffic classification.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Et-bert: A contextualized datagram representation with pre-training transformers for encrypted traffic classification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.896149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.267969Z digest=sha256:b6d0b31beb581ea2e46c3e53d5f32d1b8375a83595a70df0045142b96473c207

Observation 602324bb-d5ad-459d-9e75-1357f935d447 · outbound

This paper cites Cyberbench: A multi-task benchmark for evaluating large language models in cybersecurity.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Cyberbench: A multi-task benchmark for evaluating large language models in cybersecurity

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.886939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.271027Z digest=sha256:8bf1d98b9c3018901fe39ca63c919cb5775adc4500acec986b21772a48d83e84

Observation 7addd55e-ce6c-400f-837c-9d12548168d0 · outbound

This paper cites Visual-rft: Visual reinforcement fine-tuning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Visual-rft: Visual reinforcement fine-tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.273922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.273922Z digest=sha256:6be9733b4981e827fe306507f8fd36a5cf9304dda495c84a96b82e0023dd14b2

Observation fe77814f-b49e-4ca1-a9d0-27a9fc5cc743 · outbound

This paper cites Contrasting centralized and decentralized critics in multi-agent reinforcement learning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Contrasting centralized and decentralized critics in multi-agent reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.872569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.276936Z digest=sha256:1774090eb912afae6bc0a73e155569dfe1b69af16ae33a0f5f33b9f36a43e57d

Observation 3113a622-7603-43f8-a95c-d5faabd251b3 · outbound

This paper cites Eureka: Human-level reward design via coding large language models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Eureka: Human-level reward design via coding large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.279999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.279999Z digest=sha256:32b26b943ab0c7b5f16a359e6d2689578a513dcf3d267fadcacdeae67b9704a0

Observation 19bfc18e-0300-400e-9f37-06b0db33acbd · outbound

This paper cites Ray: A distributed framework for emerging {AI} applications.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Ray: A distributed framework for emerging {AI} applications

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.283080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.283080Z digest=sha256:770f29425766a08ff931b70c89d49bc5287128674794b55531ec1ecf2d0ba7a1

Observation 3457e4e5-f661-4b23-8580-db69c17dbeaf · outbound

This paper cites Experience with emerald to date.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Experience with emerald to date

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.853501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.285979Z digest=sha256:66714be0fc58d8cc22a32232c323ea22b2e2f4f03a1bc9d5e7a3c1b31cfe49a5

Observation bbccf353-cf5e-42ab-b437-e87fa106e2e5 · outbound

This paper cites Towards a high fidelity training environment for autonomous cyber defense agents.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Towards a high fidelity training environment for autonomous cyber defense agents

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.843871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.289324Z digest=sha256:3c4cecd96ae2bfc006d57b976e61fabec69774f4314a62ddc3b0a983d9bc8d25

Observation 6de97952-a83d-4464-a26e-8cf90eb500df · outbound

This paper cites Proximal Policy Optimization Algorithms.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Proximal Policy Optimization Algorithms

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.292586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.292586Z digest=sha256:cca604bffea7054e02039e2a6c3a4de8af47f39dae91ce8bf397a86dbcdbf26c

Observation a004ecfb-09e6-44c3-986f-43b6cbf9ea81 · outbound

This paper cites Nyu ctf bench: A scalable open-source benchmark dataset for evaluating llms in offensive security.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Nyu ctf bench: A scalable open-source benchmark dataset for evaluating llms in offensive security

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.834168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.295843Z digest=sha256:bf8596c5b2fe6e02f3ab121e8a2727b2c44617c4ca27243afee2782b0f9e994f

Observation 3a10cc6c-cbfd-46c2-a99b-629db9076462 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.298765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.298765Z digest=sha256:ec26d36c78ec74f8c8b96c306cefd1b8a059c3b6f473b0a60310940e32467cd1

Observation 97c601a8-7810-4d7f-93d1-c46ca867f66b · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Hybridflow: A flexible and efficient rlhf framework

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.301997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.301997Z digest=sha256:ef5611195496c49c9f2b762f02cf47bf1338ee932d7ee1ae08eb9c3a378685dd

Observation b6fbeb53-181b-4e61-9d1c-dbe37dd0d1ce · outbound

This paper cites Hierarchical multi-agent reinforcement learning for cyber network defense.Reinforcement Learning Journal, 6:790–810, 2025.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Hierarchical multi-agent reinforcement learning for cyber network defense.Reinforcement Learning Journal, 6:790–810, 2025

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.819801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.305358Z digest=sha256:74c84b1da27f70b72288e686a6361739fce8c08de980a44e36844aa965e96c58

Observation 0d1bded1-97d5-40eb-8495-315f36c22d2b · outbound

This paper cites A taxonomy of intrusion response systems.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) A taxonomy of intrusion response systems

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.810253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.308716Z digest=sha256:7265b57417506edbf83bb8c3b79171ea33fdaa22786576d2ba2e7be054d9d7c8

Observation 2f99b4cb-176e-44b6-8d9e-9af9daa5eac8 · outbound

This paper cites Redsage: A cybersecurity generalist LLM.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Redsage: A cybersecurity generalist LLM

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.800806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.311692Z digest=sha256:5fc1b59fba27a1eb83cf4247e3b397baef284b4ea113595504d99036352e4e36

Observation 294053bb-943e-470b-9ada-fd18ddb5bad4 · outbound

This paper cites Cyberbattlesim, 2021.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Cyberbattlesim, 2021

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.791420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.314759Z digest=sha256:fb6018b68a5d470ea9b6777d07b4542012dd86f2ef15d502484bc84d44a3cf09

Observation 1782e2f7-4dd3-4e2a-9cab-129cb7519f89 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Qwen2.5: A party of foundation models, September 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.317755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.317755Z digest=sha256:3125504223be5d0982315ac030ac92a307cc6e17584961f54b4649df2f79d3cf

Observation 52f80ab4-0d91-46e9-be5e-44d05215a9f5 · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.320654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.320654Z digest=sha256:b269e1f5a76ed91df2a18201882d55602d529efeca12b58c1b3427a9f98dcebd

Observation c66d7f38-f65f-4667-bcd0-2e9622ff5b4e · outbound

This paper cites SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.324049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.324049Z digest=sha256:c27da8e4ef0b71d2c98cc1ee9219aaa9fc56043e57e8d9e42d2187740deebd32

Observation 2c23d796-9356-45ec-a517-60c7e6bfa970 · outbound

This paper cites SymRTLO: Enhancing RTL code optimization with LLMs and neuron-inspired symbolic reasoning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) SymRTLO: Enhancing RTL code optimization with LLMs and neuron-inspired symbolic reasoning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.776191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.327331Z digest=sha256:5015abf0fb54cf0a30767c10d43a911056a498d1e803d5cf83bf96dd32a6e657

Observation 5dfeadbe-4bab-4196-a7e0-b38991215f8c · outbound

This paper cites Cyber- gym: Evaluating AI agents’ real-world cybersecurity capabilities at scale.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Cyber- gym: Evaluating AI agents’ real-world cybersecurity capabilities at scale

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.330596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.330596Z digest=sha256:fded9ca68674387e6ccbc1cba1f9205011633262866500951635d75b947a4a42

Observation 7d765deb-4598-47d6-9abe-0b9c9fe38835 · outbound

This paper cites SWE-RL: Advancing LLM reasoning via reinforcement learning on open software evolution.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) SWE-RL: Advancing LLM reasoning via reinforcement learning on open software evolution

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.762294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.333659Z digest=sha256:bb02f1a087fbe55d0ef34fa93b54b0552b3f8875d5be357c873c7f195d7117ea

Observation 7967ed37-304e-4d30-8fef-a0d5f8595db7 · outbound

This paper cites Autogen: Enabling next-gen LLM applications via multi-agent conversations.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Autogen: Enabling next-gen LLM applications via multi-agent conversations

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.752947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.336740Z digest=sha256:34092bd36e4c1b38ce9e62797ddeae39d05aa2df368fe6c14e97e7946f5bafe8

Observation eee7a847-1c5c-464e-8a27-af119d650e66 · outbound

This paper cites Qwen2 Technical Report.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Qwen2 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.339747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.339747Z digest=sha256:99010ecbd9e064562249f59c20201e67e87ea0f919e736ac601549d8a5431c83

Observation c3698d0f-5488-44a5-81a1-c9642f22a97c · outbound

This paper cites Intercode: Standardizing and benchmarking interactive coding with execution feedback.Advances in Neural Information Processing Systems, 36:23826–23854, 2023.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Intercode: Standardizing and benchmarking interactive coding with execution feedback.Advances in Neural Information Processing Systems, 36:23826–23854, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.343099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.343099Z digest=sha256:77c875b34d8aa258e4cb5ac360fe1da82aeb4ce8ac9050e6ce86edb83835331f

Observation f1dc883e-3319-46a9-bdee-54080f81ded9 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) React: Synergizing reasoning and acting in language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.346037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.346037Z digest=sha256:84e880c695b1afd95bfc8987e8b28387470df85330da541e75a18491276d3a86

Observation 1e522d46-fdde-43c2-bcf0-37c6e7d88933 · outbound

This paper cites Primus: A pioneering collection of open-source datasets for cybersecurity LLM training.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Primus: A pioneering collection of open-source datasets for cybersecurity LLM training

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.734052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.348903Z digest=sha256:6a2d0c2ba1ba398f972c508704a02298ea8ea7efb49d652929078cc8803a6ea1

Observation 1caec6f9-8983-4b60-8d5d-737a854fc042 · outbound

This paper cites ACECODER: Acing coder RL via automated test-case synthesis.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) ACECODER: Acing coder RL via automated test-case synthesis

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.724919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.351778Z digest=sha256:8ddcae35a5ebb1184024a87b3064db48ac0f8c9187acce3afa2e063681563d43

Observation 367b9b5e-445a-46ef-b1b6-e52272cd4fa2 · outbound

This paper cites Ho, and Percy Liang.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Ho, and Percy Liang

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.354780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.354780Z digest=sha256:19dc164e68c1e17e72d0d0fbaa6471cbb4e558ab703252789f2ba7a1931c9bd6

Observation 13fc7108-c6b0-42d2-9814-7d4992243321 · outbound

This paper cites Abdi, William Blum, and Muhammad Abdul-Mageed.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Abdi, William Blum, and Muhammad Abdul-Mageed

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.710547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.357824Z digest=sha256:e5005ebeaafa6d4328200b97a9dc649fcab4fbef389312a8340c089825b7aa4f

Observation e13e38df-282a-43f7-a6e7-425f58fe3047 · outbound

This paper cites Yet another traffic classifier: A masked autoencoder based traffic transformer with multi- level flow representation.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Yet another traffic classifier: A masked autoencoder based traffic transformer with multi- level flow representation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.700836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.360906Z digest=sha256:ae7183df42c7ae6b2ad425e47692a17f70b84f04433768ab19dc3dfd703a3486

Observation 0ad2c876-7da4-447a-9c50-02f548644f98 · outbound

This paper cites Curran Associates Inc., Red Hook, NY , USA, 2019.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Curran Associates Inc., Red Hook, NY , USA, 2019

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.690889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.363877Z digest=sha256:ca85f487b48e4b049912358ecf82322b0079466555319ec20330baa81c46f505

Observation 7c3dd15d-4a6b-4b36-974e-70e9254d75d6 · outbound

This paper cites More than just functional: LLM-as-a-critique for efficient code generation.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) More than just functional: LLM-as-a-critique for efficient code generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.681374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.366857Z digest=sha256:a931e4986a45038ebd3732c9891b43c2b5a7e20d4a61ee7cfd8237a81603dac7

Observation 9750a2ff-fd21-4e2c-8c61-455322623ec2 · outbound

This paper cites CVE-bench: A benchmark for AI agents’ ability to exploit real-world web application vulnerabilities.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) CVE-bench: A benchmark for AI agents’ ability to exploit real-world web application vulnerabilities

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.671444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.369963Z digest=sha256:c906d7b5c3c4053f8647b1655b72126309ce94dd95dcbca17e5c5a3dd76ad5e8

Observation bf259fbf-ae1a-4a4e-8ced-017136dbe910 · outbound

This paper cites Teams of LLM agents can exploit zero-day vulnerabilities.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Teams of LLM agents can exploit zero-day vulnerabilities

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.661621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.372961Z digest=sha256:dee113f47aa8e438dab9a1b2cd2951afa3ea66452563e1645c272faef4114d7e

Observation c46cf609-583b-4ba5-b8a1-54ebd9e9fbf7 · outbound

This paper cites Cyber-zero: Training cybersecurity agents without runtime.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Cyber-zero: Training cybersecurity agents without runtime

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T19:58:17.476910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:58:17.376083Z digest=sha256:882eb706e02c990e20ab2da751248e88056d4120eef93cadc920567651fe0a4b

Pith citing papers

No inbound Pith citation observations are available.