Pith. sign in

Paper Citation Record · LEDGER

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety

As of 9 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2508.03864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.03864 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:17:13.027364Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T18:58:53.183734Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T19:00:30.421345Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy39
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd9bd894-7267-43b7-8062-5e0173a34f7c · outbound

This paper cites Improving retrieval-augmented generation through multi-agent reinforcement learning, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Improving retrieval-augmented generation through multi-agent reinforcement learning, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:21.797033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:08.034162Z digest=sha256:1685a2e13f8c69915d767b7025429f7a8f566ea34bbfb949ff9b82a4648a0e38

Observation 563fde29-6a1e-49ba-9089-0e7ebe47144d · outbound

This paper cites Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:21.614175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:08.175865Z digest=sha256:7f1018b960dc16a1d1e59e899682b64cc7d613edb13c8512d0c693f3b2bcf80d

Observation c6b559a9-3b6a-49c1-b5e5-5dda9b771f89 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:08.292580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:08.292580Z digest=sha256:9c79899a33b2f77c7283aa269aea1a9bec2aaded22581651de3c238333f7135c

Observation 3baf1983-0e9d-49ef-9693-4d51b3a54e74 · outbound

This paper cites Multilingual jailbreak challenges in large language models, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Multilingual jailbreak challenges in large language models, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:21.390561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:08.401691Z digest=sha256:750910bbde4de065e1c5de7754ee4ec2b66ddf36bb755794ca0e4fffa64b2bc0

Observation be55f796-742a-447e-b510-eb7230a1e39a · outbound

This paper cites A practical memory injection attack against llm agents, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety A practical memory injection attack against llm agents, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:21.121642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:08.541064Z digest=sha256:0fd6a41720f967cc56f2c98278091f3bcb7e7a700b0cd5a7df9fbdcc1d622967

Observation 84c03e86-b527-479e-8d22-668b94824e67 · outbound

This paper cites Peerguard: Defending multi-agent systems against backdoor attacks through mutual reasoning,.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Peerguard: Defending multi-agent systems against backdoor attacks through mutual reasoning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:20.777722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:08.674109Z digest=sha256:8c53c62cd00b0807b9349a8d68efe4243ca6fd1c60f72b7f06aa2188204094b6

Observation 00c19c84-6117-4bf4-8b22-ab6470007ec3 · outbound

This paper cites Not what you’ve signed up for: Compromising real-world llm-integrated ap- plications with indirect prompt injection, 2023.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Not what you’ve signed up for: Compromising real-world llm-integrated ap- plications with indirect prompt injection, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:20.414779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:08.750735Z digest=sha256:d4934ebf034f9fe33c62b5026e7715294506f22667aa3ca3c6068c3d10bf66fa

Observation a8e70d4c-72fb-4780-a4f6-80e62addbec8 · outbound

This paper cites Llm multi-agent systems: Challenges and open problems, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Llm multi-agent systems: Challenges and open problems, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:20.053410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:08.835851Z digest=sha256:2065e33b64d1afad50149eb0b259d835bcad920ccd4fb3c62c107ed96d0bdd69

Observation 438d1589-163d-4f7c-865f-9bedbfe11013 · outbound

This paper cites Red-teaming llm multi-agent systems via commu- nication attacks, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Red-teaming llm multi-agent systems via commu- nication attacks, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:19.752036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:08.902413Z digest=sha256:04027f878b167468b66bd3762c402158632c5c104d409de4e7cb8da004255fc3

Observation c67e3a34-d65b-43b3-b3c5-4a5afc1c68e9 · outbound

This paper cites Measuring mathematical problem solving with the math dataset, 2021.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Measuring mathematical problem solving with the math dataset, 2021

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:19.453212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:09.019799Z digest=sha256:91488002c52aa0afc2cef4aad1524e79dc29a1950ef4741eb21fa24e9c10cf2e

Observation 2a341dd6-e818-45f6-8273-9fa6e85da5a1 · outbound

This paper cites Llama guard: Llm-based input-output safeguard for human-ai con- versations, 2023.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Llama guard: Llm-based input-output safeguard for human-ai con- versations, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:19.144675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:09.190821Z digest=sha256:b73083f04d5873a7b3f0876b9d1a1508072521e6d817402e87ae5d943e6f955d

Observation c633c3ff-51e4-466e-ba03-ead9e4384812 · outbound

This paper cites Search- r1: Training llms to reason and leverage search engines with reinforcement learning, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Search- r1: Training llms to reason and leverage search engines with reinforcement learning, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:18.829033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:09.279321Z digest=sha256:f5c713aeb2d7d8812494c88beb049e140fe04072e6583af12c17f028f4ccbbd5

Observation b8bf145c-aa84-451a-86ca-f7c230af4dd8 · outbound

This paper cites an unresolved cited work.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:17:18.524994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:09.394713Z digest=sha256:97684dc18513da75292b16b8e6023f9ab484d6ef44e9c9946e640633b9a0db66

Observation d2312e4f-b055-4103-a084-17f8bd5ffb13 · outbound

This paper cites Trust re- gion policy optimisation in multi-agent reinforcement learn- ing, 2022.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Trust re- gion policy optimisation in multi-agent reinforcement learn- ing, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:18.351198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:09.477475Z digest=sha256:af9ca3a99bd422677f07a27ffd6fd92d6959683e72c24d0e2e8acc61df0aeebe

Observation 3abee350-f220-43c3-a556-2c46afef842e · outbound

This paper cites Prompt infection: Llm-to- llm prompt injection within multi-agent systems, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Prompt infection: Llm-to- llm prompt injection within multi-agent systems, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:18.131372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:09.615213Z digest=sha256:0b4d7af2573efb47255975fbdb06461a3266e73d4c920f5abf7e8f111bee8944

Observation 0d027fcb-cd73-457d-815b-59f282d5bedf · outbound

This paper cites Deepinception: Hypnotize large language model to be jailbreaker, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Deepinception: Hypnotize large language model to be jailbreaker, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.931362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:09.701866Z digest=sha256:3d7eca19a65d3b10ed36dd5eceb42ebbf402d08ed1e97ac53e4ab71cc354c649

Observation 3aaee5ab-09c9-41fd-a037-a51cb2150694 · outbound

This paper cites Tf-attack: Transferable and fast adversarial attacks on large language models, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Tf-attack: Transferable and fast adversarial attacks on large language models, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.769154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:09.794018Z digest=sha256:ec7f706a135411ac79631d12f3af2586aed66660f2750459b97cc1420385f649

Observation 9ea0f7f3-6939-47e1-af4d-b57dbbb0846e · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.613247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:09.885976Z digest=sha256:f2a878d4f75392dc7507aee9261e7f8db3ecfbcdc5e8739a6a5f733e733b4ebc

Observation dc5de2ff-03af-4183-821d-9d360b7cd6cc · outbound

This paper cites Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jail- break attacks, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jail- break attacks, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.448546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:09.968243Z digest=sha256:3879e84fc04c5b21bbc4870c17f6b80d4baeddda0a09234698cc0bed6ebaa9b3

Observation 6b73d23f-c53a-4e18-9395-705fc64b6b13 · outbound

This paper cites Codechameleon: Personalized encryption frame- work for jailbreaking large language models, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Codechameleon: Personalized encryption frame- work for jailbreaking large language models, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.286041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:10.089362Z digest=sha256:e2cce95f903907444685df62c8341373e674741f5ada402e2b8826104cc2e300

Observation fc82ba90-e4b7-4cd4-a863-28a7e9356d1e · outbound

This paper cites Harm- bench: A standardized evaluation framework for automated red teaming and robust refusal, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Harm- bench: A standardized evaluation framework for automated red teaming and robust refusal, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.210645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:10.209297Z digest=sha256:2e2bea5e0b14cb42c0a81286ea8cf69cfa8c5906651220fa818552d6c1a845ff

Observation 172ba3d1-7225-4121-877c-619737d097b6 · outbound

This paper cites Metaspatial: Reinforcing 3d spa- tial reasoning in vlms for the metaverse.arXiv preprint arXiv:2503.18470, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Metaspatial: Reinforcing 3d spa- tial reasoning in vlms for the metaverse.arXiv preprint arXiv:2503.18470, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:10.290349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:10.290349Z digest=sha256:dd1810c8ea498619081596a790449ea968a0d926a82f167515187381408c51b6

Observation dcfa71a7-5459-452b-b77c-4839d76b6df8 · outbound

This paper cites Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:10.380062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:10.380062Z digest=sha256:3fc5669868506a126c5b24cb43a6e1290312bd037a2c0f1d8b9554e1d2a25a0d

Observation a9cd9caf-e8d6-4c21-8bd6-cf11e8a8ad7b · outbound

This paper cites Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:10.475488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:10.475488Z digest=sha256:4737d586ae562bfb1708b8b0225e1bc3fcd85cecfd410bd02379392e4d203c22

Observation 7d953f70-da50-44d8-913c-25ea686208ab · outbound

This paper cites Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-Action.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-Action

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:10.574459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:10.574459Z digest=sha256:b8ec0a6b5171f8c0cdca99a9f94b0bfaada1ba69968686530904c6827c12d686

Observation 26f531b3-0c99-4667-9d91-3e65103ea2e0 · outbound

This paper cites Do code llms understand design patterns? In2025 IEEE/ACM Inter- national Workshop on Large Language Models for Code (LLM4Code), pages 209–212.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Do code llms understand design patterns? In2025 IEEE/ACM Inter- national Workshop on Large Language Models for Code (LLM4Code), pages 209–212

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.070627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:10.694884Z digest=sha256:fa49bb00d92538ed9d5ad886a42a0a9426e90b0461e287f548a54964abffa416

Observation dd25c245-3dfc-493f-ac18-7da47383aca2 · outbound

This paper cites Yu, Manling Li, and Han Liu.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Yu, Manling Li, and Han Liu

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.906588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:10.828016Z digest=sha256:43d817fb50a10552ecbfa01bfd13049c2c9da4c04cc2adf52b1a8d67c8d42de3

Observation afb8e609-e62e-4f80-8620-b874fa7db811 · outbound

This paper cites Maporl: Multi-agent post-co-training for collaborative large language models with reinforcement learning, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Maporl: Multi-agent post-co-training for collaborative large language models with reinforcement learning, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.737349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:10.912276Z digest=sha256:89c06d70e9a9661d77872427cdb3b2ace6ca3ce0f6620d90af8a4b8993e3903f

Observation 0fbbc2b8-4635-4abb-9151-e7ad4b34d2a3 · outbound

This paper cites Red teaming the mind of the machine: A systematic evaluation of prompt injection and jailbreak vul- nerabilities in llms, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Red teaming the mind of the machine: A systematic evaluation of prompt injection and jailbreak vul- nerabilities in llms, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.581769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:11.050008Z digest=sha256:4a3c10311ca5b238e1bb6c4c83933bcb912e701dea0fed7784677d5bd067ad9c

Observation 30435057-2a00-48ea-bb19-38d223ce91a7 · outbound

This paper cites Qwen2.5 technical report, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Qwen2.5 technical report, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.462728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:11.156744Z digest=sha256:2d92e940ddb1ac7e9ed77bd13f80bc5e1327f94872cec90d48c1699fdd6a20d2

Observation f5f5cdb3-ce4f-4001-aad7-796e434fe262 · outbound

This paper cites Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning, 2018.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning, 2018

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.270429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:11.291722Z digest=sha256:7258228907de6feab1694d0aa8cae9a307bf215e51d6212ddbebcb1d4e560973

Observation cd0c0ce3-4ff0-46e8-9a8c-a4a2626487e5 · outbound

This paper cites Proximal policy optimization algo- rithms, 2017.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Proximal policy optimization algo- rithms, 2017

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.069560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:11.390773Z digest=sha256:a6a4f16786606f60c63f4949668e7d1c51190564927f9257d4ac88f8c7c986b7

Observation 7313c8ca-3796-41ce-866b-0df1da0a3aac · outbound

This paper cites Exfiltration of personal information from chatgpt via prompt injection, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Exfiltration of personal information from chatgpt via prompt injection, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:15.898471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:11.468085Z digest=sha256:16586cbb2d4d56e97b57d0628debe224f82c8b388bccaa897c21b8ff63a55887

Observation e234953e-7426-45a9-ab15-f4820562e85d · outbound

This paper cites Sciscigpt: Advancing human- ai collaboration in the science of science.arXiv preprint arXiv:2504.05559, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Sciscigpt: Advancing human- ai collaboration in the science of science.arXiv preprint arXiv:2504.05559, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:11.579252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:11.579252Z digest=sha256:1c53979f00bc4aef9d0561ae62fbc64bb5e457ae8cf58e710cc33e6246235f6e

Observation ee7eff76-68f7-44ac-b859-43b5a7ce914d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:11.673708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:11.673708Z digest=sha256:c9cbd183af70af9c2b17be083fd4f701992fe39fc55a8216f29f8b7726643e6a

Observation 5bc28a52-9f95-45d8-be1a-dfaabc07693a · outbound

This paper cites Prompt injection attack to tool selection in llm agents, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Prompt injection attack to tool selection in llm agents, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:15.712033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:11.775209Z digest=sha256:6f724ea7dba4a991a4800b38f0ce4a1c06b2d716c2007b40752ce25259cb7d5e

Observation 8ce99069-7b5a-4ac8-afea-8b737881ab91 · outbound

This paper cites an unresolved cited work.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:17:15.484083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:11.868960Z digest=sha256:d3bf8da7a6f331b771dba672b63d9bb17a23866b443b6d3d5749f15b9a706eee

Observation c2759c17-5d28-48fd-a776-64d70bbd1d72 · outbound

This paper cites Multi- agent systems execute arbitrary malicious code, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Multi- agent systems execute arbitrary malicious code, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:15.296371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:11.980418Z digest=sha256:131446b44f26efdeaf2d211d29290f74ae9a2263e76dd6e93ccc52244e2b7b36

Observation a735f8f2-8ca2-4fb3-b43f-be9f120ba091 · outbound

This paper cites Lyu, and Maarten Sap.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Lyu, and Maarten Sap

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:15.122813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:12.091016Z digest=sha256:2444e114b0900f5bef331638f92b90a941d92aa7f66239efb32e5daa2e75c62d

Observation 8534057d-95d5-4384-b11b-5f76d7b748fb · outbound

This paper cites Rema: Learning to meta-think for llms with multi-agent reinforcement learning, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Rema: Learning to meta-think for llms with multi-agent reinforcement learning, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:14.972763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:12.199699Z digest=sha256:8e31341cf8cce25a1a1df655cdb473130550907798e956f7c9026b1d621e734e

Observation 5f09d10d-b5ea-470b-9983-6738b143438c · outbound

This paper cites G- safeguard: A topology-guided security lens and treatment on llm-based multi-agent systems, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety G- safeguard: A topology-guided security lens and treatment on llm-based multi-agent systems, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:14.794982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:12.318504Z digest=sha256:c8516d575f310f57c778e6e61ae6010b28ffe84366dc7b715baabe671173e37a

Observation f5b6ad89-c513-4b99-82eb-a695d9e10541 · outbound

This paper cites Unleashing the emergent cogni- tive synergy in large language models: A task-solving agent through multi-persona self-collaboration, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Unleashing the emergent cogni- tive synergy in large language models: A task-solving agent through multi-persona self-collaboration, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:14.643499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:12.409382Z digest=sha256:f019d067a7520415dd387d97ba441e89122c6556251d1cf5dc91a04710ae7d62

Observation 5db696b9-5838-47c4-b9da-cf88ac3b0ded · outbound

This paper cites an unresolved cited work.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:17:14.448121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:12.471241Z digest=sha256:958cf4f36eb1fc4e7672dc948ce2c4950d1d45a49f94a5664ce39f3888ebeaf6

Observation ecd6e62b-a10c-4132-826c-e8bee689369d · outbound

This paper cites Beyond self-talk: A communication-centric survey of llm-based multi-agent systems, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Beyond self-talk: A communication-centric survey of llm-based multi-agent systems, 2025

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:14.190529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:12.573521Z digest=sha256:cd3b69af381451bfe3faaf8fbfa07b82d9cea101797a906f951f9ffaae17529f

Observation a476655e-230b-413d-96f5-3ce7b67cff24 · outbound

This paper cites Jailbreak attacks and defenses against large language models: A survey, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Jailbreak attacks and defenses against large language models: A survey, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:14.002521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:12.724619Z digest=sha256:db7f12d9afb8eef7f9226700abb6901cad61c1c91d24682b133a6e7170815591

Observation 75ea36b8-32cb-4a7b-88be-1357fe5c1fc2 · outbound

This paper cites The surprising effec- tiveness of ppo in cooperative, multi-agent games, 2022.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety The surprising effec- tiveness of ppo in cooperative, multi-agent games, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:13.852226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:12.800113Z digest=sha256:5e14b578b19ec53c1b9447270ec28394f15a63d95428a1bdb37e497a238c7ca8

Observation 52a24b10-38cc-42b5-88fb-d5aa4bd7049b · outbound

This paper cites Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents,.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:13.688356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:12.875825Z digest=sha256:215da683278f59cf4ca75494ac160b3e0868e86bc758849948147cfbb9a32bc3

Observation bc6697c0-445f-4a7b-8189-e958168c64c1 · outbound

This paper cites Corba: Contagious recur- sive blocking attacks on multi-agent systems based on large language models, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Corba: Contagious recur- sive blocking attacks on multi-agent systems based on large language models, 2025

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:13.539126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:12.971097Z digest=sha256:dfb93144ee26a12804954a82962141afa3ee18a7074ed0a48615965c46df5251

Observation 54529aa9-6f4a-4b55-b105-ef3c47614e96 · outbound

This paper cites Successful defense on JailBreakV 1.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Successful defense on JailBreakV 1

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:13.390325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:17:13.027364Z digest=sha256:e41e77feeada5cd98c548eb46b2a7a97e3ba840706d895aa0cb2a32e3f73177f

Pith citing papers

Observation 76e6783f-e53e-49ae-baf6-8a3465b90fb8 · inbound

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs cites this paper.

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:00:30.423456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T18:58:53.183734Z digest=sha256:7f958cc3ae3b7bca0a0e3e98c5ca2a88f963bf4e56b476eb60bee973986d1aab