Pith. sign in

Paper Citation Record · LEDGER

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety

As of 9 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2508.03864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.03864 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:17:13.027364Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T18:58:53.183734Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T19:00:30.421345Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy39
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd9bd894-7267-43b7-8062-5e0173a34f7c · outbound

This paper cites Improving retrieval-augmented generation through multi-agent reinforcement learning, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Improving retrieval-augmented generation through multi-agent reinforcement learning, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:21.797033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:08.034162Z digest=sha256:2cc7209079b707d1fb6c0c7e769ef6c08795b10900f632255808183cdbee1207

Observation 563fde29-6a1e-49ba-9089-0e7ebe47144d · outbound

This paper cites Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:21.614175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:08.175865Z digest=sha256:a741f1ea4a34d20f2e924381ddc2f089dddf380b706aea84b8f23ba48f170e37

Observation c6b559a9-3b6a-49c1-b5e5-5dda9b771f89 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:08.292580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:08.292580Z digest=sha256:9c79899a33b2f77c7283aa269aea1a9bec2aaded22581651de3c238333f7135c

Observation 3baf1983-0e9d-49ef-9693-4d51b3a54e74 · outbound

This paper cites Multilingual jailbreak challenges in large language models, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Multilingual jailbreak challenges in large language models, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:21.390561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:08.401691Z digest=sha256:55f1fe3529e67bef108c49d6c8cf6b3ee948973016888c3c1fc82eb5918c4239

Observation be55f796-742a-447e-b510-eb7230a1e39a · outbound

This paper cites A practical memory injection attack against llm agents, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety A practical memory injection attack against llm agents, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:21.121642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:08.541064Z digest=sha256:e974e1293137de52416b3dc2a95d0e1e13cbd4c3fb8971c36f9fde11ac30a8b6

Observation 84c03e86-b527-479e-8d22-668b94824e67 · outbound

This paper cites Peerguard: Defending multi-agent systems against backdoor attacks through mutual reasoning,.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Peerguard: Defending multi-agent systems against backdoor attacks through mutual reasoning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:20.777722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:08.674109Z digest=sha256:6bb7feb878ba67a1912cc6b68de0603fd82bfdcb984e0e49935e806541d188ba

Observation 00c19c84-6117-4bf4-8b22-ab6470007ec3 · outbound

This paper cites Not what you’ve signed up for: Compromising real-world llm-integrated ap- plications with indirect prompt injection, 2023.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Not what you’ve signed up for: Compromising real-world llm-integrated ap- plications with indirect prompt injection, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:20.414779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:08.750735Z digest=sha256:a17721efc3d87146c10cca81858bc95e69503a90fe7b455d01fba04ffe13d9f1

Observation a8e70d4c-72fb-4780-a4f6-80e62addbec8 · outbound

This paper cites Llm multi-agent systems: Challenges and open problems, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Llm multi-agent systems: Challenges and open problems, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:20.053410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:08.835851Z digest=sha256:146c014f30190e6ff4e5c1586c9944aa29803b48e35e8e093d6aac76adcaa583

Observation 438d1589-163d-4f7c-865f-9bedbfe11013 · outbound

This paper cites Red-teaming llm multi-agent systems via commu- nication attacks, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Red-teaming llm multi-agent systems via commu- nication attacks, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:19.752036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:08.902413Z digest=sha256:96c3ac2c51a0f2389b34ba663d0a291b8ee057dac3d983575a84a87cf43fcfb9

Observation c67e3a34-d65b-43b3-b3c5-4a5afc1c68e9 · outbound

This paper cites Measuring mathematical problem solving with the math dataset, 2021.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Measuring mathematical problem solving with the math dataset, 2021

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:19.453212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:09.019799Z digest=sha256:0e2db899f55872448158d8f0b93f49370e30a0035bf39ff9a51247214ee8dc7c

Observation 2a341dd6-e818-45f6-8273-9fa6e85da5a1 · outbound

This paper cites Llama guard: Llm-based input-output safeguard for human-ai con- versations, 2023.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Llama guard: Llm-based input-output safeguard for human-ai con- versations, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:19.144675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:09.190821Z digest=sha256:b2aa4d5bd35d1c3702211620e00bf72908105f463e4f23f2ed02ffcac36b1f61

Observation c633c3ff-51e4-466e-ba03-ead9e4384812 · outbound

This paper cites Search- r1: Training llms to reason and leverage search engines with reinforcement learning, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Search- r1: Training llms to reason and leverage search engines with reinforcement learning, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:18.829033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:09.279321Z digest=sha256:30a602f49e3eddaa9a2f543136678821ab7a2eceb4fba4207e5ce9346b012290

Observation b8bf145c-aa84-451a-86ca-f7c230af4dd8 · outbound

This paper cites an unresolved cited work.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:17:18.524994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:09.394713Z digest=sha256:23a20f36a942fb7fcb453398f5135ca7d85f7baa2685b0fac3dc12099de593bd

Observation d2312e4f-b055-4103-a084-17f8bd5ffb13 · outbound

This paper cites Trust re- gion policy optimisation in multi-agent reinforcement learn- ing, 2022.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Trust re- gion policy optimisation in multi-agent reinforcement learn- ing, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:18.351198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:09.477475Z digest=sha256:a300f4b593bf524797c36d86df676f1321c68e6fa62996ea97b02fb26b8f939c

Observation 3abee350-f220-43c3-a556-2c46afef842e · outbound

This paper cites Prompt infection: Llm-to- llm prompt injection within multi-agent systems, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Prompt infection: Llm-to- llm prompt injection within multi-agent systems, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:18.131372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:09.615213Z digest=sha256:2ba3d5a541b7139a78299731adad4d7f041a5b24ed305cd0b16dc21807c19e5a

Observation 0d027fcb-cd73-457d-815b-59f282d5bedf · outbound

This paper cites Deepinception: Hypnotize large language model to be jailbreaker, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Deepinception: Hypnotize large language model to be jailbreaker, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.931362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:09.701866Z digest=sha256:e4b00dc01116aa5b162331f4d7f466f8d643a422ccb7eb75758dd58fd051f306

Observation 3aaee5ab-09c9-41fd-a037-a51cb2150694 · outbound

This paper cites Tf-attack: Transferable and fast adversarial attacks on large language models, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Tf-attack: Transferable and fast adversarial attacks on large language models, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.769154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:09.794018Z digest=sha256:eaa6902fa8c6d89906597ed826c37046500bd468b520ac4538d7b71ad4c84425

Observation 9ea0f7f3-6939-47e1-af4d-b57dbbb0846e · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.613247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:09.885976Z digest=sha256:16f6cebef849ecf1364b2a422f0fbd650a260a3fb5e90398c1c398c8c396def0

Observation dc5de2ff-03af-4183-821d-9d360b7cd6cc · outbound

This paper cites Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jail- break attacks, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jail- break attacks, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.448546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:09.968243Z digest=sha256:3ae2341acb0130de97a2b1896387d5500a9f98b126de05acd85fc72a784d8bf5

Observation 6b73d23f-c53a-4e18-9395-705fc64b6b13 · outbound

This paper cites Codechameleon: Personalized encryption frame- work for jailbreaking large language models, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Codechameleon: Personalized encryption frame- work for jailbreaking large language models, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.286041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:10.089362Z digest=sha256:8aba0e1fe49220e3ca1c5ef63c8dafeb3e650bd253c3461f83578d3ad1a95586

Observation fc82ba90-e4b7-4cd4-a863-28a7e9356d1e · outbound

This paper cites Harm- bench: A standardized evaluation framework for automated red teaming and robust refusal, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Harm- bench: A standardized evaluation framework for automated red teaming and robust refusal, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.210645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:10.209297Z digest=sha256:2a0599c531ae10c180cbb1d35e15da94ead55cbb6a49a4db9a61bb654d1ad0d7

Observation 172ba3d1-7225-4121-877c-619737d097b6 · outbound

This paper cites Metaspatial: Reinforcing 3d spa- tial reasoning in vlms for the metaverse.arXiv preprint arXiv:2503.18470, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Metaspatial: Reinforcing 3d spa- tial reasoning in vlms for the metaverse.arXiv preprint arXiv:2503.18470, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:10.290349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:10.290349Z digest=sha256:dd1810c8ea498619081596a790449ea968a0d926a82f167515187381408c51b6

Observation dcfa71a7-5459-452b-b77c-4839d76b6df8 · outbound

This paper cites Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:10.380062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:10.380062Z digest=sha256:3fc5669868506a126c5b24cb43a6e1290312bd037a2c0f1d8b9554e1d2a25a0d

Observation a9cd9caf-e8d6-4c21-8bd6-cf11e8a8ad7b · outbound

This paper cites Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:10.475488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:10.475488Z digest=sha256:4737d586ae562bfb1708b8b0225e1bc3fcd85cecfd410bd02379392e4d203c22

Observation 7d953f70-da50-44d8-913c-25ea686208ab · outbound

This paper cites Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-Action.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-Action

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:10.574459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:10.574459Z digest=sha256:b8ec0a6b5171f8c0cdca99a9f94b0bfaada1ba69968686530904c6827c12d686

Observation 26f531b3-0c99-4667-9d91-3e65103ea2e0 · outbound

This paper cites Do code llms understand design patterns? In2025 IEEE/ACM Inter- national Workshop on Large Language Models for Code (LLM4Code), pages 209–212.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Do code llms understand design patterns? In2025 IEEE/ACM Inter- national Workshop on Large Language Models for Code (LLM4Code), pages 209–212

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:17.070627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:10.694884Z digest=sha256:3eae647429e2d18daf00b6bd1dcd01e5ebb45c4b6b637e5d447de53a766b2b51

Observation dd25c245-3dfc-493f-ac18-7da47383aca2 · outbound

This paper cites Yu, Manling Li, and Han Liu.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Yu, Manling Li, and Han Liu

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.906588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:10.828016Z digest=sha256:98c652a205035b8eaeae6d142e9d1fb8906ad4e9e4de00dd56d6d7de55c617e9

Observation afb8e609-e62e-4f80-8620-b874fa7db811 · outbound

This paper cites Maporl: Multi-agent post-co-training for collaborative large language models with reinforcement learning, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Maporl: Multi-agent post-co-training for collaborative large language models with reinforcement learning, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.737349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:10.912276Z digest=sha256:f98a38c12ddb79731d54c83725c94da60013b38a9d4abeac1d9a94a678b460d4

Observation 0fbbc2b8-4635-4abb-9151-e7ad4b34d2a3 · outbound

This paper cites Red teaming the mind of the machine: A systematic evaluation of prompt injection and jailbreak vul- nerabilities in llms, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Red teaming the mind of the machine: A systematic evaluation of prompt injection and jailbreak vul- nerabilities in llms, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.581769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:11.050008Z digest=sha256:c9b8c40ea073e1113a9dad506ea76eb550e77a09e1c0aeedee82720d3f37bc9c

Observation 30435057-2a00-48ea-bb19-38d223ce91a7 · outbound

This paper cites Qwen2.5 technical report, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Qwen2.5 technical report, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.462728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:11.156744Z digest=sha256:ed5c5a1a47e4bdc1cc9e090dee56dea115f19cd3538d900f2201f5136efc43f7

Observation f5f5cdb3-ce4f-4001-aad7-796e434fe262 · outbound

This paper cites Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning, 2018.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning, 2018

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.270429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:11.291722Z digest=sha256:96c883c7f28f8fbaa3c03bb8d13c76ab7baf7ddb76b86ef477107a841b3ccec8

Observation cd0c0ce3-4ff0-46e8-9a8c-a4a2626487e5 · outbound

This paper cites Proximal policy optimization algo- rithms, 2017.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Proximal policy optimization algo- rithms, 2017

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:16.069560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:11.390773Z digest=sha256:f0941f660899890290a429605f580c1c357e7ff1d2187a112c343b7700dbce7c

Observation 7313c8ca-3796-41ce-866b-0df1da0a3aac · outbound

This paper cites Exfiltration of personal information from chatgpt via prompt injection, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Exfiltration of personal information from chatgpt via prompt injection, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:15.898471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:11.468085Z digest=sha256:b3c2fade36ef7f2a6c45700c70f283dfcd4ad0a10428f85e2bd794b43798ae28

Observation e234953e-7426-45a9-ab15-f4820562e85d · outbound

This paper cites Sciscigpt: Advancing human- ai collaboration in the science of science.arXiv preprint arXiv:2504.05559, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Sciscigpt: Advancing human- ai collaboration in the science of science.arXiv preprint arXiv:2504.05559, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:11.579252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:11.579252Z digest=sha256:1c53979f00bc4aef9d0561ae62fbc64bb5e457ae8cf58e710cc33e6246235f6e

Observation ee7eff76-68f7-44ac-b859-43b5a7ce914d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:11.673708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:11.673708Z digest=sha256:c9cbd183af70af9c2b17be083fd4f701992fe39fc55a8216f29f8b7726643e6a

Observation 5bc28a52-9f95-45d8-be1a-dfaabc07693a · outbound

This paper cites Prompt injection attack to tool selection in llm agents, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Prompt injection attack to tool selection in llm agents, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:15.712033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:11.775209Z digest=sha256:90a0cf5d5975eb20dbd8ea9219c8c756ba77533f5026edacafdbb153eb10a213

Observation 8ce99069-7b5a-4ac8-afea-8b737881ab91 · outbound

This paper cites an unresolved cited work.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:17:15.484083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:11.868960Z digest=sha256:52b24f9657fad4facdd92dc4a7a16a8f483134a7f1ad559c9fe4ab0ed62b0a3d

Observation c2759c17-5d28-48fd-a776-64d70bbd1d72 · outbound

This paper cites Multi- agent systems execute arbitrary malicious code, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Multi- agent systems execute arbitrary malicious code, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:15.296371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:11.980418Z digest=sha256:22e617925d8a3ac8629cfefa1826c9c512264e58edfe3c82c0f3f8d2da9cdd75

Observation a735f8f2-8ca2-4fb3-b43f-be9f120ba091 · outbound

This paper cites Lyu, and Maarten Sap.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Lyu, and Maarten Sap

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:15.122813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:12.091016Z digest=sha256:5d0dc24f9b2ad44effcdf61e3876d19dcf86c149f2bd6ec576bbbe24620a6722

Observation 8534057d-95d5-4384-b11b-5f76d7b748fb · outbound

This paper cites Rema: Learning to meta-think for llms with multi-agent reinforcement learning, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Rema: Learning to meta-think for llms with multi-agent reinforcement learning, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:14.972763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:12.199699Z digest=sha256:8f49618f06d2cfedfab86653ea58fca9a9b25ae75e8bbc65d41db65a9fba9ba0

Observation 5f09d10d-b5ea-470b-9983-6738b143438c · outbound

This paper cites G- safeguard: A topology-guided security lens and treatment on llm-based multi-agent systems, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety G- safeguard: A topology-guided security lens and treatment on llm-based multi-agent systems, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:14.794982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:12.318504Z digest=sha256:34def287f2f3b84a8ed9fbe18a937b0ef866a9530f4479d31a764a9bb8f6c450

Observation f5b6ad89-c513-4b99-82eb-a695d9e10541 · outbound

This paper cites Unleashing the emergent cogni- tive synergy in large language models: A task-solving agent through multi-persona self-collaboration, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Unleashing the emergent cogni- tive synergy in large language models: A task-solving agent through multi-persona self-collaboration, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:14.643499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:12.409382Z digest=sha256:dfd9007188eab91337971ef8fb87ec5f13e14ac09de45769c1a900550731197c

Observation 5db696b9-5838-47c4-b9da-cf88ac3b0ded · outbound

This paper cites an unresolved cited work.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:17:14.448121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:12.471241Z digest=sha256:ea753976b0b5e9fea01ccf7ef338e96b5da16916973181ff1b81ff91eea676c2

Observation ecd6e62b-a10c-4132-826c-e8bee689369d · outbound

This paper cites Beyond self-talk: A communication-centric survey of llm-based multi-agent systems, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Beyond self-talk: A communication-centric survey of llm-based multi-agent systems, 2025

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:14.190529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:12.573521Z digest=sha256:041358d85c25bca655d679e3c23d3b89e0323bf78a12d9ec23390eb7eecffc71

Observation a476655e-230b-413d-96f5-3ce7b67cff24 · outbound

This paper cites Jailbreak attacks and defenses against large language models: A survey, 2024.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Jailbreak attacks and defenses against large language models: A survey, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:14.002521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:12.724619Z digest=sha256:efbf9a628bab0550dfd440677ad5d4459b25c856fde48afd46f5db7bfa1a84e0

Observation 75ea36b8-32cb-4a7b-88be-1357fe5c1fc2 · outbound

This paper cites The surprising effec- tiveness of ppo in cooperative, multi-agent games, 2022.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety The surprising effec- tiveness of ppo in cooperative, multi-agent games, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:13.852226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:12.800113Z digest=sha256:fb1e19f289691dc6dea7ef82b0474b145c1ed4633027b7890d7bc04fe759c533

Observation 52a24b10-38cc-42b5-88fb-d5aa4bd7049b · outbound

This paper cites Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents,.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:13.688356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:12.875825Z digest=sha256:aa40356d8e506127c814bfbcf56f84f1124962c7e0220667a6e1b085ce0e24b3

Observation bc6697c0-445f-4a7b-8189-e958168c64c1 · outbound

This paper cites Corba: Contagious recur- sive blocking attacks on multi-agent systems based on large language models, 2025.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Corba: Contagious recur- sive blocking attacks on multi-agent systems based on large language models, 2025

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:13.539126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:12.971097Z digest=sha256:c6a135fb355c26dce09b8beda32739bf26d1024435e2f8a9d3f5a093c5b9be89

Observation 54529aa9-6f4a-4b55-b105-ef3c47614e96 · outbound

This paper cites Successful defense on JailBreakV 1.

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety Successful defense on JailBreakV 1

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:13.390325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T04:17:13.027364Z digest=sha256:8f21f30d9ba8f4c66596ba1aa1003c3e0124a1660bfc3e0b1004ba95a23bebdc

Pith citing papers

Observation 76e6783f-e53e-49ae-baf6-8a3465b90fb8 · inbound

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs cites this paper.

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:00:30.423456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T18:58:53.183734Z digest=sha256:75f020f1b571e35c71230f58add129dd636d683186ec9500152db842ce284903