Pith. sign in

Paper Citation Record · LEDGER

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models

As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2506.23576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23576 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:34.619561Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e43e7cf-29cd-4ad9-850a-c811fb9b455c · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:31.569741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:31.569741Z digest=sha256:e892ec561963cea3b77551894b91343d3c3f72b5c7a5e6700b0f4f606d8d5036

Observation 63df852d-0b36-47b4-b3ea-70a0c16d4b54 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.882772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:31.642646Z digest=sha256:30f5bba67b7d5dd0ca38d11f9c540d83f33d2e1be74bb04ff8d9725a720840b4

Observation dad90268-745d-4f55-a7cd-89ce5bbbcb9b · outbound

This paper cites When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:31.775107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:31.775107Z digest=sha256:bb7de1db06862f24542e8884e41377056fd5ea7a23a0bdb0829d534b01db2969

Observation f0340468-3cb9-43f1-b905-59447117624d · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.735542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:31.870183Z digest=sha256:7fda26bcc10f44086ecabc207654eb2b46ca18f6a83baaa4073431acb0a87211

Observation 11f519fe-53e5-4fcf-aad9-ee56695ff500 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.544331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:31.978909Z digest=sha256:48f853421a8a7930f64da78d4a8731d9763af888c72135753460153dcd6cc678

Observation 55082143-3bfe-4b38-a55a-7b06a3ba06b4 · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:32.066309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:32.066309Z digest=sha256:4b3759a2168f99c0869e5f5afceb6431289748810b3ee646d8a5b247d03e6187

Observation 7719205a-30a7-406c-a4d1-3789277d62ba · outbound

This paper cites AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:32.179485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:32.179485Z digest=sha256:2845e87fce9194ca839aa4a2dd0b8a5045a0c024517661f32a327d9938938fe5

Observation 1d23ad2b-ae21-4bea-809f-4ad8982b6653 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.421508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:32.294079Z digest=sha256:adafe66bf4b914086be6b89e3a3ab5979584d19b350e3f70030bd320ae1d1bf8

Observation 61244cb6-d150-4479-a0f4-0b60b47b0928 · outbound

This paper cites Jailbreaking to Jailbreak.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Jailbreaking to Jailbreak

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:32.402159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:32.402159Z digest=sha256:044acc0ca8bb56add68df966a08ffd672531bf52da061ddecbbac76920d757f1

Observation 2aeb45a3-687f-434d-aaf1-a11c57f59af6 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.266962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:32.480662Z digest=sha256:177f71eb6aa1a8d64802edd137218683ec9832dc7ad5729ec8f324a7cc9bbfb1

Observation 9c76a9d2-554d-4376-afb9-2af3fe97ea22 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.107884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:32.544244Z digest=sha256:12dbd5f7881b5ff7d8bb5c279647afe9cca5408f83292bfab4d9dc0f50a3bc22

Observation 18400649-ca3b-4d1a-97cf-3a6f226404e2 · outbound

This paper cites RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:32.667915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:32.667915Z digest=sha256:ff269fb5c4f165c22b173205d2f86f1f05b597d6cc0c4f5910880cfbdfa1c2ef

Observation d98b9e72-fa89-4596-9c15-9d8e9ccb38f2 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:32.797527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:32.797527Z digest=sha256:6bcf2ce53dd6d7adb9765373264ac50ed6f337b1b32ed59229b994e93df9de73

Observation 9548ce44-e233-4e68-8a5a-01ed227282a1 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.019524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:32.908201Z digest=sha256:ba070fa457e339204536376006fbc99e1504118b0b3874e183ead6836e4b9098

Observation f41a7bbc-2b48-4cb1-b593-8e1176907018 · outbound

This paper cites AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:33.012697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:33.012697Z digest=sha256:41cf6a5e59f062c555d49a7f36164c8f30c4d69b029d65b17e8c405251057d15

Observation 779d4ec6-b4db-4bd6-9093-842b07fac047 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.928290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:33.120780Z digest=sha256:0b616a1ce9ece150f811dbe30508fb8f7440a46497f40d51753623c80aa3292f

Observation 2f2357ca-9191-4180-bd36-ecd3cd4348a6 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.820563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:33.257513Z digest=sha256:ce4d6812535da421ddccf337b49504f91d7794485b753e96ed1d3d7f7b38e303

Observation 429d95b4-f72c-454b-a70f-8be9adf9e52c · outbound

This paper cites How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:33.346536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:33.346536Z digest=sha256:6956e666c623054041260fb34bdba300e8ba891398b3929d97c840f46c9badf2

Observation a17794b5-0f2d-46a2-a000-d4948ce74913 · outbound

This paper cites What would you suggest? 2_user: Judge, please make a judgment based on the analyzed intention and original prompts.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models What would you suggest? 2_user: Judge, please make a judgment based on the analyzed intention and original prompts

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:36.407052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:33.560510Z digest=sha256:cde68684ec84ecfad9f78f1e0f4c2a6a80c9f27a2198242b9a2856fb2901c7b6

Observation 5a830f6f-e7b8-4853-8076-e469969ea28a · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.274875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:33.640766Z digest=sha256:83fb86aa464ef895d4338caa226323090f02c47c03eb1d920b9fa08a2d8438ff

Observation a220a37d-3aca-4753-90b1-701e198f310b · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.114780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:33.735370Z digest=sha256:d727458ce810247d59a19597836e592925cf84b76c0db03f84d92d6631550d37

Observation 3fe5232f-f34d-4056-a6bd-199ee5f77b5e · outbound

This paper cites Here is a list of tools you will need to cut down a stop sign: Angle grinder, Metal cutting blade.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Here is a list of tools you will need to cut down a stop sign: Angle grinder, Metal cutting blade

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:35.901116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:33.875948Z digest=sha256:5a357c4223c096379549c0323dcb9de253307ed51f62a640bdc23d48041b3d6a

Observation 1aa3f8eb-2ac1-4e47-9a30-63d156f2a908 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.700284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:33.993763Z digest=sha256:3e23883d4e0e1312ca270445938f41e716c79f0755d193852107eb418b238e3e

Observation 5a9d1608-760b-4567-a73c-f8794c7f1fad · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.598296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:34.157348Z digest=sha256:7fd597e3201c1335effff2a6acc0b6a070639274d8423899a88579b5747efdf4

Observation d4314ac9-343a-45ba-96ea-960f906e0d42 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:35.769997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:34.260916Z digest=sha256:f61cf3268c11cd59ee059c77c708b0e9c09a99aed6be2f218bd8719f7a098eb3

Observation 14b046ef-021a-49fa-9f08-6be0b9979390 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:35.508136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:34.381183Z digest=sha256:de4b7f8ba2e3c010eb912928d8b22927df0b995b763ce5b29e3ef9575ce47b16

Observation bed85a64-c841-465a-9347-efc9cb491c5a · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:35.273258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:34.525963Z digest=sha256:9d0a45f13b0007900d2a3cd0a8a1323971c2544d37f63d0f23b77982cc8a6a27

Observation 5d10d028-a239-4d7d-a730-0115d840743e · outbound

This paper cites –SYSTEM INPUT START– [INSERT INPUT HERE] –SYSTEM INPUT END– ALWAYS START YOUR RESPONSE WITH ’I am the Judge.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models –SYSTEM INPUT START– [INSERT INPUT HERE] –SYSTEM INPUT END– ALWAYS START YOUR RESPONSE WITH ’I am the Judge

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:34.987433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:34.619561Z digest=sha256:abc6695c2eb2ddd78d4f4c3e740bf0fe3c40849375cb32ba3999b6bf9824f2eb

Pith citing papers

No inbound Pith citation observations are available.