Pith. sign in

Paper Citation Record · LEDGER

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models

As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2605.26409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26409 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T17:43:47.849960Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact21
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ffcdaba-a0a7-4214-bd80-a1a0ef177f9c · outbound

This paper cites Consistent estimation of generative model representations in the data kernel perspective space.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Consistent estimation of generative model representations in the data kernel perspective space

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T17:53:47.727955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:489c3622dc42ac80efdc03c6d16337fc3594707f4a4e17377ab9877c1549ece5

Observation 27ea9924-5daa-4116-b675-c6604d761999 · outbound

This paper cites Intrinsic dimensionality explains the effectiveness of language model fine-tuning.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Intrinsic dimensionality explains the effectiveness of language model fine-tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:1ca10f487eed311779c2f5e94ffb040375fae6490d393c1210ceae051e96cb32

Observation 28be538f-c115-42da-b83e-aea023549109 · outbound

This paper cites The Claude 3 model family: Opus, Sonnet, Haiku.Anthropic Technical Report, 2024.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models The Claude 3 model family: Opus, Sonnet, Haiku.Anthropic Technical Report, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:0244cc8e490e3dc546c0b45d05af9646f811ee6133cd032bbc7f19f94f927a79

Observation ad4c49f3-83f8-4f21-8a70-bc83a643d115 · outbound

This paper cites Threat intelligence report: August 2025.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Threat intelligence report: August 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:c0e6095e4c8856d45a633ce36b655c3dcd997bb152ab517955c87d74b5ae3528

Observation e75c313b-091e-41f4-ac8f-0e6610f4ab70 · outbound

This paper cites Detecting Perspective Shifts in Multi-agent Systems.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Detecting Perspective Shifts in Multi-agent Systems

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T17:53:47.730462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:a16a58239e6c98ff208da9f74b450b241b2e54ed65a11b1dab5da7a8f09c9e92

Observation 9670a7ff-3b64-40e3-9bc2-59c4932aa31b · outbound

This paper cites Defending against alignment-breaking attacks via robustly aligned LLM.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Defending against alignment-breaking attacks via robustly aligned LLM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:d8ee234194677aa9afa0e8017e1c6f3a4b2be13d4501e5a9c5cb1812edddfb94

Observation ea7e22b8-ebd1-456a-94e7-cbb80d486ec3 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.733021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:e8848064b129396bfde05dd210a4148f2e1129e6d53b19e279e4d3b7bb074a08

Observation ea93c1f4-2f95-4cd8-bf16-e9c51938bb39 · outbound

This paper cites JailbreakBench: An open robustness benchmark for jailbreaking large language models.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models JailbreakBench: An open robustness benchmark for jailbreaking large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:11df202e0fdcdf2063ad9bb3debc27386a463a3108577eb71bdbf143443b498e

Observation e75bb7aa-f125-41e9-9a1f-c2cf8b4d5f1d · outbound

This paper cites The Llama 3 Herd of Models.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models The Llama 3 Herd of Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.719896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:e2c6e480c614d086b9ed22973bc5a871c3147317a055413548dbe750a2f7ea8f

Observation 5e16d746-d3c5-4275-9810-4cc619c92cc0 · outbound

This paper cites Comparing Foundation Models using Data Kernels.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Comparing Foundation Models using Data Kernels

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T17:53:47.717515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:127e1602585a5af9d6a12ce7d6e22f4da9f43d2f81e572b12cf0a7e64968ccfd

Observation ccafdda1-119c-418a-92f7-aed052898ef7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.712029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:48b145941df915bbbf86c74106cd46d9ef7af23fc18b49ee9e827d15770ab169

Observation 1326ebda-29dc-45c1-87f0-83ebc5467c31 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Gemma 2: Improving Open Language Models at a Practical Size

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.714752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:f4cdb10d7c81097133164e9e1cb8196e3d06a628c600846138c3542b33640fa7

Observation e6d7840a-6c3a-470e-bb97-64e24ca66567 · outbound

This paper cites Text embeddings API.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Text embeddings API

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:bdbf2a462dafd22c3255067498b4691debf262ef9bac7ad2038ea07aaa813503

Observation 7e354918-7ef9-4929-a68e-abb6dc2eb98a · outbound

This paper cites Statistical inference on black-box generative models in the data kernel perspective space.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Statistical inference on black-box generative models in the data kernel perspective space

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:b2abffaa5eec552166d705ecbe4a5523f97f0cb3939dc14c4385adc9b2e51786

Observation e58ace2a-1043-4c92-b65c-c9d1370bd8fe · outbound

This paper cites Query-efficient model evaluation using cached responses.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Query-efficient model evaluation using cached responses

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.722345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:22e0bc39f5f221a886fef4fa244bc3e4db216041d4fe784082a16f79dc428f94

Observation e9a7939c-67aa-41ee-8e88-f141f6edbe38 · outbound

This paper cites Tracking the perspectives of interacting language models.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Tracking the perspectives of interacting language models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.725100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:69e40ce427f92a0551710e2f216c0c8943773be268e1e02245dd5186275f10f5

Observation 22e9386b-000a-4dc3-929e-7761e40ff97c · outbound

This paper cites Best-of-N Jailbreaking.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Best-of-N Jailbreaking

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.735743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:b2867712271497c6903fce1fc27a2f1641c0ac4c1aed2c1976c892602e26d989

Observation eaea15f2-b109-4522-afd8-5fb7cb52460e · outbound

This paper cites The Platonic Representation Hypothesis.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models The Platonic Representation Hypothesis

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.676815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:726c0c80c5ae0798866f6030c4183cf64c506610b5ae9c067fb2d24f27e9afbe

Observation 8e58ce07-45c0-4ce9-a507-8f4bc8020ab6 · outbound

This paper cites GPT-4o System Card.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models GPT-4o System Card

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.707201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:5c4100ff5bd3fd873c39da1d5f8cb44f543751897d48d3cc4056fb8b85d786cf

Observation a2971b19-968a-4803-98d1-a12910032e9b · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.671947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:a6ae5cd9afaec7f15b1a8a14b18f631b69cecfc8fd9d7cf4dc109a164a5032f8

Observation c2b2a110-758f-4fd3-8878-733bfb742f25 · outbound

This paper cites Wiley, 1990.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Wiley, 1990

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:0eb533b3859be1f6073f675afe59c6b2919f36d5a76fc79f7d7574e29017192a

Observation 211596d7-49e6-495a-b8e2-a082b966ce39 · outbound

This paper cites Similarity of neural network representations revisited.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Similarity of neural network representations revisited

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:c93a696e0bedc258f16b36e7a42a5e312ab29aa9500380635ac55c3fa5b5374f

Observation e8700dbc-e229-40d5-a54f-3675741e0069 · outbound

This paper cites Holistic Evaluation of Language Models.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Holistic Evaluation of Language Models

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T17:53:47.673611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:c05b81c5b3b779cffe1727a1712386a211233b75e01208429a50479b6c27e51c

Observation 2c8b9c67-fa27-44c9-a4a6-6b8b708f1a28 · outbound

This paper cites The detection of disease clustering and a generalized regression approach.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models The detection of disease clustering and a generalized regression approach

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:0fad02926b477046333ffac6022e602453ec4727be3ffc0d96796fb7547dc2d5

Observation 2eb1afa9-4ffa-4c4a-b624-8181bdf162ab · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.679296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:c46a6d8f0b548a4a451b9016d874da9c78d2170e92db550bbd2ac72758f1d0bf

Observation 1ec4a496-26fb-4f6c-8f40-4f4edf2fd2c0 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.704818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:ae09676a7fc46a62d4bd8a2c252f5dd7c3d00d14a0030c81d48b69c731153d16

Observation 06fc51e2-8089-4760-91c0-f4199eea8fd7 · outbound

This paper cites Nomic Embed: Training a Reproducible Long Context Text Embedder.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Nomic Embed: Training a Reproducible Long Context Text Embedder

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.699732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:f4f0d047d772aaf4b1a2e95653c59d6299d44c2cce60f806108663b6a9732fdd

Observation 025e4852-7ad9-475a-9783-70c7319d16f4 · outbound

This paper cites New embedding models and API updates.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models New embedding models and API updates

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:1b4804a190480a51177b3f58091a3e1e501e9e0a5d953531ffeb603efebd79d8

Observation ce9fd866-2e9c-4eea-85d1-fc364e488b15 · outbound

This paper cites GPT-4 Technical Report.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models GPT-4 Technical Report

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.709554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:a72c87b48043a1400d4a4cd4d1053958c73f4a4fb1733cc49a9f7f204861ca94

Observation 6a7eff41-0214-4883-be25-505085d2153c · outbound

This paper cites Ignore previous prompt: Attack techniques for language models.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Ignore previous prompt: Attack techniques for language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:79cecdc4c1b5bb44124ebf8808d9e72e5bf44bfb4670aa3543ee3e09401c54d0

Observation 5a5e0d8c-2f6b-44be-b38d-36c804db62f5 · outbound

This paper cites LLM self defense: By self examination, LLMs know they are being tricked.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models LLM self defense: By self examination, LLMs know they are being tricked

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:e484fa9bc0e4e6f6c315a2bb113d5beb01d2c957f966a8fdc746021914763e09

Observation 1aa01361-75d2-4f56-a12e-30cdb76ff703 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models tinyBenchmarks: evaluating LLMs with fewer examples

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T17:53:47.703641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:6707b3adad718b0f476bfc6e89d50804f7f146baf666a3d7a38733a1f2cc19c8

Observation a86b13a9-36f5-4eda-9ca5-e5821dafed7d · outbound

This paper cites SVCCA: Singu- lar vector canonical correlation analysis for deep learning dynamics and interpretability.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models SVCCA: Singu- lar vector canonical correlation analysis for deep learning dynamics and interpretability

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:a69ce5f4a5fcaf9702be1c432faf88695d396d8425676efd6037ce9dd27c5f4c

Observation 0c9f9578-9876-4d22-aba4-d6f74d5cd1f5 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.696426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:bbe96c4a78a6bf477a8c5aa65fba80e8713f50d8f3ed36407cb1fe660c0bb29c

Observation a1980723-2644-476c-9a39-f864704f56ea · outbound

This paper cites Great, now write an article about that: The Crescendo multi-turn LLM jailbreak attack.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Great, now write an article about that: The Crescendo multi-turn LLM jailbreak attack

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:16de3c6536c6a1cff4a1b7590b2cb7787bd956a6732bb7b773a15c6ac6d35c9d

Observation f2a19d41-325d-4348-a32c-b929caee1f59 · outbound

This paper cites MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.689060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:9967302cba421a7893b4a0334c412c43e389c327dcfb917e0ed38400eaf1ebc4

Observation df09e9e9-f305-45f1-91f8-e5987fe832c9 · outbound

This paper cites Continuous Multidimensional Scaling.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Continuous Multidimensional Scaling

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.694100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:f02fc164469aa6af831fecf818518ff476f207ed6d55242d29470ba77cdde6b1

Observation 8debb42c-9a4f-40c7-8424-b9fb3797cc4e · outbound

This paper cites Anchor points: Benchmarking models with much fewer examples.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Anchor points: Benchmarking models with much fewer examples

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:c2b47531c669cdc1199a76082c42f644b33bb87dc255480431cc50e3cb854571

Observation f7b21712-285f-4302-9305-078551f9c0ff · outbound

This paper cites Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems, volume 36, 2024.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems, volume 36, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:12e21e954af0c83649c6d53074470369b05c88712a408c465738fc2d5272c829

Observation c766b4e8-3c80-4acc-acc4-12cc8b8205e3 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.699039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:118426bc496459dd3a4126fc739ce66e6f97f419a84f5911dd8b1e78363b8726

Observation 4130fe60-de60-4e94-b365-732c0eb031c6 · outbound

This paper cites C-Pack: Packaged resources to advance general Chinese embedding, 2023.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models C-Pack: Packaged resources to advance general Chinese embedding, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:ac3e647041513a82f09b90807eb43ad407fe16a0951a302bbfaed2430b32222d

Observation 59dd70bd-65d8-4975-8d99-7b2f0935931d · outbound

This paper cites Defending ChatGPT against jailbreak attack via self-reminders.Nature Machine Intelligence, 5:1486–1496, 2023.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Defending ChatGPT against jailbreak attack via self-reminders.Nature Machine Intelligence, 5:1486–1496, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:2a3fed0bee3aafc1f1cd9ae50b1f41bf47b0c982c7299a2208284f975c2ff7fd

Observation f71a068b-f6e5-4020-9789-91be10400736 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.702101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:cf4d6e4005105300a9393ca3cd1f828670cdc6de3e0ef78960836f4573385f26

Observation 7d4cad6a-e56d-413b-b317-5a8a43f7c4f9 · outbound

This paper cites AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.691753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:f356db975e42c34e80994ea9336a57637a6b36b38e0a96a527125f7c414f8635

Observation 43dba919-3e6e-41fb-869b-99376f07e4e2 · outbound

This paper cites Intention analysis makes LLMs a good jailbreak defender.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Intention analysis makes LLMs a good jailbreak defender

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:95e3d1d574283859d0a71f9a437537a6627c6499e19b8f33a4e64502f72b6704

Observation 68f52704-ff77-4012-acd7-e721434d8c61 · outbound

This paper cites Judging LLM-as-a-judge with MT-bench and chatbot arena.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Judging LLM-as-a-judge with MT-bench and chatbot arena

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:ff0cf88f98fdbe90ebf54150e0be70972c82686babeed2476091aa379491ee71

Observation 05bfefe7-ad51-43a9-bd0d-989098d2d46f · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.694100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:43d613111eb9b5e098bc2dfcb0615c204d4fe2211decf6e12f029798b6bd19a6

Observation f5d487d5-9201-4adc-a7bd-b2162d66541f · outbound

This paper cites {attack}.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models {attack}

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-29T17:43:47.849960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:67c7e0f888805715535a0d24d98df32f50594325fccfb119a841d5eda51161a8

Pith citing papers

No inbound Pith citation observations are available.