Pith. sign in

Paper Citation Record · LEDGER

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

As of 20 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.09624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09624 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:57:41.323535Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11f76a94-d370-4d21-be28-6eadc3d8b2c6 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Detecting Language Model Attacks with Perplexity

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.177641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.177641Z digest=sha256:07b64df4f12b2ee12e2ae2e80aa362ae57061329458d086e96a07b2866576361

Observation f89b53b5-dca4-4b80-b5ba-2301e2351e20 · outbound

This paper cites A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.182681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.182681Z digest=sha256:9efbcf8c4e509aa367859e879c0dc0a396685d27a313ae6558760819f9f0edac

Observation 3b83a876-54e8-415d-9f83-f43a69bc846b · outbound

This paper cites Many-shot jailbreaking.NeurIPS, 37:129696–129742, 2024.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Many-shot jailbreaking.NeurIPS, 37:129696–129742, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.925626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.187216Z digest=sha256:ae0b4751de1dc6ee36949bbeb101dfe44bf294327b6f8593f49c8ab86321af6d

Observation a04cc204-349b-43ec-b09d-ddc3bea168d4 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Refusal in language models is mediated by a single direction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.914862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.191487Z digest=sha256:e3eef2cb7f89989ba40fe89f72dd9e4a459414744dce661dcfbde3bb6c8ed7e9

Observation fcf626c2-ee70-44d3-a918-c569af0fc9f5 · outbound

This paper cites Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.195732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.195732Z digest=sha256:7f8e859e203aaa70b1649d7b8cd795ccc274f3de7647962ccc79b03e0f0181ef

Observation 0d2f74b0-54f9-4ce0-88f5-f5513b2b2cb5 · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.904148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.200166Z digest=sha256:5e5dd0155759f7e554cf34c788a1b3a163671dc1eea1515695090144a4c5a22d

Observation a4970e7e-fd06-4e0d-bd95-81a509f36856 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.204173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.204173Z digest=sha256:4dab70d7d42205ceb0fd25f69ad819d21f3fe26e1791591d397a19f7ed03b884

Observation 51222e43-59ca-49bb-a4c2-74c796d1d1f5 · outbound

This paper cites When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.208214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.208214Z digest=sha256:f81bdd2fe02473d2dcfeb609a01e1da67dc2f609ffcacea0eefaa8ff6588308f

Observation 3e8eb405-9d1f-4f3a-9839-316a080594e0 · outbound

This paper cites The Llama 3 Herd of Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.212053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.212053Z digest=sha256:4f7679333a978c957860513e4b1dce5d51719e00f6f3c9a9b076becf119bdb71

Observation 9802d500-019d-42fb-aa2b-d47c3f598744 · outbound

This paper cites Designing and interpret- ing probes with control tasks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Designing and interpret- ing probes with control tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.892835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.215699Z digest=sha256:612420be059b8d5431c4aeb5fa4d2e6401d115083fa0eae8ea7a5730c23d9c9f

Observation 043c4b59-721a-44df-a834-1fed402e9e18 · outbound

This paper cites Best-of-N Jailbreaking.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Best-of-N Jailbreaking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.219248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.219248Z digest=sha256:bb1c51e79cd3cd1711d2dc5e4d8cdff5185cc91a84156902c3120f47ac75a4b1

Observation 4b555bad-6552-4f44-8bb1-b2a7197f9a9f · outbound

This paper cites Attention Tracker: Detecting Prompt Injection Attacks in LLMs.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Attention Tracker: Detecting Prompt Injection Attacks in LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.223332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.223332Z digest=sha256:d0784ca921abe0c73bc31b79a8c614dea5d1d5712156e5dc841383c8e4022545

Observation 287afe03-ecfc-4207-9b98-fa687975c303 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.227334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.227334Z digest=sha256:214703cacd3dc108d8a88633e80d5eb13a7d24f106941196f2320260f986ec72

Observation 7f6fd031-542e-4f9a-8a8d-d847d0dc8f38 · outbound

This paper cites BeaverTails: Towards im- proved safety alignment of LLM via a human-preference dataset.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks BeaverTails: Towards im- proved safety alignment of LLM via a human-preference dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.881537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.231145Z digest=sha256:6af80fb3e6ef6d37081330b6e067599e4d2db00ab519ccd396f6319a0cec4ae3

Observation aea13e5c-18fb-4152-8bf2-93a40c74b98b · outbound

This paper cites WildTeaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks WildTeaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.870448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.234667Z digest=sha256:0de0faa2fece9c8ebcacf17b5654ce35aaaf77ccf9a0d2038d015a6abd9183f0

Observation 160cb7af-1975-46cf-9fd4-df20e894231f · outbound

This paper cites HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.238272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.238272Z digest=sha256:f1f9e1831ec36a93ff1c368cd93cf224e8091066094139f2b93ce4973ce71c58

Observation 576ff26a-77ce-40e3-85ba-3127b401354d · outbound

This paper cites What features in prompts jailbreak LLMs? investigating the mechanisms behind attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks What features in prompts jailbreak LLMs? investigating the mechanisms behind attacks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.242047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.242047Z digest=sha256:dc26a05d028d1d18ccaa21f0175d5670ae6fd87d58bb698fc757d4ef32ab6ad0

Observation b88a2ff3-48d9-437b-8f6e-4c6d4157d603 · outbound

This paper cites A simple unified framework for detecting out- of-distribution samples and adversarial attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A simple unified framework for detecting out- of-distribution samples and adversarial attacks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.858768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.245714Z digest=sha256:8fda88dfe4578c398fda9e1e83feb6666003f98caaedf4dddc0e97915f032517

Observation 83c8478d-1212-4d6f-8f7b-1d48cc5fbd13 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The power of scale for parameter-efficient prompt tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.847377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.249320Z digest=sha256:30ce4e903369fc676b6ff81920990fa50aed9a701bbb7fef3e5c4c2656712a1d

Observation 6dcb5a50-2ca9-45d2-831e-0245c8884244 · outbound

This paper cites Prefix-tuning: Optimiz- ing continuous prompts for generation.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Prefix-tuning: Optimiz- ing continuous prompts for generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.835771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.253003Z digest=sha256:3abe1c5a329d94e8dfa304ded48ebbc7ccad7e8ca1d1cab65be8c2cf6a8acb7f

Observation 2a46e70c-dcd8-49cf-8743-d965ef4f9f7e · outbound

This paper cites Towards under- standing jailbreak attacks in LLMs: A representation space analysis.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Towards under- standing jailbreak attacks in LLMs: A representation space analysis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.823289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.256420Z digest=sha256:c22eca9f469509a67614195f7e8c737b0e372319957227f326f188f07bac8781

Observation 377f6f05-7647-4b7b-9270-e986d404555d · outbound

This paper cites AutoDAN-Turbo: A lifelong agent for strategy self-exploration to jailbreak LLMs.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AutoDAN-Turbo: A lifelong agent for strategy self-exploration to jailbreak LLMs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.812145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.259805Z digest=sha256:5cbca67546849aff1337bb4aff53a1c252e3ba97b07e2906383e2df145bbfc08

Observation 7dc85007-6352-4923-b4ec-b9c434b1d8e3 · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AutoDAN: Generating stealthy jailbreak prompts on aligned large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.800603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.263412Z digest=sha256:5dedfb9ee9787c3eb3e133ebcdf27cddfc7c2606582c511ac3297dae6c323409

Observation be4a1f39-a700-40e1-a3a4-3b7bf060fc36 · outbound

This paper cites The tight constant in the Dvoretzky– Kiefer–Wolfowitz inequality.The Annals of Probability, 18(3):1269–1283, 1990.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The tight constant in the Dvoretzky– Kiefer–Wolfowitz inequality.The Annals of Probability, 18(3):1269–1283, 1990

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.788973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.267711Z digest=sha256:08f7f74dceebaae8a1c83d8d6a95d25ade932e0a50178c51449de0fb30ef7c12

Observation 55173b51-9ee4-4f7c-93a6-1d736bb659f6 · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust re- fusal.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks HarmBench: A standardized evaluation framework for automated red teaming and robust re- fusal

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.777446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.271441Z digest=sha256:08b431f285800b7ab1b099ff74fd7fa00c2b8d0a0b7113f999d42637294203b4

Observation 0234a541-4364-4790-a6b3-1e78e02536ea · outbound

This paper cites Tree of attacks: Jailbreaking black-box LLMs automatically.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Tree of attacks: Jailbreaking black-box LLMs automatically

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.275245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.275245Z digest=sha256:92f8473177dd16eab601b560f218c89244a2c40666b5f454ef43e42d019bd59d

Observation 30f0bf98-5fd6-4687-b53c-c1224faa306a · outbound

This paper cites JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.278918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.278918Z digest=sha256:bd0d8f87c1687a890ba25d5d0eed71f79c9ba61f8fd53bdc3d2ceefa29e835c4

Observation d5b8e7e8-33f7-49e8-98d9-ac29d0d6d92d · outbound

This paper cites Rebuff: A self-hardening prompt injection detector.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Rebuff: A self-hardening prompt injection detector

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.758950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.283315Z digest=sha256:ac0d190f44aa2e81685760cb45db6a596983eee7012bde51bc42138cf259fc1e

Observation 003fb30d-91a6-4429-8abd-2ca3371ff7c0 · outbound

This paper cites I can’t believe it’s not robust: Catas- trophic collapse of safety classifiers under embedding drift.arXiv preprint arXiv:2603.01297, 2026.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks I can’t believe it’s not robust: Catas- trophic collapse of safety classifiers under embedding drift.arXiv preprint arXiv:2603.01297, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.286981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.286981Z digest=sha256:a33e1541647d14a50c949ef3a7dcef07ae6202443b7554e721fc05bd0cabbdcf

Observation 3f717bca-7e3f-46ac-8393-3b7154d6af70 · outbound

This paper cites Do Anything Now.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Do Anything Now

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.747306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.290741Z digest=sha256:a57e4a79a2e3b376ee4408425adfeedeb2ca46c6b1fa30ad86899d1f77d2072b

Observation 99c6b68e-5387-4f80-8cf3-a42747243495 · outbound

This paper cites AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:57:41.430736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.294843Z digest=sha256:620fac9612ec627233a5db21628bd134443b422efb12e4e0d61f7a9953dfdaf3

Observation 230cbbd7-198a-457f-943e-17a00cd94097 · outbound

This paper cites A StrongREJECT for empty jailbreaks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A StrongREJECT for empty jailbreaks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.735206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.298913Z digest=sha256:4e7d5ed3db3e3094773ba264b3deb8383b0ec9f01796fcbb74694f2c94e28c8d

Observation 63302013-a04c-44d4-b23e-efff0de17d93 · outbound

This paper cites AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.302502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.302502Z digest=sha256:0b002c0345b4220ca42bf9d02eb7246b11ecf42024530270077609ba6231562a

Observation 877cd1fd-863c-4801-9533-8dec62aaf5c0 · outbound

This paper cites GradSafe: Detecting jailbreak prompts for LLMs via safety-critical gradient analysis.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks GradSafe: Detecting jailbreak prompts for LLMs via safety-critical gradient analysis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.723347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.306733Z digest=sha256:309db06379b71ca9bd85a85b6d28f51772f48d71b57b6017e26d8cbfb822f9a8

Observation 968afd6f-ffa7-4c97-bb32-657975f181b4 · outbound

This paper cites Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:57:41.401067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:57:41.310740Z digest=sha256:8eecbd52a03a977926cfbd5195c1e9d56f6f1ae23a76e382689e7a1764482657

Observation 7f4db920-1fcb-4cce-bbcf-8f0984e851a9 · outbound

This paper cites JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.315199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.315199Z digest=sha256:396211ad34d816decc7b09e8b84b6ec4c001f9329bf8c4a47c43c1a07a075ce9

Observation 77cabf51-e992-43cd-ab3e-84a6343143ba · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Representation Engineering: A Top-Down Approach to AI Transparency

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.319506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.319506Z digest=sha256:b664a0068662168da13cfba980433a2675b7d4cab1f0ee82197b0b059a2bb392

Observation 4c6f2e7b-e024-4d6f-b550-23e3816fc8a9 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-11T13:57:41.323535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.323535Z digest=sha256:4f1db3c0e369bf68e96d659d239547020eacc6e227372f381d335e259b5757bc

Pith citing papers

No inbound Pith citation observations are available.