Pith. sign in

Paper Citation Record · LEDGER

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

As of 13 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.09624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09624 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:57:41.323535Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11f76a94-d370-4d21-be28-6eadc3d8b2c6 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Detecting Language Model Attacks with Perplexity

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.177641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.177641Z digest=sha256:667225cc16ab35842bdcbde159d43fb55c20f57c60fb7ef1302b23f1d139375c

Observation f89b53b5-dca4-4b80-b5ba-2301e2351e20 · outbound

This paper cites A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.182681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.182681Z digest=sha256:2c86feb8f06bbdb7a549ccbd75b3185aeab7298e53456444843def6f79c1d836

Observation 3b83a876-54e8-415d-9f83-f43a69bc846b · outbound

This paper cites Many-shot jailbreaking.NeurIPS, 37:129696–129742, 2024.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Many-shot jailbreaking.NeurIPS, 37:129696–129742, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.925626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.187216Z digest=sha256:036c7d05178759e27997c063d9d72ba0fe9587d124d3cb1f6c39326a09628328

Observation a04cc204-349b-43ec-b09d-ddc3bea168d4 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Refusal in language models is mediated by a single direction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.914862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.191487Z digest=sha256:d35218e574dbaebc26583d3f6c0c4359503b787ba4c038ad22c6c70679a8b009

Observation fcf626c2-ee70-44d3-a918-c569af0fc9f5 · outbound

This paper cites Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.195732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.195732Z digest=sha256:8584471dfc78c39162cf0b679172f016e4604bd5d45ed6b2c762d04a3c786cf8

Observation 0d2f74b0-54f9-4ce0-88f5-f5513b2b2cb5 · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.904148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.200166Z digest=sha256:2d9a855d362b64fba2a63dc0cd1ffa6a9595998149cba00a1c4096c2b10655e6

Observation a4970e7e-fd06-4e0d-bd95-81a509f36856 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.204173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.204173Z digest=sha256:6192a4c4d9d82791bb55dd524d5b96141c28bc0bd924c33778c49e46dad85f6d

Observation 51222e43-59ca-49bb-a4c2-74c796d1d1f5 · outbound

This paper cites When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.208214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.208214Z digest=sha256:e276b904e3c3ca35fe457350583db343618e1a5526c6e7efe1449051ec671975

Observation 3e8eb405-9d1f-4f3a-9839-316a080594e0 · outbound

This paper cites The Llama 3 Herd of Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.212053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.212053Z digest=sha256:84a5e766c338af1a02514c9d683ec2c3173fa13bd5d94866ee0e6d31882417c2

Observation 9802d500-019d-42fb-aa2b-d47c3f598744 · outbound

This paper cites Designing and interpret- ing probes with control tasks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Designing and interpret- ing probes with control tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.892835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.215699Z digest=sha256:ec0caa9b22ec37af6c2fdcbc681e1a3e699f90c0d56bc7dde515adab69593893

Observation 043c4b59-721a-44df-a834-1fed402e9e18 · outbound

This paper cites Best-of-N Jailbreaking.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Best-of-N Jailbreaking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.219248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.219248Z digest=sha256:27b10111a465dcdbf7fc925987d4fca6a45d63d666985ea55b52a9c8660258dd

Observation 4b555bad-6552-4f44-8bb1-b2a7197f9a9f · outbound

This paper cites Attention Tracker: Detecting Prompt Injection Attacks in LLMs.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Attention Tracker: Detecting Prompt Injection Attacks in LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.223332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.223332Z digest=sha256:6038567c74e1d00e47c4770b3caebf71cbae2f702e45089b81a36db06fadfb62

Observation 287afe03-ecfc-4207-9b98-fa687975c303 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.227334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.227334Z digest=sha256:1430ef9f012d76768451eebc1f4148ebd80275e2882f6644cab2f908f629f2f2

Observation 7f6fd031-542e-4f9a-8a8d-d847d0dc8f38 · outbound

This paper cites BeaverTails: Towards im- proved safety alignment of LLM via a human-preference dataset.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks BeaverTails: Towards im- proved safety alignment of LLM via a human-preference dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.881537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.231145Z digest=sha256:7083242da53ba22d15fb1ec8af64102cf9161c231e2a8251bf651bcd40a95787

Observation aea13e5c-18fb-4152-8bf2-93a40c74b98b · outbound

This paper cites WildTeaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks WildTeaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.870448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.234667Z digest=sha256:7036c6be83900069ac7bce1d5218458ecc7b4892345e9f7f0409a7c86ad92b4e

Observation 160cb7af-1975-46cf-9fd4-df20e894231f · outbound

This paper cites HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.238272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.238272Z digest=sha256:99d28bc49c62c66bf9d7bda1b3563ea81a208bfc260c5a9a5ee63f6d72fdd7dd

Observation 576ff26a-77ce-40e3-85ba-3127b401354d · outbound

This paper cites What features in prompts jailbreak LLMs? investigating the mechanisms behind attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks What features in prompts jailbreak LLMs? investigating the mechanisms behind attacks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.242047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.242047Z digest=sha256:70ef36cc63bd27ef102d9d74529fbee3befc85964b8ee7918bf3a091888f5fc8

Observation b88a2ff3-48d9-437b-8f6e-4c6d4157d603 · outbound

This paper cites A simple unified framework for detecting out- of-distribution samples and adversarial attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A simple unified framework for detecting out- of-distribution samples and adversarial attacks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.858768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.245714Z digest=sha256:673fced2aa1086eeb2613bdb9a547303492e736b68c46fe3f33faf696c7a7021

Observation 83c8478d-1212-4d6f-8f7b-1d48cc5fbd13 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The power of scale for parameter-efficient prompt tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.847377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.249320Z digest=sha256:5b711cc8978e7bade8363d0ecc65f1db750f23c14e8f80b405f5ef698bc3bff5

Observation 6dcb5a50-2ca9-45d2-831e-0245c8884244 · outbound

This paper cites Prefix-tuning: Optimiz- ing continuous prompts for generation.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Prefix-tuning: Optimiz- ing continuous prompts for generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.835771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.253003Z digest=sha256:7cdee5654c167fbe8a74c4627f21c34093c646a1c0955c3a443076492794b3d4

Observation 2a46e70c-dcd8-49cf-8743-d965ef4f9f7e · outbound

This paper cites Towards under- standing jailbreak attacks in LLMs: A representation space analysis.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Towards under- standing jailbreak attacks in LLMs: A representation space analysis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.823289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.256420Z digest=sha256:e388372ea455137792df767b063a0fe7c5347b820ec8fd51c4dfe615f9ba385a

Observation 377f6f05-7647-4b7b-9270-e986d404555d · outbound

This paper cites AutoDAN-Turbo: A lifelong agent for strategy self-exploration to jailbreak LLMs.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AutoDAN-Turbo: A lifelong agent for strategy self-exploration to jailbreak LLMs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.812145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.259805Z digest=sha256:792e7ad4b21691a53f681a85d4dba9e4c35ff1669716473906422f3c33de2f30

Observation 7dc85007-6352-4923-b4ec-b9c434b1d8e3 · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AutoDAN: Generating stealthy jailbreak prompts on aligned large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.800603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.263412Z digest=sha256:25cf14bebad8421f0f163fbf357f539281fc5a4f8a3a16e0dfd31cff14efaaa9

Observation be4a1f39-a700-40e1-a3a4-3b7bf060fc36 · outbound

This paper cites The tight constant in the Dvoretzky– Kiefer–Wolfowitz inequality.The Annals of Probability, 18(3):1269–1283, 1990.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The tight constant in the Dvoretzky– Kiefer–Wolfowitz inequality.The Annals of Probability, 18(3):1269–1283, 1990

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.788973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.267711Z digest=sha256:f7178b5a4568be87dba994528c2a619d2513eb65ad6a114d28188788be39c3e9

Observation 55173b51-9ee4-4f7c-93a6-1d736bb659f6 · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust re- fusal.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks HarmBench: A standardized evaluation framework for automated red teaming and robust re- fusal

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.777446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.271441Z digest=sha256:78bc0f0cd5b3f5a08ce3390dd1f74b25f780bfca2f8474d79c7043d5dce7f281

Observation 0234a541-4364-4790-a6b3-1e78e02536ea · outbound

This paper cites Tree of attacks: Jailbreaking black-box LLMs automatically.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Tree of attacks: Jailbreaking black-box LLMs automatically

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.275245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.275245Z digest=sha256:bb2a74c0e781ec831795a431085464a3c00d854b6b95f7f83b658737397c7d9d

Observation 30f0bf98-5fd6-4687-b53c-c1224faa306a · outbound

This paper cites JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.278918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.278918Z digest=sha256:daea9e4cb478ecdaa3997e3b09d26b528c39d1d6ea5901e99f75c2d9a1c50261

Observation d5b8e7e8-33f7-49e8-98d9-ac29d0d6d92d · outbound

This paper cites Rebuff: A self-hardening prompt injection detector.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Rebuff: A self-hardening prompt injection detector

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.758950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.283315Z digest=sha256:d5ca2fffe87dfba37cc8be3fecc81e6c9646e1ce342302ecfb4c349ef94624cc

Observation 003fb30d-91a6-4429-8abd-2ca3371ff7c0 · outbound

This paper cites I can’t believe it’s not robust: Catas- trophic collapse of safety classifiers under embedding drift.arXiv preprint arXiv:2603.01297, 2026.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks I can’t believe it’s not robust: Catas- trophic collapse of safety classifiers under embedding drift.arXiv preprint arXiv:2603.01297, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.286981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.286981Z digest=sha256:38952dc569277a21b3f3f0b4d6090fdbed7dc225eceb832167e7e115ef6d0bd0

Observation 3f717bca-7e3f-46ac-8393-3b7154d6af70 · outbound

This paper cites Do Anything Now.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Do Anything Now

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.747306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.290741Z digest=sha256:3ec7fb8a4f70403fa710f9d6db039cdec4841a85414389210873db16c18d4991

Observation 99c6b68e-5387-4f80-8cf3-a42747243495 · outbound

This paper cites AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:57:41.430736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.294843Z digest=sha256:f3875a2e21c1c94eeabbdecea32314b815723f160e2a931557078d3ff06b7525

Observation 230cbbd7-198a-457f-943e-17a00cd94097 · outbound

This paper cites A StrongREJECT for empty jailbreaks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A StrongREJECT for empty jailbreaks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.735206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.298913Z digest=sha256:c671b15171ced2b6bb1d0ff587911f24a229bb192fbeb007d9badffcd7c9078c

Observation 63302013-a04c-44d4-b23e-efff0de17d93 · outbound

This paper cites AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.302502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.302502Z digest=sha256:cee133ca827761c3519fb1d84c976aee7f00b8bc15809a80bfcfb17189034b15

Observation 877cd1fd-863c-4801-9533-8dec62aaf5c0 · outbound

This paper cites GradSafe: Detecting jailbreak prompts for LLMs via safety-critical gradient analysis.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks GradSafe: Detecting jailbreak prompts for LLMs via safety-critical gradient analysis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.723347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.306733Z digest=sha256:7861276e75985d7a8f65bc7682ac006615f5378ccaabc8d0785cff92b433d668

Observation 968afd6f-ffa7-4c97-bb32-657975f181b4 · outbound

This paper cites Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:57:41.401067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:57:41.310740Z digest=sha256:921cacd5ccfd8915008d8526ce746ec381cad312851bd18a28e244e6bbe72b1b

Observation 7f4db920-1fcb-4cce-bbcf-8f0984e851a9 · outbound

This paper cites JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.315199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.315199Z digest=sha256:71f629d03891c9b446715623f599a935ad95dcaf54d1ff88747f0a4d950b4978

Observation 77cabf51-e992-43cd-ab3e-84a6343143ba · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Representation Engineering: A Top-Down Approach to AI Transparency

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.319506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.319506Z digest=sha256:c045e8c65c8c9c3c4434353302bb68406fb56c668165b1ee9ea65d0429350884

Observation 4c6f2e7b-e024-4d6f-b550-23e3816fc8a9 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-11T13:57:41.323535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.323535Z digest=sha256:fc778c4b5ac43760e9da1413ff4fe3e0a7cc358c4b02ed832aeaa9dcba08e50c

Pith citing papers

No inbound Pith citation observations are available.