Pith. sign in

Paper Citation Record · LEDGER

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

As of 13 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.09624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09624 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:57:41.323535Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11f76a94-d370-4d21-be28-6eadc3d8b2c6 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Detecting Language Model Attacks with Perplexity

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.177641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.177641Z digest=sha256:7ad17f75baefc418d55f7ec2813a21f900407181009c7b7d8302ecfbe9611f00

Observation f89b53b5-dca4-4b80-b5ba-2301e2351e20 · outbound

This paper cites A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.182681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.182681Z digest=sha256:ca2340133e70b392526483535e810f10823c6510af3e2ba5ba3386612daf1961

Observation 3b83a876-54e8-415d-9f83-f43a69bc846b · outbound

This paper cites Many-shot jailbreaking.NeurIPS, 37:129696–129742, 2024.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Many-shot jailbreaking.NeurIPS, 37:129696–129742, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.925626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.187216Z digest=sha256:87ffcda76a8c346be81b7ecd3c9f1f0c0a962c1c99c59419770813f997fd71f4

Observation a04cc204-349b-43ec-b09d-ddc3bea168d4 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Refusal in language models is mediated by a single direction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.914862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.191487Z digest=sha256:b735c487a16f4a588106005d5f1a4bd266887f24ba549e2f5c3c2c15938366a5

Observation fcf626c2-ee70-44d3-a918-c569af0fc9f5 · outbound

This paper cites Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.195732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.195732Z digest=sha256:71b24a81bcf5d219370765c98f140e3c22440d74e102fbb7d1c49cd6c582a701

Observation 0d2f74b0-54f9-4ce0-88f5-f5513b2b2cb5 · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.904148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.200166Z digest=sha256:31e8d2e8e9ee4c8a6fe83bac91894fcb0e86378018b155854b08a531bab5ba8e

Observation a4970e7e-fd06-4e0d-bd95-81a509f36856 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.204173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.204173Z digest=sha256:dbfd93375c9142d6c1557417118285593a3e58d6fdf11ed9c345bedbe906a3c8

Observation 51222e43-59ca-49bb-a4c2-74c796d1d1f5 · outbound

This paper cites When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.208214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.208214Z digest=sha256:0ba0e428e3e25bcdb8aedb1d37249e9d53719f501302fc00294dca85c2d3822d

Observation 3e8eb405-9d1f-4f3a-9839-316a080594e0 · outbound

This paper cites The Llama 3 Herd of Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.212053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.212053Z digest=sha256:a142d85be7d6093289a1286c4b3773d3a3da08e045095419f3bc6d7348d6b725

Observation 9802d500-019d-42fb-aa2b-d47c3f598744 · outbound

This paper cites Designing and interpret- ing probes with control tasks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Designing and interpret- ing probes with control tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.892835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.215699Z digest=sha256:f63b2aae2dd06f0bac9695a3296213388db5de1fadf1031a175717ac6ab69926

Observation 043c4b59-721a-44df-a834-1fed402e9e18 · outbound

This paper cites Best-of-N Jailbreaking.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Best-of-N Jailbreaking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.219248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.219248Z digest=sha256:aa475c17947f113fa78cca98c1e2b2032126f1184f31de45c4348bd14ba0bc66

Observation 4b555bad-6552-4f44-8bb1-b2a7197f9a9f · outbound

This paper cites Attention Tracker: Detecting Prompt Injection Attacks in LLMs.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Attention Tracker: Detecting Prompt Injection Attacks in LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.223332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.223332Z digest=sha256:c89907feb7f411234842958764fa203580d79a7d987b7e32655379eb27485880

Observation 287afe03-ecfc-4207-9b98-fa687975c303 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.227334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.227334Z digest=sha256:af137f6dd77476bfc55da7a36a9fb2124021c9fbe1b2baa1969e4f41d4e869e8

Observation 7f6fd031-542e-4f9a-8a8d-d847d0dc8f38 · outbound

This paper cites BeaverTails: Towards im- proved safety alignment of LLM via a human-preference dataset.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks BeaverTails: Towards im- proved safety alignment of LLM via a human-preference dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.881537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.231145Z digest=sha256:5061de155ce297284b470649aa4720508dbb7a2bd4387714d880eb6adb5a6e0a

Observation aea13e5c-18fb-4152-8bf2-93a40c74b98b · outbound

This paper cites WildTeaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks WildTeaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.870448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.234667Z digest=sha256:fe1b08030ae3f15941f4c4f5a4dbfcd0c474e6570072000b3bc8e6031929878f

Observation 160cb7af-1975-46cf-9fd4-df20e894231f · outbound

This paper cites HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.238272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.238272Z digest=sha256:c78150ee7ca93b4edc8539d967521f6e6bb73cdab72b146bfccc1fe3524f4229

Observation 576ff26a-77ce-40e3-85ba-3127b401354d · outbound

This paper cites What features in prompts jailbreak LLMs? investigating the mechanisms behind attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks What features in prompts jailbreak LLMs? investigating the mechanisms behind attacks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.242047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.242047Z digest=sha256:faad6fe7506bbd488d18f05ed974ad564881528df8fd4035e5ba7e6b9be997b4

Observation b88a2ff3-48d9-437b-8f6e-4c6d4157d603 · outbound

This paper cites A simple unified framework for detecting out- of-distribution samples and adversarial attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A simple unified framework for detecting out- of-distribution samples and adversarial attacks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.858768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.245714Z digest=sha256:d196f74e3c366274dba7d96b797501e78e55186b490d71133d3679370e5e0b16

Observation 83c8478d-1212-4d6f-8f7b-1d48cc5fbd13 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The power of scale for parameter-efficient prompt tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.847377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.249320Z digest=sha256:eb48d85002eecfcc923e48b03f2a57655014e25e14a8481113424efc095533d2

Observation 6dcb5a50-2ca9-45d2-831e-0245c8884244 · outbound

This paper cites Prefix-tuning: Optimiz- ing continuous prompts for generation.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Prefix-tuning: Optimiz- ing continuous prompts for generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.835771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.253003Z digest=sha256:017c4ae90e6f2577feac78e0ba829d688acd7d878664900a87d83134d5944e6e

Observation 2a46e70c-dcd8-49cf-8743-d965ef4f9f7e · outbound

This paper cites Towards under- standing jailbreak attacks in LLMs: A representation space analysis.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Towards under- standing jailbreak attacks in LLMs: A representation space analysis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.823289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.256420Z digest=sha256:b632fd10efca4ff348690ee21f7b0d55d323570ea04ffe83302559d7acaeccc4

Observation 377f6f05-7647-4b7b-9270-e986d404555d · outbound

This paper cites AutoDAN-Turbo: A lifelong agent for strategy self-exploration to jailbreak LLMs.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AutoDAN-Turbo: A lifelong agent for strategy self-exploration to jailbreak LLMs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.812145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.259805Z digest=sha256:141bcd3c91abbdced57ae697c52d03074b53521e554c38a84e17066d3ddd4186

Observation 7dc85007-6352-4923-b4ec-b9c434b1d8e3 · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AutoDAN: Generating stealthy jailbreak prompts on aligned large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.800603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.263412Z digest=sha256:d5deabc911b2b0996cea1bf9a2a626d503bb89cb83aed352dcdcc97f3b4ed526

Observation be4a1f39-a700-40e1-a3a4-3b7bf060fc36 · outbound

This paper cites The tight constant in the Dvoretzky– Kiefer–Wolfowitz inequality.The Annals of Probability, 18(3):1269–1283, 1990.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The tight constant in the Dvoretzky– Kiefer–Wolfowitz inequality.The Annals of Probability, 18(3):1269–1283, 1990

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.788973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.267711Z digest=sha256:78a8f19ba7f1e50834b1dc211aab5ee18f90f93be61df60d6af8f43b4878113a

Observation 55173b51-9ee4-4f7c-93a6-1d736bb659f6 · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust re- fusal.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks HarmBench: A standardized evaluation framework for automated red teaming and robust re- fusal

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.777446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.271441Z digest=sha256:e7d03e037f139f381aca23134ddc828cc2ce4beb6a6b47ba9ad8e96d79714f28

Observation 0234a541-4364-4790-a6b3-1e78e02536ea · outbound

This paper cites Tree of attacks: Jailbreaking black-box LLMs automatically.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Tree of attacks: Jailbreaking black-box LLMs automatically

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.275245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.275245Z digest=sha256:704e8ae54a271eab9eda96858b84d9300eed99929c55cfd605882636ea470aea

Observation 30f0bf98-5fd6-4687-b53c-c1224faa306a · outbound

This paper cites JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.278918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.278918Z digest=sha256:0e843a2b1dcfe6f5edc25233b015c5d3030d394b565081aaab179c119c2e61ed

Observation d5b8e7e8-33f7-49e8-98d9-ac29d0d6d92d · outbound

This paper cites Rebuff: A self-hardening prompt injection detector.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Rebuff: A self-hardening prompt injection detector

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.758950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.283315Z digest=sha256:9d77174c454024d1b390ab19ed51627a91668bd0c71e52767b920a2af3847cf9

Observation 003fb30d-91a6-4429-8abd-2ca3371ff7c0 · outbound

This paper cites I can’t believe it’s not robust: Catas- trophic collapse of safety classifiers under embedding drift.arXiv preprint arXiv:2603.01297, 2026.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks I can’t believe it’s not robust: Catas- trophic collapse of safety classifiers under embedding drift.arXiv preprint arXiv:2603.01297, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.286981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.286981Z digest=sha256:f6af2b22855cabb8fb964c19ded494aa447b7e7c19b88f62bb9dd68ba93b70df

Observation 3f717bca-7e3f-46ac-8393-3b7154d6af70 · outbound

This paper cites Do Anything Now.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Do Anything Now

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.747306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.290741Z digest=sha256:a180053efee910c596ad768e3d17bf1867c4d79fff8e812c400829ceae58aa2b

Observation 99c6b68e-5387-4f80-8cf3-a42747243495 · outbound

This paper cites AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:57:41.430736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.294843Z digest=sha256:d025af0b6158f8d3b35748d233b52f05867cb31b0b5e9ca937880ccf58983690

Observation 230cbbd7-198a-457f-943e-17a00cd94097 · outbound

This paper cites A StrongREJECT for empty jailbreaks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A StrongREJECT for empty jailbreaks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.735206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.298913Z digest=sha256:d2900ca7731c6f73ad364837df481e9dd4cfb7a9e1c585169c058b1f354dfea9

Observation 63302013-a04c-44d4-b23e-efff0de17d93 · outbound

This paper cites AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.302502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.302502Z digest=sha256:5fddbc575e853c529e0ab7a4aeae5e8f346325eaa3dd32a97dff417455484495

Observation 877cd1fd-863c-4801-9533-8dec62aaf5c0 · outbound

This paper cites GradSafe: Detecting jailbreak prompts for LLMs via safety-critical gradient analysis.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks GradSafe: Detecting jailbreak prompts for LLMs via safety-critical gradient analysis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.723347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.306733Z digest=sha256:1e59a6796dd24011b008f664d512a953f23e6826b27b57d877af3834414a5860

Observation 968afd6f-ffa7-4c97-bb32-657975f181b4 · outbound

This paper cites Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:57:41.401067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:57:41.310740Z digest=sha256:99206e8c15a4bebf458e83c3e8d4d7ca2dc835e1825b377c864cdcba29e84cd6

Observation 7f4db920-1fcb-4cce-bbcf-8f0984e851a9 · outbound

This paper cites JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.315199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.315199Z digest=sha256:fa95b59f4a11a01a996020699fcb14a0d6fcc02cb8d43e2c908a6fb8a86bf227

Observation 77cabf51-e992-43cd-ab3e-84a6343143ba · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Representation Engineering: A Top-Down Approach to AI Transparency

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.319506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.319506Z digest=sha256:d3249c9eb9f3a76e67e0c4d255e9eb8e91bfa24068b345405f37c43c94e6f142

Observation 4c6f2e7b-e024-4d6f-b550-23e3816fc8a9 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-11T13:57:41.323535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.323535Z digest=sha256:12f080d61b374479769f728fb5f47f2d3c06b99ffb70cfd05a8dfadcefd6272c

Pith citing papers

No inbound Pith citation observations are available.