Pith. sign in

Paper Citation Record · LEDGER

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2510.20129.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.20129 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T05:22:30.509147Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact20
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 60d19697-c414-403d-8613-47f1d2a45592 · outbound

This paper cites GPT-4 Technical Report.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.588103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:c3b3a1cfd0054e68ce28f807b051f59e1b9da75283c33de867ce6f5775f2fb35

Observation 62a58878-b48d-4562-bf88-5231492a62b3 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Detecting Language Model Attacks with Perplexity

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.601980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:2f21f2d1628bdc67cfea084794c176f5daa2bbf08aee8244329d5ca054a657be

Observation d0fb6236-87cb-456b-af2e-ef2b9a3d95de · outbound

This paper cites Efficient Training of Language Models to Fill in the Middle.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Efficient Training of Language Models to Fill in the Middle

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.581354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:41f11b4365a2f1303eaca65443137fc78a03af57170df022d425b2412ff48ade

Observation b0aa928f-7f2a-492b-87c5-85c78b430438 · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big? InProceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models On the dangers of stochastic parrots: Can language models be too big? InProceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.975897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:bb51276b895bf3be16c110337fc60559422dc672610840483919192ce22e4e07

Observation 80d9fd52-404a-445d-8d57-deb9b2c81485 · outbound

This paper cites Single-pass Detection of Jailbreaking Input in Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Single-pass Detection of Jailbreaking Input in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.598504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:4f409d713b4a2741a23d972fa39faa3803be6996a33bcab47e5bb76eb0d1293e

Observation 8ac98264-c6e6-4e9f-bfb7-50c9b0c50203 · outbound

This paper cites A survey on evaluation of large language models.ACM transactions on intelligent systems and technology, 15(3):1–45, 2024a.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models A survey on evaluation of large language models.ACM transactions on intelligent systems and technology, 15(3):1–45, 2024a

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.969547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:b0f162ca5a939ee18bf0261df5dace4613bce63aac923673407319e3f1559a40

Observation a33a43cc-c162-4524-a220-dcbf48b5f7a0 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.972940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:6606572edeed8af430f178fd530c2ef8ebc2516d26dbc2f57d0ad72b7774724c

Observation bdcdeac0-827d-4bbc-9058-f8fe59b29b4a · outbound

This paper cites Attack Prompt Generation for Red Teaming and Defending Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.594563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:ac0e3ffe69cbdfc12ad1b8c5c8b7465931d7ec40e1f7d83caa6d4596882ec5cb

Observation 46a53fd7-b283-434a-8f45-3a266a9a2acb · outbound

This paper cites Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.605699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:640e950574ee3e76204f6a3e23e0bbe805722df171a667b39dd66a300c2b8481

Observation 5653751b-b50d-4246-aec2-17d7a6739f09 · outbound

This paper cites ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.591261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:2d1d4326da1365406a449a29bd6aad1c82675b1ad1eed11362716ab416615100

Observation 3334fb1b-1f0b-4af7-83ed-f8955a7ebb35 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.584578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:dc5f30774154b4586b549cead3c2b6023a4b26b98c59bce0e1ee5858d4ecd9a6

Observation c39e4aac-3abc-478b-9ea1-51ed769a62ea · outbound

This paper cites Scaling Laws for Neural Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Scaling Laws for Neural Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.608709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:002282b843a0c0864f7ec945633ddd1016f9c7115b9cceb9493417d3571b926d

Observation 692ae75b-197e-4f6e-8e29-e1d2b887ca24 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:37:07.279605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:00b4b0e609307aebc71ed3efdb576334796853fb3172f631e6cabfc367dcf8e6

Observation 3630742f-6a01-4f88-983f-14b53afa37ef · outbound

This paper cites Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T05:25:54.619328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:566d2eaa67e83f8e876f813e391965d5a820cb238dc9d214f12af446c7ab0033

Observation a4f16710-d197-41d0-9f06-c86237e9b564 · outbound

This paper cites The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.623202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:94d987d8f62dc68f5b0cda6ee7b1382f3db223992f887ead645b11cff4d245a8

Observation 7604e27f-5245-40be-ac1c-8aba2983e0c0 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.615726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:f66a416dee5b226948b60e00e05223e5f66f86737ec7fe6f4e19b59f7e5e3d0e

Observation eddb566c-e329-4d43-988d-fec1b861e745 · outbound

This paper cites FlipAttack: Jailbreak LLMs via Flipping.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models FlipAttack: Jailbreak LLMs via Flipping

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:16.351675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:63bebd0e6c92cb0549ffe01036e4a331a79710ea344909413f5f040b047191e3

Observation f005551b-3535-4fbc-8b50-44a2164ce102 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.612403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:ebed93a5954d7dc3565113bd6c4895fe49df41a4dfff441096b54647c05abb8e

Observation 8b9af65f-eef9-4e39-b148-2c7c4b2f4287 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.630924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:7ce1416af27631fdf52f99268fe58d3a61f7d472938200cbbcf3572039d0fccd

Observation f4a23f62-8ce5-41e7-8783-73cd59afbd12 · outbound

This paper cites Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.651445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:1aa91610b81c18ce7dc78afbd87d0ca76228b983b6443829f9cbe368581043f1

Observation fcdc1f0f-0310-4bd3-85ac-c878661eb97d · outbound

This paper cites Aligning Large Language Models via Fine-grained Supervision.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Aligning Large Language Models via Fine-grained Supervision

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.634495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:83e620c9090de1864e2f01d2947331c3aa11cfd19f2bbab6ea3f9eba266834f7

Observation e1dc0f50-33f0-44c8-a642-1a38297e42ae · outbound

This paper cites SQL Injection Jailbreak: A Structural Disaster of Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models SQL Injection Jailbreak: A Structural Disaster of Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.641114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:478f9ceda6eb1ed70e11d2d4520c6d143299e0f7a0ac295e96856474e412ea61

Observation 47c0bf16-1594-4405-a165-94c5d2ae83f1 · outbound

This paper cites A Survey of Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models A Survey of Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.637716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:83c53255542c1d7434af17bed451e7cad8d442e56eac252e34ba1daa0e9cce82

Observation bde3143e-20e0-43e6-a32c-d846de78f1bd · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.644521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:cb5d1e72a3ee02b5b6eb5b5b9269dc05c2828f4e9cd7564991f6a47a233669ee

Observation 26b76b3f-b214-4dfe-876f-9241edfc35ca · outbound

This paper cites comment- ing out.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models comment- ing out

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.981075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:01ff8c99e87afe01d71b1297c691c67b141eab7b62773c8a22638d45e6069620

Observation 0ff281eb-4bc6-480f-8262-4ba0b439132c · outbound

This paper cites Can I”, significantly outperform prefixes that carry a strong personal or instructional intent, like “Can you teach me to do this to others.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Can I”, significantly outperform prefixes that carry a strong personal or instructional intent, like “Can you teach me to do this to others

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.983873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:abb5565051a278020125296221e902647ba84387c2f78d308ce998bd2ae25a5f

Observation e22bf799-5ac6-4635-9ffa-62842cc5e0f9 · outbound

This paper cites role-playing.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models role-playing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.978587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:2e2618e81d3191bd9646abd9596dcf0d677f1787692a936366b46fdc2863bdcb

Pith citing papers

No inbound Pith citation observations are available.