Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T05:22:30.509147Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2510.20129.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T05:22:30.509147Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 60d19697-c414-403d-8613-47f1d2a45592 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62a58878-b48d-4562-bf88-5231492a62b3 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Detecting Language Model Attacks with Perplexity
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0fb6236-87cb-456b-af2e-ef2b9a3d95de · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Efficient Training of Language Models to Fill in the Middle
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0aa928f-7f2a-492b-87c5-85c78b430438 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models On the dangers of stochastic parrots: Can language models be too big? InProceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80d9fd52-404a-445d-8d57-deb9b2c81485 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Single-pass Detection of Jailbreaking Input in Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ac98264-c6e6-4e9f-bfb7-50c9b0c50203 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models A survey on evaluation of large language models.ACM transactions on intelligent systems and technology, 15(3):1–45, 2024a
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a33a43cc-c162-4524-a220-dcbf48b5f7a0 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bdcdeac0-827d-4bbc-9058-f8fe59b29b4a · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Attack Prompt Generation for Red Teaming and Defending Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46a53fd7-b283-434a-8f45-3a266a9a2acb · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5653751b-b50d-4246-aec2-17d7a6739f09 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3334fb1b-1f0b-4af7-83ed-f8955a7ebb35 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c39e4aac-3abc-478b-9ea1-51ed769a62ea · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Scaling Laws for Neural Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 692ae75b-197e-4f6e-8e29-e1d2b887ca24 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3630742f-6a01-4f88-983f-14b53afa37ef · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4f16710-d197-41d0-9f06-c86237e9b564 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7604e27f-5245-40be-ac1c-8aba2983e0c0 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eddb566c-e329-4d43-988d-fec1b861e745 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models FlipAttack: Jailbreak LLMs via Flipping
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f005551b-3535-4fbc-8b50-44a2164ce102 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b9af65f-eef9-4e39-b148-2c7c4b2f4287 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4a23f62-8ce5-41e7-8783-73cd59afbd12 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fcdc1f0f-0310-4bd3-85ac-c878661eb97d · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Aligning Large Language Models via Fine-grained Supervision
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1dc0f50-33f0-44c8-a642-1a38297e42ae · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models SQL Injection Jailbreak: A Structural Disaster of Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47c0bf16-1594-4405-a165-94c5d2ae83f1 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models A Survey of Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bde3143e-20e0-43e6-a32c-d846de78f1bd · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26b76b3f-b214-4dfe-876f-9241edfc35ca · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models comment- ing out
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ff281eb-4bc6-480f-8262-4ba0b439132c · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Can I”, significantly outperform prefixes that carry a strong personal or instructional intent, like “Can you teach me to do this to others
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e22bf799-5ac6-4635-9ffa-62842cc5e0f9 · outbound
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models role-playing
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.