Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:02:26.737170Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2412.18171.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:02:26.737170Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c677173e-3af0-4e00-9e18-99c5cee50232 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Jailbreak chat, 2023
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2bdcac9a-1dcc-40fa-a639-35a2edad71c8 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Many-shot jailbreaking
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ed4625c8-4b10-4fcf-99c2-8e9cd67c9e9f · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models A General Language Assistant as a Laboratory for Alignment
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56408827-6197-47a7-9f04-0525fcbcc1d5 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff102511-77bf-4ea2-baf5-28278e09193d · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e67b2e6f-940d-4cd7-b311-a9fea92fd7ea · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Zico Kolter
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8650b5b1-c9da-4ab9-a493-c439f712182c · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645c60a9-f761-4276-8439-7c42712e7687 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d791c1c-8193-4918-8d71-e8b36fb2084b · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e0fba1-631f-4d0a-8258-a37b74dfb516 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models In conversation with Artificial Intelligence: aligning language models with human values
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8e2bfd-f142-4877-9547-c03c44f799ff · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Certifying LLM Safety against Adversarial Prompting
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b140455-dd06-4910-b29d-8ab790cfb8e2 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Hashimoto
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 075521c8-5206-4dc9-8aef-627e16dff426 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a911f54d-ba12-4e5b-8d85-e511243ac224 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ad325b4-1779-4699-9b1e-c7a1caff0cd2 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models GPT-4 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee63dac9-75b2-40a1-a3c9-fc2b32be6e06 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a547d90-4968-4506-bec3-884520f865a9 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f71874a8-ebff-49e2-9cef-d38c56c5993a · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Meta llama guard 2
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f3ed272b-8902-4467-8ba7-cbc6f834e443 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f12a02-e1fd-413c-b2b6-125962a7e667 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models On adaptive attacks to adversarial example defenses
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c05bef07-2c6a-456a-982d-140b59fd6193 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d013308c-9204-477e-baa9-9634be07d53f · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Jailbroken: How Does LLM Safety Training Fail?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cfd5c5e-551a-4de1-8066-5a1b12d2d127 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca6d1489-5ced-474a-8e7d-577eed24b3d7 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Defending chatgpt against jailbreak attack via self-reminders
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96517cfd-8eb1-44e3-9180-34792e08cc7a · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Intention Analysis Makes LLMs A Good Jailbreak Defender
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa3158fe-f488-43ef-95d9-92f6ee752848 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Weak-to-Strong Jailbreaking on Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da5016f6-9b03-45a2-9b57-6e64ba5be173 · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c61f57e-1e38-46d9-a149-a2537f73272a · outbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.