Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T13:53:10.999594Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 4 inbound Pith citation observations for arXiv:2501.18626.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T13:53:10.999594Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:12:27.776350Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T23:32:15.547339Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3b069cd7-e081-4bec-ac55-eebcb87a4939 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Does Refusal Training in LLMs Generalize to the Past Tense?
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23002035-bef3-4797-a923-04f22aee1eba · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 042f2853-3088-45a8-aaac-7252d334bd5c · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 882387c2-eda5-4729-8a8c-91dfb740634a · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1bd78ad-d57e-4e45-a751-bbc4dda800c5 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Building Guardrails for Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8989374-b901-4edf-9035-23c83495dff9 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Safeguarding Large Language Models: A Survey
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc78a156-1abb-4934-9031-be723e29ce18 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 647a6428-80be-4819-9eba-3a2f173ecdf9 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Gemma 2: Improving Open Language Models at a Practical Size
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f48098f6-712a-4a75-a291-d2f115c90b7f · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a19c080-0f53-46d8-8cf9-fca5c74bcbba · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 597a9711-fba7-4108-862f-9f757296b2f3 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Single Character Perturbations Break LLM Alignment
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 07d0002d-e631-41e8-9474-aedc25cb55ba · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation db1a50b7-41b6-444a-b628-0b4e09d9dd28 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 414db0aa-2e9b-45ac-8cac-5ad4a6b01815 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2843845-7d5c-4311-9773-a9786f78c2f6 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5e7e662b-aebf-460c-8456-7850ac22fd68 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation edb41915-0583-433b-8db2-4aaf4f1f779d · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4550525d-c4cf-4bd7-b094-71c336fe6311 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1f5f8411-04e3-4d97-a8da-aa007534dd25 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d996eb7a-c653-4074-8f10-2eb43f381f03 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bd221061-12f2-4cfa-978a-c44a7c9e5a80 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39032472-e589-4a3a-87ad-4d434a19a32b · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs do anything now
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 554f58e4-b208-4242-b7ae-3201d02cd7b0 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 01e275ba-f362-444f-95ad-03271447018e · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c7a1fa-7a82-42ee-ab41-5dd45ea8bcab · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75fa16ee-f3bf-4d7b-8c33-0918fe9d9b6a · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Jailbroken: How Does LLM Safety Training Fail?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e865d0b-8f1c-458d-8295-91538a363dbb · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d93a15bb-7882-4f30-bd4f-ef0b857eaebe · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs An LLM can Fool Itself: A Prompt-Based Adversarial Attack
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 110524e7-28a8-4eb9-a44d-05c89944aca7 · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a8d6227-1051-4734-8e63-64b05a5429fa · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs online" 'onlinestring :=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e52257d-8ebf-4cb3-a449-7f97d3fdcd2a · outbound
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs write newline
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 642b46d7-505c-43bd-b4f6-d10351af91ce · inbound
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2661a4dc-7c8d-4897-a17c-5d5a869f2b21 · inbound
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a1439e-624a-4b3c-80d5-9c78bf3d68ed · inbound
POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad93d65a-7726-4ec3-9e37-9d3dc4335df1 · inbound
LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.