Pith. sign in

Paper Citation Record · LEDGER

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models

As of 13 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2412.18171.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18171 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:02:26.737170Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c677173e-3af0-4e00-9e18-99c5cee50232 · outbound

This paper cites Jailbreak chat, 2023.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Jailbreak chat, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:27.210335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:02:26.602837Z digest=sha256:79754eff570f8249d2889fee7bb2afd454b34243d7fecf52740e0d274abd01d7

Observation 2bdcac9a-1dcc-40fa-a639-35a2edad71c8 · outbound

This paper cites Many-shot jailbreaking.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Many-shot jailbreaking

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:27.193034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:02:26.608044Z digest=sha256:0cdad6f522137b79d9ecd1f9ce65e7b684c437f706ea909e25aa9c7c846158b0

Observation ed4625c8-4b10-4fcf-99c2-8e9cd67c9e9f · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models A General Language Assistant as a Laboratory for Alignment

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.612486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.612486Z digest=sha256:32a79c381d9631a6096411700be3f709660c1ee4bfaaaf5f4991225cfa015a4d

Observation 56408827-6197-47a7-9f04-0525fcbcc1d5 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.617889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.617889Z digest=sha256:2cf1bc5c51cf561ef29f2ef2d61bb2eef02f8087eaac29bc3f2f152acc6d8f8e

Observation ff102511-77bf-4ea2-baf5-28278e09193d · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.622921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.622921Z digest=sha256:96171491e8c0691c5ed71aed2a153e0d5b732bb05f1265e76705a54cb71890f9

Observation e67b2e6f-940d-4cd7-b311-a9fea92fd7ea · outbound

This paper cites Zico Kolter.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Zico Kolter

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.628462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.628462Z digest=sha256:fb20f521f700d35299b463720296bc04f60ee8fab56ebfc929674a4f58b163cb

Observation 8650b5b1-c9da-4ab9-a493-c439f712182c · outbound

This paper cites Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.636480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.636480Z digest=sha256:f89a72d8580f699194bb00223c046428ddfed67a4d41cd8975db3c65203fcfbf

Observation 645c60a9-f761-4276-8439-7c42712e7687 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.641984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.641984Z digest=sha256:c676e25e43c7e79ebd7cc1d85c650e0dcfea9c935e8a14c60bccb8030498b84f

Observation 3d791c1c-8193-4918-8d71-e8b36fb2084b · outbound

This paper cites Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.647309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.647309Z digest=sha256:8d7763b8802f8f85d31ae85ce7ee0a6652e4c800c8999219095a2f6213f0de32

Observation 97e0fba1-631f-4d0a-8258-a37b74dfb516 · outbound

This paper cites In conversation with Artificial Intelligence: aligning language models with human values.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models In conversation with Artificial Intelligence: aligning language models with human values

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.652362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.652362Z digest=sha256:b248fd93449db232a8b811c480c81b9b08a23298dbbabbf215b1ca86e4906423

Observation 1c8e2bfd-f142-4877-9547-c03c44f799ff · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Certifying LLM Safety against Adversarial Prompting

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.657326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.657326Z digest=sha256:8eadbbec61e8b16f38f3fe2a2ec77e4f714cd581ffbd4c8b905c450ada343918

Observation 7b140455-dd06-4910-b29d-8ab790cfb8e2 · outbound

This paper cites Hashimoto.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Hashimoto

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.662444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.662444Z digest=sha256:b0244174d0ae364da73206857e618031399cd98b1c75980dd7f7d9369fa07402

Observation 075521c8-5206-4dc9-8aef-627e16dff426 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.666965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.666965Z digest=sha256:ec5182dec760df82f3d8b2c345499216ac70338501d76f12cdd482518e5d6756

Observation a911f54d-ba12-4e5b-8d85-e511243ac224 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.671653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.671653Z digest=sha256:f06da91c088bbea0abcfc9081d28b8f47c819af32f35810c30c37b113641562d

Observation 9ad325b4-1779-4699-9b1e-c7a1caff0cd2 · outbound

This paper cites GPT-4 Technical Report.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models GPT-4 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.676403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.676403Z digest=sha256:02c15dde1604435b6afbf9d3924c5c342b1b1d2555f82b4876e235cc947b72ee

Observation ee63dac9-75b2-40a1-a3c9-fc2b32be6e06 · outbound

This paper cites an unresolved cited work.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.680836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.680836Z digest=sha256:3923be649f444d8f226d71ed2067cfd16ed10b857d2002146c69b287873958b9

Observation 6a547d90-4968-4506-bec3-884520f865a9 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.685499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.685499Z digest=sha256:8df6ecba5bcdb68f9ad1e83e9afc572fcc422259f13c5a28cc9e2020f17f9ad0

Observation f71874a8-ebff-49e2-9cef-d38c56c5993a · outbound

This paper cites Meta llama guard 2.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Meta llama guard 2

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:27.143996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:02:26.690411Z digest=sha256:85564f818d6eb6f0e841268b25b7d5dd75a917c6a1d5cb5d61c5b7eb2e4d44f4

Observation f3ed272b-8902-4467-8ba7-cbc6f834e443 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.694550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.694550Z digest=sha256:cedea960bc07738c38e7a7227248adc10586bc0bd967cf7acfbcfb9225893017

Observation 24f12a02-e1fd-413c-b2b6-125962a7e667 · outbound

This paper cites On adaptive attacks to adversarial example defenses.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models On adaptive attacks to adversarial example defenses

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:27.127315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:02:26.699213Z digest=sha256:51736925fd2f1616eee10fa020151b225829402aa94d321939edeeb274b20d56

Observation c05bef07-2c6a-456a-982d-140b59fd6193 · outbound

This paper cites The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.704127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.704127Z digest=sha256:6ccaf336d9a2ede4ac6b5a51078480f2f60ac6f7345b3b524c6fdf13e6a9c71c

Observation d013308c-9204-477e-baa9-9634be07d53f · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Jailbroken: How Does LLM Safety Training Fail?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.708733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.708733Z digest=sha256:84b6b18c34ad83fdada18e409b25a8548507646cd82248db389a511e471baa3c

Observation 4cfd5c5e-551a-4de1-8066-5a1b12d2d127 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.713244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.713244Z digest=sha256:9e1a4a731a78f2a35b9bcb3911278539f0fc8eef98c4e936539b5530dcce9e73

Observation ca6d1489-5ced-474a-8e7d-577eed24b3d7 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Defending chatgpt against jailbreak attack via self-reminders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.717626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.717626Z digest=sha256:a4135991dac9bbdc5cbab58d098d91c063539e290df08a9cde02c820f6b1e2ca

Observation 96517cfd-8eb1-44e3-9180-34792e08cc7a · outbound

This paper cites Intention Analysis Makes LLMs A Good Jailbreak Defender.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Intention Analysis Makes LLMs A Good Jailbreak Defender

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.721836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.721836Z digest=sha256:7a2dd5b21984053eaa2e2abfb7dcafdf610f382e7e5787063fa716a23fc17d34

Observation aa3158fe-f488-43ef-95d9-92f6ee752848 · outbound

This paper cites Weak-to-Strong Jailbreaking on Large Language Models.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.726743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.726743Z digest=sha256:36e15689bd0d183272fbb28508d249f7275c2b8b5f63f4bd4f743c42310557e5

Observation da5016f6-9b03-45a2-9b57-6e64ba5be173 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.731586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.731586Z digest=sha256:328420fba863e2880c093af6c4c1b7e7164b2910262b63ddeeeac911c5e07103

Observation 5c61f57e-1e38-46d9-a149-a2537f73272a · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.737170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.737170Z digest=sha256:629d493d6a0373bfb0bbeb9855a4f22107ecf331ab6bfeae33f3d57113648acb

Pith citing papers

No inbound Pith citation observations are available.