Pith. sign in

Paper Citation Record · LEDGER

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

As of 9 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2607.24392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24392 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T16:06:06.687929Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b310d6db-cb48-4ee2-bb07-75581854612d · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:03.615181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:03.615181Z digest=sha256:14cfd104887bbc399701eca27489be53c56d12be1601f241ab6f10c937981034

Observation a0f0e8e9-48cd-486a-8129-9dd083cabf0f · outbound

This paper cites GPT-4 Technical Report.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:03.689934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:03.689934Z digest=sha256:a56f81bb41904ec2c149a1e7a5fcedc2bc4a6a8bdcf2a5515953bfa86b4dbd8c

Observation b509488a-218b-4b99-83aa-3e9b2ac83105 · outbound

This paper cites an unresolved cited work.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:03.783409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:03.783409Z digest=sha256:986732c3470969873fdd80771a877f829a55963b6ebb6357a12264f9d5bab745

Observation f41706c2-33b6-48e5-86df-aece5f6d7ede · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Detecting Language Model Attacks with Perplexity

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:03.882350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:03.882350Z digest=sha256:9974e461fce5383d4199328317062a4a073b199994e8b255be352f12dbab862f

Observation aeeedf14-617a-4aed-8c71-cef91edcac68 · outbound

This paper cites Language Models are Few-Shot Learners.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:04.089010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:04.089010Z digest=sha256:59cc55cd0e77687da87374acc43f0d8a57d3f041bc3135626b8940b77a003d62

Observation 253a77cf-dfb4-4256-b0df-ebfab08a46fc · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:04.193653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:04.193653Z digest=sha256:9a11d096fa1f52dd1ea57618acea84ff77d6efc28f11b75382e978984b16ed49

Observation 75a6e83e-7179-488d-8857-cf1a5ded795f · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:04.281114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:04.281114Z digest=sha256:4c0fec863fab83ac688712fc5ca93c6a1c28de57849c30ee13265c91953d0ea9

Observation 15352712-e2f2-4e54-9f8d-cb9f642e26d9 · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:04.364227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:04.364227Z digest=sha256:72b7bfb2dcf4faf33644d1ea61aa71f3fd6e38cb02590cb53f8a378df60e47b4

Observation 7f055f78-ed1d-47af-9e52-cff5c7013f2b · outbound

This paper cites an unresolved cited work.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:04.511164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:04.511164Z digest=sha256:b265910bf358fcfb6ca2f5362bf442203d99fe1b7de3cd310df16ac01bad8350

Observation 88a8a4ce-aa82-44e0-91ed-98d7ce5e593c · outbound

This paper cites TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:04.707133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:04.707133Z digest=sha256:eb4b11a8593aafc7d1961542bbc15b78087a89b98221da1199f70347b6de658c

Observation f69015bc-e1bd-4c3f-a985-fdb2956ef072 · outbound

This paper cites The robots are coming: Exploring the implications of openai codex on introductory program- ming.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs The robots are coming: Exploring the implications of openai codex on introductory program- ming

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:04.784376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:04.784376Z digest=sha256:2d08e643252efc87284ada0985195ee2434effb95b0031b4ba188389509499a3

Observation ff5872d1-564f-43b3-995c-f00c0877dac7 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:04.876439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:04.876439Z digest=sha256:87ba3d97affde2c58fd7f1cf8d6394df229a169c3912f08a17ff77bbc49bea99

Observation 0dab7cd4-8336-406d-9b0c-f54cbe7c67a6 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:04.944283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:04.944283Z digest=sha256:73cad62a2b2b81492afc58de961f571550ef43108362ca436810d2bbc8d00c93

Observation 5c447e52-6f8a-43f2-878c-c4f64c0f9dea · outbound

This paper cites Mistral 7B.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.004532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.004532Z digest=sha256:c78687bd11c1c2d4434db3e1cf1ad27dae08fa1adf955964c9c803eaa9bc7e8c

Observation 0accd163-a524-4631-a378-50070b71ae7e · outbound

This paper cites Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Processing Systems, 36, 2024.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Processing Systems, 36, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.089361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.089361Z digest=sha256:b32a3849c054b957c1f5ef254801e48df2a96c0ebe68f618f8158ed99dce0128

Observation b1887b75-0793-4ee5-b58c-1ef94547d7d0 · outbound

This paper cites Salad-bench: A hierarchical and comprehensive safety benchmark for large language models.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Salad-bench: A hierarchical and comprehensive safety benchmark for large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.156920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.156920Z digest=sha256:ffd60ea062f0c73d48381a5d57caa770e0f21f97bce971c2b2c9fcf53d51e1df

Observation 195d7399-377a-4024-b7bc-786eb3d8fdc8 · outbound

This paper cites Counting function estimates for coherent frames and Riesz sequences.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Counting function estimates for coherent frames and Riesz sequences

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.219804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.219804Z digest=sha256:531bc0e2c45cd82cab95ac0aef59d4beb815e7727bc1c953c3260c8b67d91f26

Observation 1486b3ba-5640-47a6-b238-8ae4f0b015bf · outbound

This paper cites Llm self defense: By self examination, llms know they are being tricked,.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Llm self defense: By self examination, llms know they are being tricked,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.279438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.279438Z digest=sha256:ec20c3dfe593b4208a59c8ce65c4af129ff5a91ba800f2d3f6c8421fb88ac5bc

Observation b33c1953-7071-425f-b150-2142b1e69023 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.405942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.405942Z digest=sha256:e8e9beac89befe9c846831b1f7e0e7dbee43e42516237421258562badf7140bc

Observation 0789fb59-fc4e-4340-8c58-463e14c888d5 · outbound

This paper cites LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.342582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.342582Z digest=sha256:a752bd4894eb5a8b6ac19d7e365dc160e56efe25d16861d5a14160455a683812

Observation 98057fa9-a562-4330-af10-eeafa3d03830 · outbound

This paper cites Trustllm: Trustworthiness in large language models.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Trustllm: Trustworthiness in large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.528177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.528177Z digest=sha256:d1c81f8a2161b139c6f09c3b5dcf0afe499297997fa88914155134d2e5520b04

Observation 65dc2ad8-1853-470a-88f6-8d4714f294c0 · outbound

This paper cites XSTest: A test suite for identifying exaggerated safety behaviours in large language models.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs XSTest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.465819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.465819Z digest=sha256:6ce5da258777e562e20aab046f0ad7f1730ae12c4985ef41ebb4f3a8dd930db9

Observation f9639ddd-ddfc-4704-bd7c-72e117b06068 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Gemma 2: Improving Open Language Models at a Practical Size

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.675062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.675062Z digest=sha256:dcc67c61235e68184c692086c5320a9939451a35abf20e21dc4cc0703387efb9

Observation a8cb5949-4aea-4280-9c85-af3750d94922 · outbound

This paper cites LLM4Decompile: Decompiling Binary Code with Large Language Models.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs LLM4Decompile: Decompiling Binary Code with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.601499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.601499Z digest=sha256:ff3d67f53376b7293f396b12e4257123c1f7e2bd6f1fdc635841cfc6e98fcb17

Observation 78b46bff-5621-4fd2-83c0-15d8f16652c3 · outbound

This paper cites Defending LLMs against jailbreaking attacks via backtranslation.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Defending LLMs against jailbreaking attacks via backtranslation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.805954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.805954Z digest=sha256:9f48680d0aa70c267bab55a92ac417d34fcdbab7c6a92d0dd956249242527921

Observation d998395c-8f7a-48c4-b26f-c5c87662bce9 · outbound

This paper cites The art of defending: A systematic evaluation and analysis of LLM defense strategies on safety and over-defensiveness.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs The art of defending: A systematic evaluation and analysis of LLM defense strategies on safety and over-defensiveness

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.742114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.742114Z digest=sha256:1b8f80447525798f46092477b81940277539548252492b99f6a84cb8f86f6d3e

Observation 7a179e9a-bb3b-445a-b683-f7a1f9e23786 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.960205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.960205Z digest=sha256:00deb527849d523f4d616ec523677ff8b1cbce01f8143bbca1dae806cd5775ad

Observation 967b658a-0064-47d9-b9c8-f210621a010e · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.868142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.868142Z digest=sha256:7ca751bd5ab5950ecf91209e6c8ca214f5dd89acf7154dbce1d12687184234eb

Observation 015aee69-d2e8-49ad-ba73-34c93b04abc1 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:06.082878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:06.082878Z digest=sha256:41d7876a031e0638eb0cc98e597ec45bcb17a692dbc2672e6cc33142bf52f982

Observation 0df3be41-3ad5-47e0-9b08-a214089f302f · outbound

This paper cites LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:06.031150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:06.031150Z digest=sha256:a818792f4ae0c77944c99cf2ebb5953ce1780426f950c1e74984ea34c314887a

Observation dcd3b31c-d92f-45fb-89ce-84553e2d4239 · outbound

This paper cites SafeDecoding: Defending against jailbreak attacks via safety-aware decoding.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs SafeDecoding: Defending against jailbreak attacks via safety-aware decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:06.239596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:06.239596Z digest=sha256:dfb9db32fff0a38b147d93e023cf1f2ad30c5372125e0bd01273debdc91024d6

Observation 7c08e564-bb40-4baa-ada9-8baa55f1d190 · outbound

This paper cites Contrastive preference optimization: Pushing the bound- aries of LLM performance in machine translation.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Contrastive preference optimization: Pushing the bound- aries of LLM performance in machine translation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:06.178608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:06.178608Z digest=sha256:6e1ac09369c7eda5fb42f1ac52ab25423cf84f74e23592d4fbf8c3cc140141db

Observation 15a06c57-1f00-44ac-ad24-2746479fd2ae · outbound

This paper cites Qwen2 Technical Report.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Qwen2 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:06.390610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:06.390610Z digest=sha256:f6ed54d5a21d574ce35f19c6403e1a7405b71f750e775ac08f8f9f069f433276

Observation c66ab8a0-923c-4170-922f-8fe067ecd526 · outbound

This paper cites A comprehensive study of jailbreak attack versus defense for large language models.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs A comprehensive study of jailbreak attack versus defense for large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:06.323695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:06.323695Z digest=sha256:24072a403af2fc75b5008ffb264bb146510982458000dee62e58c3c4916474a3

Observation eb2b5ed2-554b-4b60-913f-271f13795738 · outbound

This paper cites Intention Analysis Makes LLMs A Good Jailbreak Defender.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Intention Analysis Makes LLMs A Good Jailbreak Defender

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:06.522729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:06.522729Z digest=sha256:974c1bb92f4c94293cd370893044d9d40794179f45b847d61b2da0b3113f6b35

Observation d1add4be-80e4-4697-adb3-8a270c806bb6 · outbound

This paper cites Jailbreak open-sourced large language models via enforced decoding.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Jailbreak open-sourced large language models via enforced decoding

Reference 36

Resolution
verified exact
doi, observed 2026-07-31T16:11:53.860160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-31T16:06:06.459700Z digest=sha256:8e6120df89e35b7ffb7d23dc931c00b11d50d35270c4a1386d1e07cae8dfd61b

Observation 43f900a8-bad3-4961-a189-78e99f0bcfd0 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Instruction-Following Evaluation for Large Language Models

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-07-31T16:06:06.687929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:06.687929Z digest=sha256:b2f3fb3fa061e1de17ad4f93f5962c8d6aa48514c01660adc4a3dc1cdda9d131

Observation 5465ff8f-bbda-494c-8e48-909797b2a1fe · outbound

This paper cites Defending large language models against jailbreaking attacks through goal prioritization.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Defending large language models against jailbreaking attacks through goal prioritization

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-07-31T16:06:06.619362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:06.619362Z digest=sha256:c4c0ca49e4cd7d79eff816aa0c7acefb3794ca70c088558d89829e19a784fb4a

Observation e375fec4-ae95-4a0f-9514-9a211127c156 · outbound

This paper cites The Llama 3 Herd of Models.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:04.622310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:04.622310Z digest=sha256:253768cf4884d7053d2ae0ed451d22f8e5e27c7f1b18a1a3a30b97f8df55891c

Pith citing papers

No inbound Pith citation observations are available.