Pith. sign in

Paper Citation Record · LEDGER

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.02665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02665 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:48:49.845907Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73d5e304-eaf1-4e55-adbf-64fc2f51e490 · outbound

This paper cites Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:47.691899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:47.691899Z digest=sha256:4ab09c7f3d4006699d3b328e7542368a5452c328e8ee14e928a0999f0aedbc45

Observation 60ade032-b860-4a8c-bffc-634a86ebd9d8 · outbound

This paper cites Pappas, Florian Tramer, Hamed Hassani, and Eric Wong.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Pappas, Florian Tramer, Hamed Hassani, and Eric Wong

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:48:53.100173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:47.816131Z digest=sha256:cfb115f27169d255a60276b125288d4ec31550778849d0f106019c60f35aabbf

Observation 58f0fcc2-0fb7-4eac-9743-37fcf82efc48 · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:47.873645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:47.873645Z digest=sha256:25563a00bf54f1f0e26b384978c7a75110e480806fb5e23611d1233267e327df

Observation 5dc7053d-6fff-4f28-a9c3-00cb61985a9a · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.868473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:47.964952Z digest=sha256:4b41d9aff3ff3be468d40885d278cb8a839f9ddff844ff8bdb6a96557b4ba1dd

Observation fe14dc20-04de-4e94-9e16-600890e0727c · outbound

This paper cites Best-of-N Jailbreaking.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Best-of-N Jailbreaking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:48.102460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:48.102460Z digest=sha256:86fbed1693c27284550495b575e6220cbab68253c8a1e0eff9ff50437bcbd107

Observation 6a825f54-2d85-4f97-bced-821b5569187d · outbound

This paper cites Jacobs and Hanna Wallach.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Jacobs and Hanna Wallach

Reference 6

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-05T00:48:50.695454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.228849Z digest=sha256:69ffdc4a491eba373dabb7069595942e3c794882fa0bd0a61bb8dfc08a4cb416

Observation 784b503d-3cbf-4283-86cc-58e181a40db7 · outbound

This paper cites A Cross-Language Investigation into Jailbreak Attacks in Large Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity A Cross-Language Investigation into Jailbreak Attacks in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:48.346153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:48.346153Z digest=sha256:cb37d215f04d92c1bd6ee4dc38f71d29b7dc31290aa57aa259c4b87b2fff4695

Observation c1131312-66ab-401b-9424-f9a97a5eef6b · outbound

This paper cites Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:48:50.397734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.474079Z digest=sha256:75291dd4e27820136f168e71639923fc829b9410a653d9c4d5d7b5ffec2e5c4b

Observation 46e4f696-0677-40f2-bc22-4e2c92e88b0d · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.681403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.573825Z digest=sha256:4dbf509283d8efd1202d5e0c33ad288a23fe4e84caa0cdb2ac2533a0ea794fce

Observation bf59ab63-6167-4b78-957b-e628b03c71b0 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.434370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.689413Z digest=sha256:a01f3d900d488d91ab0be89ac3950454673a01bdc061e52d860c1137a7dc3e6d

Observation ba4de435-e0d6-4f05-b7c2-f8a12de7a683 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.149128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.797993Z digest=sha256:ef448eb6621f02d0964e470f7806031982f734892bcdb7de08a721ab608bccad

Observation c39f7f01-617c-4c84-ad77-4574af70a24c · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.827995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.913829Z digest=sha256:388d766cfb43c58820f34b1ca7dd2681e620c0a40848fe48e312b9411a07c5ae

Observation ddfb7a64-a005-4b07-ae67-0b2cca66b80b · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.587110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.069998Z digest=sha256:566b122b88e3be0f427d7ae97a8b812adf79235a581ed224cf91241154a5bed5

Observation 5c0b3f9e-5023-42cf-bfe0-6d65e5718c24 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.223511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.223511Z digest=sha256:0aa332fca961d09a5f632b9db9c35c188c78d8ae3f0e68d44ded154eb19a2723

Observation cebaace7-c0f9-4320-b9e0-41c4c5bacbfc · outbound

This paper cites When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.304875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.304875Z digest=sha256:0dea216288f47d90e32eb7741913ff0a300c8a2ecb959284aa074956e90b5976

Observation 71d8efd2-22c5-40c3-b32e-9750cd307da2 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Low-Resource Languages Jailbreak GPT-4

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.383630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.383630Z digest=sha256:1d4895e28ec8c3afa38ed7bff51e78811162d2284afe9aefc7a321d26e207dcf

Observation 90c79157-601a-4a9c-b628-c8fb11fc3899 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.352287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.488227Z digest=sha256:31bd529f144a2c87d47924401b6990353281bc93e0e79db8b0a73447ec87430c

Observation 5eb38f5c-d248-4d90-a46d-cb6fd5a2b453 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.605652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.605652Z digest=sha256:0a36d2a50c1cf146de11faa22bb1eff86382e28def69f82e9cc2b23168ac0bee

Observation 8f71313a-8b27-43e3-8cfa-672f2f284fc4 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.039113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.690266Z digest=sha256:47a1f9664ca5f91408ea2e92885fcabe93da5a9cbe5bb1da91cdacc6c4d3a2cf

Observation e3c2cb12-e5d4-4d55-8a36-ee36640ae800 · outbound

This paper cites Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:48:50.091776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.750627Z digest=sha256:b7c26bca4c6a553c138e5777d25e81a67db6dbed0c1ee021159921edfc374380

Observation 75d7b4ee-09fa-432f-8c66-4c1a3dd6297d · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.845907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.845907Z digest=sha256:48b2fed7c4a20e6059ad91260546d23ca706920044e4e40752c38503931f100e

Pith citing papers

No inbound Pith citation observations are available.