Pith. sign in

Paper Citation Record · LEDGER

Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2409.00598.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.00598 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:10:30.898210Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T13:54:58.586248Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cfb22ddd-294c-4bd9-a6ea-c1a035716acd · inbound

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models cites this paper.

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T14:51:58.694027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:51:58.694027Z digest=sha256:1dc7e703c9af6912ea31b656cade41048d3264df8bfa8d18f3d9701987dc5142

Observation 0503e350-69ab-4040-9474-d413bfc24ec8 · inbound

Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests cites this paper.

Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T19:21:19.146827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:21:19.146827Z digest=sha256:ec443547b611c82e16a76c309d98d4a28db88f9e0f807190ddc763a1c05ed95a

Observation 62ae4d59-dd8c-457d-8afb-1fa5b705cfbc · inbound

DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification cites this paper.

DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:30.898210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:10:30.898210Z digest=sha256:f53c70457901acf37fc23e4d752a880ceeabf2ecc68b9096a2170766a84fb492

Observation 4bf1d9f4-6e4d-43e9-ae6a-df2c055a56fc · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.686789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.686789Z digest=sha256:29f7f2691c0987c9ccb833f692638d09123d53720d0d655e88c634da7f13a3f4

Observation 8d3d4fb5-83c3-42de-a7a6-611ceb89fe6f · inbound

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning cites this paper.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.379859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.379859Z digest=sha256:bd130270845318fdc9379ffaa8b58bd91e46fc8e6de68479472e9a79597f48f6

Observation 7acea42f-75e7-47b0-adfb-176c818a5870 · inbound

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection cites this paper.

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T22:24:23.458700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:24:23.458700Z digest=sha256:0497741cfbc3a59ee444d291410dd3d00fad44ef9666ff1af3016a89ba16fe0b

Observation 98445783-ce10-4e6f-b0f2-4bd1377432d8 · inbound

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? cites this paper.

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:29:39.426113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T13:22:12.541923Z digest=sha256:0aa64e4882cb5ae074c9fdaf44f15e3aa747742ee38ea8ed6b952e0a430baed5

Observation 32161650-bd79-4f33-8b03-e244e9084edf · inbound

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability cites this paper.

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T13:54:58.587659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-08T13:50:54.083165Z digest=sha256:0a76d093a0464b3f703d8a4f0aec34d19b361346aaa77e200d27b271d11cfa68

Observation e7c07c8b-79d4-47b2-bad5-26de0370aa68 · inbound

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators cites this paper.

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T07:09:29.730590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:09:29.730590Z digest=sha256:b9ae32c54de6efa5be03e0ce360a1472f15de4931d4c9633c2cf282ee0a6d15d

Observation 73d5e304-eaf1-4e55-adbf-64fc2f51e490 · inbound

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity cites this paper.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:47.691899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:47.691899Z digest=sha256:2613b7266cfeef7ebfb9d85a546b62dc7287996a36d986cb7a862dc49443525f

Observation 9b8a67bf-7cd3-45e2-bfe7-67bd7f040adc · inbound

Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety cites this paper.

Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T14:12:28.323410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:12:28.323410Z digest=sha256:9bb547912ab38a01fb9ca0b4041cde60523e7015fa0e473ca7b80694b111b231