Pith. sign in

Paper Citation Record · LEDGER

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift

As of 10 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 3 inbound Pith citation observations for arXiv:2602.14161.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.14161 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:22:39.222267Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:09:30.417971Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:43:50.534673Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e76383d9-2896-4b36-8446-a6c8bea3a9e1 · outbound

This paper cites Unveiling decision-making in LLMs for text classification: Extraction of influential and interpretable concepts with sparse autoencoders.arXiv preprint arXiv:2506.23951,.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Unveiling decision-making in LLMs for text classification: Extraction of influential and interpretable concepts with sparse autoencoders.arXiv preprint arXiv:2506.23951,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.431988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.431988Z digest=sha256:a49332d1cbe16001772cb5c15d57ec93f6c2511f29dc7854aa348f5825dfce20

Observation 92470775-d594-4e06-8be6-992f70fb6e7e · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.478153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.478153Z digest=sha256:a9fdd801d79f61c1a3d29e515155b3f8fe39516b12aa8a307ddf31e7e46e0aab

Observation b12eca75-ceb7-4ca5-9602-10cae427576a · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.615073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.615073Z digest=sha256:65d61c9725ba1996b11719e7528e8798761f8921bdbde1453d5a4a7a93d7499b

Observation ab54ac8e-75a1-4d4a-899d-0e9ddf54c45d · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.767419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.767419Z digest=sha256:bebf191108e4ac0f9d517cb1f27e2dc99cc8e4cbc230d5fa7f6f34d51c26296d

Observation 538c7225-0d0a-4c6b-ac48-d2634878d9c3 · outbound

This paper cites Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.844788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.844788Z digest=sha256:c02062cfbfc04dd9d580f69f2c3bd85f7462ca100db87b44c278a6cd62bd44a5

Observation 68e4120d-125c-4479-88e5-69c031c421b3 · outbound

This paper cites Explaining Language Models' Predictions with High-Impact Concepts.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Explaining Language Models' Predictions with High-Impact Concepts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.945796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.945796Z digest=sha256:bf8fd1d30d24df4ec1f8f37a616f215b1a413eddef6e56ea86596ebc7e9c7eb2

Observation eb5e9b94-fdc5-4e7c-997e-9570e765d67c · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Improving Alignment and Robustness with Circuit Breakers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:39.082604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:39.082604Z digest=sha256:0bb36b78b86eab19f73b3f8950f9bd222b4eff1856f0dcfda1c8f10c4d4e0eee

Observation eff26982-06f1-4e0f-9a47-8496a3d756c8 · outbound

This paper cites Subject: {subject}Body:{body}.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Subject: {subject}Body:{body}

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:39.136959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:39.136959Z digest=sha256:72c65e1323cd292e8a2362a60507caece80069b73af216f6639e2341dea9dc6e

Observation 76f18ec2-a151-4a22-890f-add5dfdf4f43 · outbound

This paper cites system” role, so we prepend system message content to the first user message. The model generates a classification (“safe.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift system” role, so we prepend system message content to the first user message. The model generates a classification (“safe

Reference 14

Resolution
malformed identifier
no resolver link, observed 2026-08-02T23:22:39.222267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:39.222267Z digest=sha256:bd3eaaabad282062b2adb3aa58e704f5053b8c046ead74c0424e2db1ee6c8192

Observation e0b1212b-14c2-415f-b0ce-d1aeb815e209 · outbound

This paper cites Baturay Saglam, Paul Kassianik, Blaine Nelson, Sajana Weerawardhena, Yaron Singer, and Amin Kar- basi.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Baturay Saglam, Paul Kassianik, Blaine Nelson, Sajana Weerawardhena, Yaron Singer, and Amin Kar- basi

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.703390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.703390Z digest=sha256:c67b23e0ce493992a85f778f1c1316379cfd77bea77a5831e7d275d5b699b22f

Observation 1793219f-8e9a-45c2-92fe-72f3cf5535db · outbound

This paper cites Are Sparse Autoencoders Useful? A Case Study in Sparse Probing.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Are Sparse Autoencoders Useful? A Case Study in Sparse Probing

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.310667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.310667Z digest=sha256:eb354cc93a620c5ff63a9834d17f8b318c826b91f2c4685ba4355c6cb9a4bf65

Observation d33871a0-14f6-4f9b-ad3e-adb08c095d09 · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.536363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.536363Z digest=sha256:e410a7b0e673ba17b7a7f65b9b10cbe229cdecb58551f923ba551c986420817b

Observation 1b5281b6-515d-4abc-ac12-33eaebfb0e99 · outbound

This paper cites Sparse autoencoder features for classifications and transferability.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Sparse autoencoder features for classifications and transferability

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.137371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.137371Z digest=sha256:6a97b1364c771faff449b3de7fd8f79619161aec1ab7b2b9f7a7e2608cbb397a

Observation a95826a6-5dd0-495a-a9f9-2bd572221255 · outbound

This paper cites DeepMind Mechanistic Interpretability Team.

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift DeepMind Mechanistic Interpretability Team

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T23:22:38.040897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:22:38.040897Z digest=sha256:ca8211772e043a4b25ebf8dcdec997dd8c258a7c8c0d84c04f2dc81c3867b5cd

Pith citing papers

Observation aca06d7a-b362-4148-8dfd-1c1afc97d3e2 · inbound

Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry cites this paper.

Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:35.156854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T08:56:10.736240Z digest=sha256:9067aea97a489e290c7242c45e1fdce5823dddb9f3224dfa88b6e379c26e99c4

Observation a5e3e63c-772b-48c1-9df4-4a376a89965b · inbound

Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals cites this paper.

Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:35.156854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T18:41:03.566306Z digest=sha256:39b13d3951c16b7607b803d509bc0169cf23b79605b7440d2917f8656df7db3e

Observation d797654d-db01-467c-bf4d-d05a795ca3ab · inbound

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators cites this paper.

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T07:09:30.417971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:09:30.417971Z digest=sha256:9087fdfaa664fc1e23269f6d14e08d13383348a18123d12104923bcc3a906549