Pith. sign in

Paper Citation Record · LEDGER

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.08029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08029 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:37:59.145162Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6aa3591f-baa5-4be8-9d5b-eddc15455b21 · outbound

This paper cites Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models , year=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models , year=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.672979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.012396Z digest=sha256:eb84d06df83d5114ea4fccfc6b151f838049c38ad265004ff8ff2f1784a8a8e1

Observation 6f4a3679-3a9f-4e9b-ad5f-e58ae87956b3 · outbound

This paper cites an unresolved cited work.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:37:59.658878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.017910Z digest=sha256:bc98fcc2c37805be1bf7d75fd921f0051816cf7f7136086c93018755fc2bdd98

Observation 0459a150-09f0-4ae5-a423-e6595df35361 · outbound

This paper cites Representation Engineering: A Top-Down Approach to.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Representation Engineering: A Top-Down Approach to

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.022697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.022697Z digest=sha256:2bfe00dab9be4896c6fbab310a69014559b2fe315bc2c65bbafddd4ac26c3a87

Observation d2264ccc-03d7-408f-a9a2-2d9634c47017 · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.027822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.027822Z digest=sha256:7aecdd97ba56979083cd7634f71fd18d4cd5b9c2ede93359a5076aeb1e28263c

Observation b60c1275-addc-4a99-ad5e-ffeb853f940a · outbound

This paper cites Llama Guard:.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Llama Guard:

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.626598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.032577Z digest=sha256:816124dd9e87e06ef3ad0bae5d6fa9b4caeefbc620088a8d3ca880ef0c51f209

Observation edb287ba-f08a-4853-a9ea-f7d7ca991962 · outbound

This paper cites ShieldGemma: Generative.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families ShieldGemma: Generative

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.611291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.037173Z digest=sha256:dbb2b6a456e11766daa66c46636a04841ebd1dafce43c51b393f920ed37bb34a

Observation 089bc59a-3eb6-486d-952f-821c59431111 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.597472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.042273Z digest=sha256:b4352acbd4884fe0b70318fdd1a569c39cf3c5e25ac08c5f31b728eb1afe2b1a

Observation 403188b2-7140-4a34-97ce-fc980f85a667 · outbound

This paper cites BeaverTails: Towards Improved Safety Alignment of.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families BeaverTails: Towards Improved Safety Alignment of

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.582454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.046851Z digest=sha256:bcf129efdf76d5c9f1d571367f083c7a49588ca524ab5d15c4769b579b2df1c9

Observation 37740555-f152-459d-bd26-6dcc8323aa5e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Advances in Neural Information Processing Systems , volume=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.567362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.051464Z digest=sha256:5815a900a4672b6f6374345685fc82b51af1ac49b96278a3801ab1f7c42a3566

Observation 6bc35da8-d369-4766-9617-de5a12a58ab5 · outbound

This paper cites an unresolved cited work.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:37:59.552713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.055967Z digest=sha256:2553fdbbd4b584b0cd05675ef3f2e414d71187ba083c2958b30236888407492f

Observation e3fd81f6-42dd-4367-b6d4-2f5d23e04722 · outbound

This paper cites arXiv preprint arXiv:2601.04603 , year=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families arXiv preprint arXiv:2601.04603 , year=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.060666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.060666Z digest=sha256:ad8800c26d02acf84d1ebe98e035615cd299c8cc5b0d7610689162fc9e45c83a

Observation fc24ee59-7394-4b51-be6f-63e0dacaacd6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Advances in Neural Information Processing Systems , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.065099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.065099Z digest=sha256:6855ecf47f356354ce4e94ef4f1443735785675ae3ad90b851c1a72a70aeb908

Observation 25c05f1d-3c48-4ed7-aa4a-4348c8aa060e · outbound

This paper cites International Conference on Learning Representations , year=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families International Conference on Learning Representations , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.069432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.069432Z digest=sha256:68e4f86f47ec77e77b14edd2308dc94df3534242a67ea465955ce3d211b9639b

Observation 2843ac01-abd3-4dd1-99b7-15d7a4d27312 · outbound

This paper cites LLMScan: Causal Scan for.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families LLMScan: Causal Scan for

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.520545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.073680Z digest=sha256:aba1e28bb276518d21ba2d6256a1e1fdccf90c53a9bf73b477677f4a7056e7fc

Observation 634ecce7-7a5e-4470-b498-3b5aa80f3571 · outbound

This paper cites Lightweight Safety Guardrails Using Fine-Tuned.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Lightweight Safety Guardrails Using Fine-Tuned

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.506408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.077918Z digest=sha256:a4d5912d7635c79710aae83e4c904e2d590e2e354d87cb42c19133baf7dff719

Observation 8dce6f83-8d76-4438-8959-87ecc1158e48 · outbound

This paper cites Jailbroken: How Does.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Jailbroken: How Does

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.082283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.082283Z digest=sha256:e7e4afa9e0a5b6d69516b64b383c458f93ccd30a790ca5a66e616862cb1592d4

Observation de493c2a-5478-498b-88df-56cf40d9b06b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.086570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.086570Z digest=sha256:5ae6c1f123d81c2b9ba8630d0a1a746a7182bcbd6f4fb8d9a26126e9e302cbfa

Observation 3dfcb42b-9a9b-4d04-b5c6-be648ef27ec9 · outbound

This paper cites Alignment faking in large language models.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Alignment faking in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.091007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.091007Z digest=sha256:40941c9afccb9cfe2bc442c24a8dcd078ffa831427f1a8e09dbad903ca07db58

Observation 5ad8465c-c8be-45a3-a110-f651cb5e2067 · outbound

This paper cites Building Guardrails for Large Language Models.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Building Guardrails for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.095718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.095718Z digest=sha256:50c55f935794dd172e0f332b716e0abca112c8276ea04c72ac2f9192f294d440

Observation 28979a47-960d-4006-b226-afbcf44d1b69 · outbound

This paper cites Attention Tracker: Detecting Prompt Injection Attacks in.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Attention Tracker: Detecting Prompt Injection Attacks in

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.483031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.099973Z digest=sha256:c7b264f77e026f423eed3b5d3e73c4fb9fe37647c97221955982fb1ea05e2222

Observation bf60ba57-34d6-4a1d-8404-d53ac7ef3e24 · outbound

This paper cites Probing Latent Subspaces in.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Probing Latent Subspaces in

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.468493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.103954Z digest=sha256:2f8c4dc2a1609ccf1ff768cfb7696e6168c543216c1d68320327d8c5a45fa1e1

Observation cdae0327-410e-4279-9bee-7c0a8fd798be · outbound

This paper cites DeepContext: Stateful Real-Time Detection of Multi-Turn Adversarial Intent Drift in.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families DeepContext: Stateful Real-Time Detection of Multi-Turn Adversarial Intent Drift in

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.455445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.107867Z digest=sha256:f7dcd6cf1f14e0be4ee2742aa47f85e308862928be60986593bf82e6ecd7b31a

Observation f3acae53-8611-4c0c-a4b3-6fb77fa7b976 · outbound

This paper cites an unresolved cited work.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.111778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.111778Z digest=sha256:51573b49389ab9c7993b2841ba47c4e2b7746f02cdbc88c57d4fefb721d32d3e

Observation d7c658da-a5b0-4787-a0ec-ee021bffb0a6 · outbound

This paper cites Qwen2 Technical Report.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Qwen2 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.116162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.116162Z digest=sha256:17a6318fdaa5b9d44410d63b37a92ebe63278edc335bb982573fe42406c4293a

Observation f609ef15-9011-4242-ae66-c3678a3a0494 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Gemma 2: Improving Open Language Models at a Practical Size

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.121029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.121029Z digest=sha256:e20242fd16a03f0ae3dc0cefd115021b8e6f2d7b233ab6594852d8616aab17b6

Observation 14f4fe65-d58d-48d2-8113-6f11b288025c · outbound

This paper cites an unresolved cited work.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:37:59.430852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.125213Z digest=sha256:dcd4086c9dab9906bcfaa6f726187130c76f670fcbd48b007be28c4981e0993c

Observation 4f640a18-b274-4a2d-a025-cf64e22cb38f · outbound

This paper cites Harmful Intent as a Geometrically Recoverable Feature of.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Harmful Intent as a Geometrically Recoverable Feature of

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.417130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.129157Z digest=sha256:ecc0feb2f0071e04bfdd5e26ec9e3df79c0ca92e3718fcdb18f9f482fb835dfc

Observation e4a529d2-ad31-4e7d-8d5d-ae1e3a121040 · outbound

This paper cites Before the Last Token: Diagnosing Final-Token Safety Probe Failures.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Before the Last Token: Diagnosing Final-Token Safety Probe Failures

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:37:59.186952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.133017Z digest=sha256:b8b12d327eb67728b0a85b25625848167f9c9ca0baaf68ad4f62d7ef17db641b

Observation 13053263-34da-4c37-9478-554028ac65da · outbound

This paper cites an unresolved cited work.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:37:59.401677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.137247Z digest=sha256:97a40110a6ec5fffdad922db0b416e10715ea40f746ebe7923385f9204235c78

Observation 036d2190-aa09-4e12-a31a-1f1d2a8aa5f9 · outbound

This paper cites ACM SIGOPS Operating Systems Review , volume=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families ACM SIGOPS Operating Systems Review , volume=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.386755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.141162Z digest=sha256:cc739c65e13220906f7df7fc50c4d3f0f92ae032a0c39999567c4c7c6b997d7f

Observation 33784954-22a3-4fe7-b39e-cfa942e27e70 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Advances in Neural Information Processing Systems , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.371701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.145162Z digest=sha256:296ba96988434e7ff67313da18fe1f746a3c8bbb30b710b2876e2e40e471774f

Pith citing papers

No inbound Pith citation observations are available.