Pith. sign in

Paper Citation Record · LEDGER

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.08029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08029 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:37:59.145162Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6aa3591f-baa5-4be8-9d5b-eddc15455b21 · outbound

This paper cites Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models , year=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models , year=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.672979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.012396Z digest=sha256:f214cb77754604e65514540a63937f84507743e01bc294c9a66be8f83c904a9f

Observation 6f4a3679-3a9f-4e9b-ad5f-e58ae87956b3 · outbound

This paper cites an unresolved cited work.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:37:59.658878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.017910Z digest=sha256:7b7fad2b74f16ed02a0d01c3cd4bd56defb9ce09b3c6f89bf7a4076201684e11

Observation 0459a150-09f0-4ae5-a423-e6595df35361 · outbound

This paper cites Representation Engineering: A Top-Down Approach to.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Representation Engineering: A Top-Down Approach to

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.022697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.022697Z digest=sha256:2bfe00dab9be4896c6fbab310a69014559b2fe315bc2c65bbafddd4ac26c3a87

Observation d2264ccc-03d7-408f-a9a2-2d9634c47017 · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.027822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.027822Z digest=sha256:7aecdd97ba56979083cd7634f71fd18d4cd5b9c2ede93359a5076aeb1e28263c

Observation b60c1275-addc-4a99-ad5e-ffeb853f940a · outbound

This paper cites Llama Guard:.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Llama Guard:

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.626598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.032577Z digest=sha256:2ce815ad534d0178a29790c907a627768264ee1302d30f91378f5642d762b32e

Observation edb287ba-f08a-4853-a9ea-f7d7ca991962 · outbound

This paper cites ShieldGemma: Generative.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families ShieldGemma: Generative

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.611291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.037173Z digest=sha256:43c735b7a84b170a5a18c815d4ccc3c8e39be6170453ae6209e017e6c29a6033

Observation 089bc59a-3eb6-486d-952f-821c59431111 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.597472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.042273Z digest=sha256:12f9ee1454e425a1ef40edc0283cb7f3ac6235aaeac33aeadb815c6dbcef10a6

Observation 403188b2-7140-4a34-97ce-fc980f85a667 · outbound

This paper cites BeaverTails: Towards Improved Safety Alignment of.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families BeaverTails: Towards Improved Safety Alignment of

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.582454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.046851Z digest=sha256:8092e872e37197e2f722d63ee14ff13d587ec6c6503f3b9abdc867ee74f0879a

Observation 37740555-f152-459d-bd26-6dcc8323aa5e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Advances in Neural Information Processing Systems , volume=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.567362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.051464Z digest=sha256:6e7263b1ecfbbc126d0d69997b2de4860dbd0fdfb9884e331b673d94e95a625d

Observation 6bc35da8-d369-4766-9617-de5a12a58ab5 · outbound

This paper cites an unresolved cited work.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:37:59.552713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.055967Z digest=sha256:1cdc1b14e8887a3eacd79276d0be1e2f7940be41ac98b497f9a1e4ac76680407

Observation e3fd81f6-42dd-4367-b6d4-2f5d23e04722 · outbound

This paper cites arXiv preprint arXiv:2601.04603 , year=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families arXiv preprint arXiv:2601.04603 , year=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.060666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.060666Z digest=sha256:ad8800c26d02acf84d1ebe98e035615cd299c8cc5b0d7610689162fc9e45c83a

Observation fc24ee59-7394-4b51-be6f-63e0dacaacd6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Advances in Neural Information Processing Systems , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.065099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.065099Z digest=sha256:6855ecf47f356354ce4e94ef4f1443735785675ae3ad90b851c1a72a70aeb908

Observation 25c05f1d-3c48-4ed7-aa4a-4348c8aa060e · outbound

This paper cites International Conference on Learning Representations , year=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families International Conference on Learning Representations , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.069432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.069432Z digest=sha256:68e4f86f47ec77e77b14edd2308dc94df3534242a67ea465955ce3d211b9639b

Observation 2843ac01-abd3-4dd1-99b7-15d7a4d27312 · outbound

This paper cites LLMScan: Causal Scan for.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families LLMScan: Causal Scan for

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.520545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.073680Z digest=sha256:45c26bf1a5b2c552c975f292ce37433163038616ed9fca006588df88f353eac3

Observation 634ecce7-7a5e-4470-b498-3b5aa80f3571 · outbound

This paper cites Lightweight Safety Guardrails Using Fine-Tuned.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Lightweight Safety Guardrails Using Fine-Tuned

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.506408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.077918Z digest=sha256:05fddda3cbdbd2a823cb255c859a839d4989193fcef253b21f3c3a983c63c620

Observation 8dce6f83-8d76-4438-8959-87ecc1158e48 · outbound

This paper cites Jailbroken: How Does.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Jailbroken: How Does

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.082283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.082283Z digest=sha256:e7e4afa9e0a5b6d69516b64b383c458f93ccd30a790ca5a66e616862cb1592d4

Observation de493c2a-5478-498b-88df-56cf40d9b06b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.086570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.086570Z digest=sha256:5ae6c1f123d81c2b9ba8630d0a1a746a7182bcbd6f4fb8d9a26126e9e302cbfa

Observation 3dfcb42b-9a9b-4d04-b5c6-be648ef27ec9 · outbound

This paper cites Alignment faking in large language models.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Alignment faking in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.091007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.091007Z digest=sha256:40941c9afccb9cfe2bc442c24a8dcd078ffa831427f1a8e09dbad903ca07db58

Observation 5ad8465c-c8be-45a3-a110-f651cb5e2067 · outbound

This paper cites Building Guardrails for Large Language Models.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Building Guardrails for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.095718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.095718Z digest=sha256:50c55f935794dd172e0f332b716e0abca112c8276ea04c72ac2f9192f294d440

Observation 28979a47-960d-4006-b226-afbcf44d1b69 · outbound

This paper cites Attention Tracker: Detecting Prompt Injection Attacks in.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Attention Tracker: Detecting Prompt Injection Attacks in

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.483031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.099973Z digest=sha256:6d69fa9eacbf3cfb4b5220a59fdd34113b0a16814893102242c357da6f8b719a

Observation bf60ba57-34d6-4a1d-8404-d53ac7ef3e24 · outbound

This paper cites Probing Latent Subspaces in.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Probing Latent Subspaces in

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.468493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.103954Z digest=sha256:fb0951c63e49b6866e64abf852e5dbdeb7e62464a09f7c01703cdd69d009c8ba

Observation cdae0327-410e-4279-9bee-7c0a8fd798be · outbound

This paper cites DeepContext: Stateful Real-Time Detection of Multi-Turn Adversarial Intent Drift in.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families DeepContext: Stateful Real-Time Detection of Multi-Turn Adversarial Intent Drift in

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.455445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.107867Z digest=sha256:d439a354b90e1d9b029774f513ede6bf8f69439bad1711f6ba6f1c76c2124bae

Observation f3acae53-8611-4c0c-a4b3-6fb77fa7b976 · outbound

This paper cites an unresolved cited work.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.111778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.111778Z digest=sha256:51573b49389ab9c7993b2841ba47c4e2b7746f02cdbc88c57d4fefb721d32d3e

Observation d7c658da-a5b0-4787-a0ec-ee021bffb0a6 · outbound

This paper cites Qwen2 Technical Report.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Qwen2 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.116162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.116162Z digest=sha256:17a6318fdaa5b9d44410d63b37a92ebe63278edc335bb982573fe42406c4293a

Observation f609ef15-9011-4242-ae66-c3678a3a0494 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Gemma 2: Improving Open Language Models at a Practical Size

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:37:59.121029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:37:59.121029Z digest=sha256:e20242fd16a03f0ae3dc0cefd115021b8e6f2d7b233ab6594852d8616aab17b6

Observation 14f4fe65-d58d-48d2-8113-6f11b288025c · outbound

This paper cites an unresolved cited work.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:37:59.430852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.125213Z digest=sha256:c9b89fd739c0f9d95bdc53c6e1af922d17a1564118d0217cf2b81a5ab116b9e6

Observation 4f640a18-b274-4a2d-a025-cf64e22cb38f · outbound

This paper cites Harmful Intent as a Geometrically Recoverable Feature of.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Harmful Intent as a Geometrically Recoverable Feature of

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.417130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.129157Z digest=sha256:ed7ae0c58c14fe11f2685300348c26fe66347dd26afbb98771f351f83a8c069a

Observation e4a529d2-ad31-4e7d-8d5d-ae1e3a121040 · outbound

This paper cites Before the Last Token: Diagnosing Final-Token Safety Probe Failures.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Before the Last Token: Diagnosing Final-Token Safety Probe Failures

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:37:59.186952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.133017Z digest=sha256:23091395858d302a69055d5399d6e0115f880292bdd01d5aaf576f7bbae16a22

Observation 13053263-34da-4c37-9478-554028ac65da · outbound

This paper cites an unresolved cited work.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:37:59.401677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.137247Z digest=sha256:31e23971508ee7efed54613a419a23c6e26a71c5cea94aff6fb232e6ce100033

Observation 036d2190-aa09-4e12-a31a-1f1d2a8aa5f9 · outbound

This paper cites ACM SIGOPS Operating Systems Review , volume=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families ACM SIGOPS Operating Systems Review , volume=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.386755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.141162Z digest=sha256:3e7d0c987cf054307ac1d8d0f443c40f309d718d334e1a2dfd02d176652f7b1b

Observation 33784954-22a3-4fe7-b39e-cfa942e27e70 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Advances in Neural Information Processing Systems , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:37:59.371701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:37:59.145162Z digest=sha256:86a2d54869312dc2de74652267576e2664eafd9fb61cd96ebebb4e88abc5e3c1

Pith citing papers

No inbound Pith citation observations are available.