Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:37:59.145162Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.08029.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:37:59.145162Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6aa3591f-baa5-4be8-9d5b-eddc15455b21 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models , year=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6f4a3679-3a9f-4e9b-ad5f-e58ae87956b3 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0459a150-09f0-4ae5-a423-e6595df35361 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Representation Engineering: A Top-Down Approach to
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2264ccc-03d7-408f-a9a2-2d9634c47017 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b60c1275-addc-4a99-ad5e-ffeb853f940a · outbound
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation edb287ba-f08a-4853-a9ea-f7d7ca991962 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families ShieldGemma: Generative
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 089bc59a-3eb6-486d-952f-821c59431111 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 403188b2-7140-4a34-97ce-fc980f85a667 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families BeaverTails: Towards Improved Safety Alignment of
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 37740555-f152-459d-bd26-6dcc8323aa5e · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Advances in Neural Information Processing Systems , volume=
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6bc35da8-d369-4766-9617-de5a12a58ab5 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e3fd81f6-42dd-4367-b6d4-2f5d23e04722 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families arXiv preprint arXiv:2601.04603 , year=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc24ee59-7394-4b51-be6f-63e0dacaacd6 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Advances in Neural Information Processing Systems , volume=
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25c05f1d-3c48-4ed7-aa4a-4348c8aa060e · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families International Conference on Learning Representations , year=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2843ac01-abd3-4dd1-99b7-15d7a4d27312 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families LLMScan: Causal Scan for
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 634ecce7-7a5e-4470-b498-3b5aa80f3571 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Lightweight Safety Guardrails Using Fine-Tuned
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8dce6f83-8d76-4438-8959-87ecc1158e48 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Jailbroken: How Does
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de493c2a-5478-498b-88df-56cf40d9b06b · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dfcb42b-9a9b-4d04-b5c6-be648ef27ec9 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Alignment faking in large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ad8465c-c8be-45a3-a110-f651cb5e2067 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Building Guardrails for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28979a47-960d-4006-b226-afbcf44d1b69 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Attention Tracker: Detecting Prompt Injection Attacks in
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bf60ba57-34d6-4a1d-8404-d53ac7ef3e24 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Probing Latent Subspaces in
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cdae0327-410e-4279-9bee-7c0a8fd798be · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families DeepContext: Stateful Real-Time Detection of Multi-Turn Adversarial Intent Drift in
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f3acae53-8611-4c0c-a4b3-6fb77fa7b976 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c658da-a5b0-4787-a0ec-ee021bffb0a6 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Qwen2 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f609ef15-9011-4242-ae66-c3678a3a0494 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Gemma 2: Improving Open Language Models at a Practical Size
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f4fe65-d58d-48d2-8113-6f11b288025c · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4f640a18-b274-4a2d-a025-cf64e22cb38f · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Harmful Intent as a Geometrically Recoverable Feature of
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e4a529d2-ad31-4e7d-8d5d-ae1e3a121040 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Before the Last Token: Diagnosing Final-Token Safety Probe Failures
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 13053263-34da-4c37-9478-554028ac65da · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 036d2190-aa09-4e12-a31a-1f1d2a8aa5f9 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families ACM SIGOPS Operating Systems Review , volume=
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 33784954-22a3-4fe7-b39e-cfa942e27e70 · outbound
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families Advances in Neural Information Processing Systems , volume=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
No inbound Pith citation observations are available.