Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:01:47.854894Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2507.21141.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:01:47.854894Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:09:32.736304Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-14T21:38:00.071469Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ecb6aa75-1eaf-4e6c-89a3-701ad29a7dff · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing On the Opportunities and Risks of Foundation Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82c5c459-c5e7-4430-96b4-4da7071c1999 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Safety-Aware Fine-Tuning of Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f360be14-326c-43cc-8947-f019ae1cba39 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb6beaf-c30a-4a6e-a4a4-0d5bf79bceec · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a528a80c-c0b7-4f4d-9440-a7b03bb7d901 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Safedpo: A simple approach to direct preference optimization with enhanced safety
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0343f1-d2dc-41e9-9564-a4d601b82724 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1939382-08e1-4005-bcf6-ae0c11ce47c8 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Enhancing LLM Safety via Constrained Direct Preference Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff0ef252-7b16-4373-8618-3decc99bb851 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a681fd9b-fb00-472c-b52d-078489649203 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 509a0729-c1e0-46ab-a923-9a4e81bc055d · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Emergent Linear Representations in World Models of Self-Supervised Sequence Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ffcc5d-2c1c-4592-ac8a-bbbb34600921 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a8fb2e9-746c-4bd9-808e-5dc127b0f233 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Interpretable steering of large language models with feature guided activation additions
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a850eefb-8758-4c6a-934d-26aa0304363f · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Linear Representations of Sentiment in Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad54f81-b217-4a3e-b17c-67d477fd1619 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Steering Language Models With Activation Engineering
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b0c430-ee7b-49a7-acd6-18f093ce0bce · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing pyvene: A library for understanding and improving PyTorch models via interventions
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3407a27e-cdb7-4591-9266-9b563e4e3180 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing URL https://aclanthology.org/2024.naacl-demo
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea24d6d9-d573-408f-b00e-5409aa37c140 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Qwen2 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c4d94f-9caa-4c5e-8b5b-e722938ad976 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e59f07d6-563b-4f03-b42d-8bf5f66c7b67 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Under review
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 323d869c-b384-4e58-b0b8-df3e41f18161 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a5a077b-c308-4fea-820a-662d0f93c565 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing doi: https:// doi.org/10.1016/S0031-3203(96)00142-2
Reference 1997
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580723f0-2f20-47b0-897a-9660e8d9a8e0 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Refusal Behavior in Large Language Models: A Nonlinear Perspective
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15050118-aeec-4901-9351-987148f61d04 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing emnlp-main.273
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad56729-0f6b-43cc-88de-f5f7a594dbb2 · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Toy Models of Superposition
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 266f749f-500e-42c1-9e67-1ad2a6323e5f · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b54dfec-7589-47ba-96ab-80ad776a304e · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing Refusal in Language Models Is Mediated by a Single Direction
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ab45db4-4393-4a33-b972-58749b0b829a · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89355c50-affb-4354-b6b5-88a10c87a02d · outbound
The Geometry of Harmfulness in LLMs through Subconcept Probing On the Origins of Linear Representations in Large Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b043eb-10a1-4ebf-9169-085a654c8545 · inbound
Before the Last Token: Diagnosing Final-Token Safety Probe Failures The Geometry of Harmfulness in LLMs through Subconcept Probing
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28a2ee74-b0c1-4dc9-a21e-8d16199d1ba7 · inbound
The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators The Geometry of Harmfulness in LLMs through Subconcept Probing
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.