Pith. sign in

Paper Citation Record · LEDGER

The Geometry of Harmfulness in LLMs through Subconcept Probing

As of 20 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2507.21141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21141 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:01:47.854894Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:09:32.736304Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T21:38:00.071469Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ecb6aa75-1eaf-4e6c-89a3-701ad29a7dff · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing On the Opportunities and Risks of Foundation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.765176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.765176Z digest=sha256:6695569813612fba165b5f6a6330c9bbd5dfd8afa42a29ddad9f68941bb18115

Observation 82c5c459-c5e7-4430-96b4-4da7071c1999 · outbound

This paper cites Safety-Aware Fine-Tuning of Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Safety-Aware Fine-Tuning of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.772974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.772974Z digest=sha256:d2ea5b76ba3da78405e3d8e43ec9f5004e3e98f3b8e2867eefc62d5c3db92d61

Observation f360be14-326c-43cc-8947-f019ae1cba39 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

The Geometry of Harmfulness in LLMs through Subconcept Probing Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.776571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.776571Z digest=sha256:445e6f4ed307fcf1518ceb254fa49d869a27bac8f80211f938e213eb50ba76db

Observation 7fb6beaf-c30a-4a6e-a4a4-0d5bf79bceec · outbound

This paper cites The Llama 3 Herd of Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.791537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.791537Z digest=sha256:5670e622a8908d9da9068ee054ed69041035e372683944af8b99ff1c9ee5ce3b

Observation a528a80c-c0b7-4f4d-9440-a7b03bb7d901 · outbound

This paper cites Safedpo: A simple approach to direct preference optimization with enhanced safety.

The Geometry of Harmfulness in LLMs through Subconcept Probing Safedpo: A simple approach to direct preference optimization with enhanced safety

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.801834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.801834Z digest=sha256:5e13a7d68a6a93578d5dcc9616c93b56a77b87ea31c62acf2ac5e9e8e49759e5

Observation 6b0343f1-d2dc-41e9-9564-a4d601b82724 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.805088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.805088Z digest=sha256:d4c8baf23b83ac5cbb9720b6568115dac7726238b7f9d85a8448be5646254aee

Observation a1939382-08e1-4005-bcf6-ae0c11ce47c8 · outbound

This paper cites Enhancing LLM Safety via Constrained Direct Preference Optimization.

The Geometry of Harmfulness in LLMs through Subconcept Probing Enhancing LLM Safety via Constrained Direct Preference Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.808584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.808584Z digest=sha256:61030af4940164eb7cbb95a35bf0de3d361c8d21eae990ce3cb0614a5448523e

Observation ff0ef252-7b16-4373-8618-3decc99bb851 · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

The Geometry of Harmfulness in LLMs through Subconcept Probing The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.812006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.812006Z digest=sha256:250668f9098fc24372dd8a358c49cda6f8be0b80cc9370a1e52e70fbcfaf552e

Observation a681fd9b-fb00-472c-b52d-078489649203 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.815767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.815767Z digest=sha256:dfdd73673827a40798aa240587813f754edaf6367f25d209d932dedf6d37aa65

Observation 509a0729-c1e0-46ab-a923-9a4e81bc055d · outbound

This paper cites Emergent Linear Representations in World Models of Self-Supervised Sequence Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Emergent Linear Representations in World Models of Self-Supervised Sequence Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.819681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.819681Z digest=sha256:3dce6e908f952c9c8823a28faee0ed4e556aea9c49d65638f09481179d65b138

Observation 70ffcc5d-2c1c-4592-ac8a-bbbb34600921 · outbound

This paper cites The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions.

The Geometry of Harmfulness in LLMs through Subconcept Probing The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.822986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.822986Z digest=sha256:72e9b48ba481358062e89829fa55c46a61ffe92a6e3b3a1085d909c77c9d9e57

Observation 0a8fb2e9-746c-4bd9-808e-5dc127b0f233 · outbound

This paper cites Interpretable steering of large language models with feature guided activation additions.

The Geometry of Harmfulness in LLMs through Subconcept Probing Interpretable steering of large language models with feature guided activation additions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.208051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:47.826803Z digest=sha256:e1a5fd6afb70400e99d9f069191a44593ebf76d8a96d0146e33c961fa732863c

Observation a850eefb-8758-4c6a-934d-26aa0304363f · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Linear Representations of Sentiment in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.830115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.830115Z digest=sha256:5d68be9b9a6ad5abd0e3a98e8ab1ed0e3c9ffb88e82c19c71d5fc36f425aff0c

Observation 4ad54f81-b217-4a3e-b17c-67d477fd1619 · outbound

This paper cites Steering Language Models With Activation Engineering.

The Geometry of Harmfulness in LLMs through Subconcept Probing Steering Language Models With Activation Engineering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.833410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.833410Z digest=sha256:dfa29903e2699a15bed4eaab9664cd035515c1f4cda09a27c83278187397fd63

Observation 47b0c430-ee7b-49a7-acd6-18f093ce0bce · outbound

This paper cites pyvene: A library for understanding and improving PyTorch models via interventions.

The Geometry of Harmfulness in LLMs through Subconcept Probing pyvene: A library for understanding and improving PyTorch models via interventions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.198757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:47.836938Z digest=sha256:1b683d55661cdd6ebe860cf3181ddc2c9dbb2343a14ec962e23c70e197d2bdf9

Observation 3407a27e-cdb7-4591-9266-9b563e4e3180 · outbound

This paper cites URL https://aclanthology.org/2024.naacl-demo.

The Geometry of Harmfulness in LLMs through Subconcept Probing URL https://aclanthology.org/2024.naacl-demo

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.189182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:47.840771Z digest=sha256:b50125487d32e15e79884ad17281d1a67b0e355bf11e1900c6a0744dd4914f9f

Observation ea24d6d9-d573-408f-b00e-5409aa37c140 · outbound

This paper cites Qwen2 Technical Report.

The Geometry of Harmfulness in LLMs through Subconcept Probing Qwen2 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.843887Z digest=sha256:2132186f4cdf4b5056636619fe2185b0c51e94c52315f4341677a09a67ad3b3d

Observation f5c4d94f-9caa-4c5e-8b5b-e722938ad976 · outbound

This paper cites From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs.

The Geometry of Harmfulness in LLMs through Subconcept Probing From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.847228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.847228Z digest=sha256:7d2a9245dfc6bc5836156575aaef81260d7535b3568bba6efdb4c19f87e07b32

Observation e59f07d6-563b-4f03-b42d-8bf5f66c7b67 · outbound

This paper cites Under review.

The Geometry of Harmfulness in LLMs through Subconcept Probing Under review

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.177822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:47.850568Z digest=sha256:1c32009e3a533e3d8ae23d9889374b2aa310d564717d28dfbef0b9a6b356bd29

Observation 323d869c-b384-4e58-b0b8-df3e41f18161 · outbound

This paper cites an unresolved cited work.

The Geometry of Harmfulness in LLMs through Subconcept Probing Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:48.167501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:47.854894Z digest=sha256:3d39d84eff5744944713b1710a756c5c81ddb42bcc427771b4acad05bf5b55e7

Observation 4a5a077b-c308-4fea-820a-662d0f93c565 · outbound

This paper cites doi: https:// doi.org/10.1016/S0031-3203(96)00142-2.

The Geometry of Harmfulness in LLMs through Subconcept Probing doi: https:// doi.org/10.1016/S0031-3203(96)00142-2

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.769206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.769206Z digest=sha256:bc94fa1decca818a4ae4c58c132a385431543a43f34613e226bf83cad700da59

Observation 580723f0-2f20-47b0-897a-9660e8d9a8e0 · outbound

This paper cites Refusal Behavior in Large Language Models: A Nonlinear Perspective.

The Geometry of Harmfulness in LLMs through Subconcept Probing Refusal Behavior in Large Language Models: A Nonlinear Perspective

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.794774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.794774Z digest=sha256:2fb80b4a1be7b9d433a98038c89e37173f69f62ba19689d1221bd9877d44f1be

Observation 15050118-aeec-4901-9351-987148f61d04 · outbound

This paper cites emnlp-main.273.

The Geometry of Harmfulness in LLMs through Subconcept Probing emnlp-main.273

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.787746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.787746Z digest=sha256:8ce053ac415e06a01e35eecf5f2f2ece4ba898d9ec2d3b10517881313d71af74

Observation 2ad56729-0f6b-43cc-88de-f5f7a594dbb2 · outbound

This paper cites Toy Models of Superposition.

The Geometry of Harmfulness in LLMs through Subconcept Probing Toy Models of Superposition

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.780339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.780339Z digest=sha256:951bf8bc50225c02b1efbc10a80b3a8ee56a405798ddde5fdddba0b08e56f88e

Observation 266f749f-500e-42c1-9e67-1ad2a6323e5f · outbound

This paper cites an unresolved cited work.

The Geometry of Harmfulness in LLMs through Subconcept Probing Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:48.218134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:47.784263Z digest=sha256:3dfed307fb0561946402ea9697ac269d9b0cb62c714e42ddecd2fe7a65e103b5

Observation 2b54dfec-7589-47ba-96ab-80ad776a304e · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

The Geometry of Harmfulness in LLMs through Subconcept Probing Refusal in Language Models Is Mediated by a Single Direction

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.757141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.757141Z digest=sha256:b3f6c2109c01144c192a044958cf7c5e96641030285f4b2f8c0b8902870495cb

Observation 9ab45db4-4393-4a33-b972-58749b0b829a · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp.

The Geometry of Harmfulness in LLMs through Subconcept Probing On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.761453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.761453Z digest=sha256:5ea8485b3d864abab99ef58e571e8543ab9e04a6540156624897dad66304bae3

Observation 89355c50-affb-4354-b6b5-88a10c87a02d · outbound

This paper cites On the Origins of Linear Representations in Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing On the Origins of Linear Representations in Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.798462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.798462Z digest=sha256:54e18b6b4c5002abebe5ffcde04118b50222820cb4069e95098bf9f6bc38d52a

Pith citing papers

Observation 05b043eb-10a1-4ebf-9169-085a654c8545 · inbound

Before the Last Token: Diagnosing Final-Token Safety Probe Failures cites this paper.

Before the Last Token: Diagnosing Final-Token Safety Probe Failures The Geometry of Harmfulness in LLMs through Subconcept Probing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.073909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T21:37:02.439454Z digest=sha256:5e233b0f64f8162638318d35d680f22ed04b2f7d9b87559d4b10b1cf4a1dd3b3

Observation 28a2ee74-b0c1-4dc9-a21e-8d16199d1ba7 · inbound

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators cites this paper.

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators The Geometry of Harmfulness in LLMs through Subconcept Probing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:09:32.736304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:09:32.736304Z digest=sha256:9e7008d606f099c35ac12a435df764b5f070ce978abdf04ba62104e9216fbace