Pith. sign in

Paper Citation Record · LEDGER

The Geometry of Harmfulness in LLMs through Subconcept Probing

As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2507.21141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21141 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:01:47.854894Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:09:32.736304Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T21:38:00.071469Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ecb6aa75-1eaf-4e6c-89a3-701ad29a7dff · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing On the Opportunities and Risks of Foundation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.765176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.765176Z digest=sha256:d406e387dea45b0d51e622da6d49237afe64de2f8c34dfc40c6a72b71e062265

Observation 82c5c459-c5e7-4430-96b4-4da7071c1999 · outbound

This paper cites Safety-Aware Fine-Tuning of Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Safety-Aware Fine-Tuning of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.772974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.772974Z digest=sha256:30ed14894ba278caf2761df0f89210781b4bb152c24bce7a9847f0b7807826da

Observation f360be14-326c-43cc-8947-f019ae1cba39 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

The Geometry of Harmfulness in LLMs through Subconcept Probing Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.776571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.776571Z digest=sha256:fd7f9d7b92a2451ae438d3cfc233490514f6316868a9214d814fe9ae72cbe7b2

Observation 7fb6beaf-c30a-4a6e-a4a4-0d5bf79bceec · outbound

This paper cites The Llama 3 Herd of Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.791537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.791537Z digest=sha256:ace6eacc3d0cc3c7fb23c847186c6e9aa2157989dcdeae1fb4e6a8572d2b476c

Observation a528a80c-c0b7-4f4d-9440-a7b03bb7d901 · outbound

This paper cites Safedpo: A simple approach to direct preference optimization with enhanced safety.

The Geometry of Harmfulness in LLMs through Subconcept Probing Safedpo: A simple approach to direct preference optimization with enhanced safety

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.801834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.801834Z digest=sha256:e4c1b605f00a49944a940c503c0b1ff35a5e0dd24ba5dd1f102741d19e4a31fa

Observation 6b0343f1-d2dc-41e9-9564-a4d601b82724 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.805088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.805088Z digest=sha256:0ff5cc40d2d38a23e5fb91785b2a2bb1a71b10ac58f9d27518e899f44d0b9ee1

Observation a1939382-08e1-4005-bcf6-ae0c11ce47c8 · outbound

This paper cites Enhancing LLM Safety via Constrained Direct Preference Optimization.

The Geometry of Harmfulness in LLMs through Subconcept Probing Enhancing LLM Safety via Constrained Direct Preference Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.808584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.808584Z digest=sha256:f77f1f0a99c78f649fb72f9960d99d795beee62419f608cff63b726d87a14e06

Observation ff0ef252-7b16-4373-8618-3decc99bb851 · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

The Geometry of Harmfulness in LLMs through Subconcept Probing The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.812006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.812006Z digest=sha256:590876e8533575c0dffa5485e94138e79eaebf4fd77ea8ce6d7a6f89ee731c29

Observation a681fd9b-fb00-472c-b52d-078489649203 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.815767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.815767Z digest=sha256:3f2d40c61d83a8f1c17b88f16aa84779062aba7171aca7de56fe05ef438aa535

Observation 509a0729-c1e0-46ab-a923-9a4e81bc055d · outbound

This paper cites Emergent Linear Representations in World Models of Self-Supervised Sequence Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Emergent Linear Representations in World Models of Self-Supervised Sequence Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.819681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.819681Z digest=sha256:b6e11ff901170e18ebf4ae33f1a469ee6de25274c148e967fb7b0ee18e3cb576

Observation 70ffcc5d-2c1c-4592-ac8a-bbbb34600921 · outbound

This paper cites The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions.

The Geometry of Harmfulness in LLMs through Subconcept Probing The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.822986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.822986Z digest=sha256:f58b55e65c0f0ae6eeecd92ffe51107f83934dc402cabdc61dcf8bf87ad67476

Observation 0a8fb2e9-746c-4bd9-808e-5dc127b0f233 · outbound

This paper cites Interpretable steering of large language models with feature guided activation additions.

The Geometry of Harmfulness in LLMs through Subconcept Probing Interpretable steering of large language models with feature guided activation additions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.208051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:01:47.826803Z digest=sha256:60f19438f03930d302fd1352a2c445c84b3e509f508e5626bef48812457b3b59

Observation a850eefb-8758-4c6a-934d-26aa0304363f · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Linear Representations of Sentiment in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.830115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.830115Z digest=sha256:c3c3c06b36272ee3d602f020f299300df8e0ed99a7112d3fa88f4ec421b030a8

Observation 4ad54f81-b217-4a3e-b17c-67d477fd1619 · outbound

This paper cites Steering Language Models With Activation Engineering.

The Geometry of Harmfulness in LLMs through Subconcept Probing Steering Language Models With Activation Engineering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.833410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.833410Z digest=sha256:856a35272ff2842e4a3ae05648df67567e6657466cd94dd6814b720970e41c7c

Observation 47b0c430-ee7b-49a7-acd6-18f093ce0bce · outbound

This paper cites pyvene: A library for understanding and improving PyTorch models via interventions.

The Geometry of Harmfulness in LLMs through Subconcept Probing pyvene: A library for understanding and improving PyTorch models via interventions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.198757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:01:47.836938Z digest=sha256:f32e755108d02774469fcc83b60408b666ddc5cdd316363936ee1b2b5c611d5d

Observation 3407a27e-cdb7-4591-9266-9b563e4e3180 · outbound

This paper cites URL https://aclanthology.org/2024.naacl-demo.

The Geometry of Harmfulness in LLMs through Subconcept Probing URL https://aclanthology.org/2024.naacl-demo

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.189182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:01:47.840771Z digest=sha256:e95f56078345484d1d730cd0c2f5953450b0d64fdeaa786d3424829054b409c0

Observation ea24d6d9-d573-408f-b00e-5409aa37c140 · outbound

This paper cites Qwen2 Technical Report.

The Geometry of Harmfulness in LLMs through Subconcept Probing Qwen2 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.843887Z digest=sha256:156295fdc42dd37662a56c26c98a0f6c59cab5782278742b41264fda67fdbb0c

Observation f5c4d94f-9caa-4c5e-8b5b-e722938ad976 · outbound

This paper cites From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs.

The Geometry of Harmfulness in LLMs through Subconcept Probing From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.847228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.847228Z digest=sha256:af2ff8a0959c7f1998e46c70a9c5b60952a1bdd9198c89c027f0d6becb06e9b4

Observation e59f07d6-563b-4f03-b42d-8bf5f66c7b67 · outbound

This paper cites Under review.

The Geometry of Harmfulness in LLMs through Subconcept Probing Under review

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.177822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:01:47.850568Z digest=sha256:a784e7ec8b7b8912995bbf2a4b0b2e9877904ffe1e19ba921b07accca63f3110

Observation 323d869c-b384-4e58-b0b8-df3e41f18161 · outbound

This paper cites an unresolved cited work.

The Geometry of Harmfulness in LLMs through Subconcept Probing Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:48.167501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:01:47.854894Z digest=sha256:6bb9b125b485142e8e8a3ef200c82e356e1132289a65bd0059dee4fe672fd70b

Observation 4a5a077b-c308-4fea-820a-662d0f93c565 · outbound

This paper cites doi: https:// doi.org/10.1016/S0031-3203(96)00142-2.

The Geometry of Harmfulness in LLMs through Subconcept Probing doi: https:// doi.org/10.1016/S0031-3203(96)00142-2

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.769206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.769206Z digest=sha256:afd184f52015b37cd002dcb33ff12dac00a9db31c2f33ed7df816403e7e1427e

Observation 580723f0-2f20-47b0-897a-9660e8d9a8e0 · outbound

This paper cites Refusal Behavior in Large Language Models: A Nonlinear Perspective.

The Geometry of Harmfulness in LLMs through Subconcept Probing Refusal Behavior in Large Language Models: A Nonlinear Perspective

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.794774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.794774Z digest=sha256:3dce63cea35b6fbec26ba2017daacf35d20ca06ac5f0a8bc50388c3c199f59a5

Observation 15050118-aeec-4901-9351-987148f61d04 · outbound

This paper cites emnlp-main.273.

The Geometry of Harmfulness in LLMs through Subconcept Probing emnlp-main.273

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.787746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.787746Z digest=sha256:de73ae8b43576d77475a9c899ca34855f42d579fbc0b3d909fe1ef83326d42b0

Observation 2ad56729-0f6b-43cc-88de-f5f7a594dbb2 · outbound

This paper cites Toy Models of Superposition.

The Geometry of Harmfulness in LLMs through Subconcept Probing Toy Models of Superposition

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.780339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.780339Z digest=sha256:9885f9381acc0e664e4e6535aea64a772b171223c51257a57d1f4d80e00b0e17

Observation 266f749f-500e-42c1-9e67-1ad2a6323e5f · outbound

This paper cites an unresolved cited work.

The Geometry of Harmfulness in LLMs through Subconcept Probing Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:48.218134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:01:47.784263Z digest=sha256:d700e846500f2667c52d13bfedf639b7d8b6196f2f22825319f8ecc8a24df57f

Observation 2b54dfec-7589-47ba-96ab-80ad776a304e · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

The Geometry of Harmfulness in LLMs through Subconcept Probing Refusal in Language Models Is Mediated by a Single Direction

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.757141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.757141Z digest=sha256:5bab8c8ebec634efe4562d19bc4e60674bbde7b014e0d055876116af65a46b04

Observation 9ab45db4-4393-4a33-b972-58749b0b829a · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp.

The Geometry of Harmfulness in LLMs through Subconcept Probing On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.761453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.761453Z digest=sha256:05a8ab90a80a1477334e5de96846d8d04b9a05ea4725570e8bfaa63906cbb9ae

Observation 89355c50-affb-4354-b6b5-88a10c87a02d · outbound

This paper cites On the Origins of Linear Representations in Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing On the Origins of Linear Representations in Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.798462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.798462Z digest=sha256:311ff80b1da588b809caba8fdc759717ed12e018b4c3d81ce250ddab5f0586b3

Pith citing papers

Observation 05b043eb-10a1-4ebf-9169-085a654c8545 · inbound

Before the Last Token: Diagnosing Final-Token Safety Probe Failures cites this paper.

Before the Last Token: Diagnosing Final-Token Safety Probe Failures The Geometry of Harmfulness in LLMs through Subconcept Probing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.073909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:37:02.439454Z digest=sha256:a41a8a2bb05c223b0c373b3aa53dc2b239f6f1c66436b63a51a0dfd08f8a1ffd

Observation 28a2ee74-b0c1-4dc9-a21e-8d16199d1ba7 · inbound

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators cites this paper.

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators The Geometry of Harmfulness in LLMs through Subconcept Probing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:09:32.736304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:09:32.736304Z digest=sha256:69269e63beb34170f7870efb3d73537337f44dac8e16d82d30f98d0f5a474550