Pith. sign in

Paper Citation Record · LEDGER

Interpretable Steering of Large Language Models with Feature Guided Activation Additions

As of 17 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 13 inbound Pith citation observations for arXiv:2501.09929.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09929 v3

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:34:26.980516Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:35:04.713574Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation d87ac576-f7f9-4db6-9b55-bd6313842b14 · outbound

This paper cites Improving Steering Vectors by Targeting Sparse Autoencoder Features.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Improving Steering Vectors by Targeting Sparse Autoencoder Features

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:26.909588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:34:26.909588Z digest=sha256:46f41c5b8e2b1bd16714231017f99f67bd062016cf9898dfd68c78941104d016

Observation dae8fa78-5f68-49cc-b9f6-fde2532d1909 · outbound

This paper cites an unresolved cited work.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:26.914583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:34:26.914583Z digest=sha256:d2b49995daa83ec5ffec819939d5fe4299b6c80cfb3a292dfe3d3dba32bbebbd

Observation 3fb11239-5157-4fb7-8f87-384c9bf0e60c · outbound

This paper cites <bos>I think this is a photo of a giant squid attacking a Russian submarine, and it is one of the most Incredible Aliens captured in Antarctica! These mind.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions <bos>I think this is a photo of a giant squid attacking a Russian submarine, and it is one of the most Incredible Aliens captured in Antarctica! These mind

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:27.243131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T19:34:26.971071Z digest=sha256:1d4c1c32189f4e53694e150378b48e3aec9c66bbc0e1efd6d8572b65fd994b47

Observation c6660f73-7181-4e0c-b75d-3f93070becb9 · outbound

This paper cites Kiho Park, Yo Joong Choe, and Victor Veitch.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Kiho Park, Yo Joong Choe, and Victor Veitch

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:27.298229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T19:34:26.924549Z digest=sha256:9e105e8a0652da9a656178d0c33bca47c374eb0c97e25f2e316210a722e9de9e

Observation 9604e859-cb4e-4ed3-8088-fca4d0d0d53b · outbound

This paper cites 10 Published at Building Trust Workshop at ICLR 2025 Gonçalo Paulo, Alex Mallen, Caden Juang, and Nora Belrose.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions 10 Published at Building Trust Workshop at ICLR 2025 Gonçalo Paulo, Alex Mallen, Caden Juang, and Nora Belrose

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:27.283520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T19:34:26.928833Z digest=sha256:fba0c6dc0a5cdce0d69176c2792843db4acd6baa70e477a570d242d79c50eb6f

Observation bf2712e3-c2a5-4c81-826a-ad5abe3ff685 · outbound

This paper cites Automatically Interpreting Millions of Features in Large Language Models.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Automatically Interpreting Millions of Features in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:26.933320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:34:26.933320Z digest=sha256:a20e25b6d2f7686b7793e9d8094289a93fae92f4efaa0c3cf00fc1406af48d16

Observation 46f99a6f-72eb-473b-b2d3-27b888fcaef2 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Gemma 2: Improving Open Language Models at a Practical Size

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:26.939993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:34:26.939993Z digest=sha256:69cf45f069adfda2bd3e61d972282f2e6704d5d218007c84c740ab5dad7dc965

Observation e2217090-5ff5-4b2b-80c3-0152c913626c · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Gemma 2: Improving Open Language Models at a Practical Size

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:26.944662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:34:26.944662Z digest=sha256:765da2a9efd9ca95d0bd7b6c3111064eca6c46642d78cad39f38bc05f9932027

Observation 48140458-dea3-4ee5-ad4e-29366c4d5a3c · outbound

This paper cites Polysemanticity and Capacity in Neural Networks.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Polysemanticity and Capacity in Neural Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:26.949010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:34:26.949010Z digest=sha256:50f6f4e8b3f0b16a81c69de34a3fe376fecd8301ff11f539e2bedfa307c606fc

Observation 173e66f3-3c14-4d4d-b6b7-bb39e76ef4fd · outbound

This paper cites Steering Language Models With Activation Engineering.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Steering Language Models With Activation Engineering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:26.953810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:34:26.953810Z digest=sha256:9e39c51d1c6b41ced97d84004d511bc8ebc0c765a2b057b147089cbc0fecd79a

Observation 22ad1bad-899b-4aa7-a43a-2947b904b3c0 · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:26.959178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:34:26.959178Z digest=sha256:6b61b35060d0a7290dae0ace55b5fe257c61ee0f4f906e57bf794ff018c0ee3f

Observation 6a22c022-9f4a-4eff-b2c3-42d62027012c · outbound

This paper cites an unresolved cited work.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:27.269169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T19:34:26.963315Z digest=sha256:3b02a7a9b0723b30284b05e2d0ef90fb313b7fe17a7f937cd8cf6f8959d432ca

Observation 847f23d4-b30a-45cc-b955-88105f8c1a17 · outbound

This paper cites The government is hiding the truth about alien contact.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions The government is hiding the truth about alien contact

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:27.256087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T19:34:26.967448Z digest=sha256:6354aa96ea311229e6a7018bd0d446e93f7716e2082ab2d7ccd10f3b0d636b2e

Observation 9dbcc039-e3b8-4a68-938a-3451a94a6799 · outbound

This paper cites me" in different contexts -1.039 2605 References to presence or absence of evidence Rollouts at Scale = 80 (Optimal Scale): 17 Published at Building Trust Workshop at ICLR 2025.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions me" in different contexts -1.039 2605 References to presence or absence of evidence Rollouts at Scale = 80 (Optimal Scale): 17 Published at Building Trust Workshop at ICLR 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:27.229066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T19:34:26.974852Z digest=sha256:53730b87f74091e70144f6955330855a7adf3455fee00141d010ec8debfa1b25

Observation 6dda3c0d-9531-41ce-b537-007f64dda116 · outbound

This paper cites Evaluate the text based on the criterion.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Evaluate the text based on the criterion

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:27.213787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T19:34:26.980516Z digest=sha256:b8834c683c2df6dedae832a8e46a82203c7ca924fd0aac3f885650b2022e75e0

Observation 3a57c897-f4e2-4c75-b725-34e073feaf9b · outbound

This paper cites Aaron Gokaslan and Vanya Cohen.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Aaron Gokaslan and Vanya Cohen

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:27.311448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T19:34:26.919961Z digest=sha256:a7d17c0403fb8797348dffa58e5a000ef7ab83e49f9a072e2eb92ed0239323e8

Observation 5773f753-99df-4e71-819e-1462c9cd80d5 · outbound

This paper cites Sergey Chalnev, Michael Siu, and Alexander Conmy.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Sergey Chalnev, Michael Siu, and Alexander Conmy

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:27.323585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T19:34:26.904960Z digest=sha256:f4fdde2bacde9776f877eecef74411fca332f38ac1e9fae40105d524e520ba60

Observation 54764d6d-e786-41f5-ac49-97dc4ff1bbaf · outbound

This paper cites an unresolved cited work.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:27.335785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T19:34:26.900107Z digest=sha256:a5f943bd36dcf25007d7d1b255a867473eaee1d896b22ad8e04afed9468adaec

Pith citing papers

Observation f8c0b27d-97a8-4cc4-89bc-2fe2adc6e762 · inbound

EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models cites this paper.

EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:35:04.713574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:35:04.713574Z digest=sha256:267a999f10d3a991e6c8a8bf08b6d84166a55a51bc1d6a6d3886b4fd03a8e088

Observation 9594690c-bd84-4876-9545-2840a176f7c3 · inbound

Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models cites this paper.

Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T21:45:40.619559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T21:44:36.351517Z digest=sha256:1646773907a4e1f0e8b8d29bf2127fb5ccf71b46c5bdedf9bb90d8481485df50

Observation 8c32c189-ad64-4b1d-b692-2cb0ceaa4b47 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.869974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:a0da2374a4ee46babd4ed83878ed1f69bcb6461a06d0aa14aa47b3e958697f3f

Observation b0e6d66a-db95-425d-95a8-a6c9b7bcc935 · inbound

On Emotion-Sensitive Decision Making of Small Language Model Agents cites this paper.

On Emotion-Sensitive Decision Making of Small Language Model Agents Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:50.716346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:58:10.190022Z digest=sha256:cba53b27ee58ed4ca078b9bb0379ede287f683b55fba9feb6877616501c63ece

Observation 0aa614dc-950f-4a7e-8cec-070917eb79ce · inbound

Continuous Interpretive Steering for Scalar Diversity cites this paper.

Continuous Interpretive Steering for Scalar Diversity Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:00.779131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:08:02.296583Z digest=sha256:02cd28811b8b7ee62c335ffd5c5a161808e950e364b963e5965d920a3da02132

Observation 5107464f-c7c5-472f-93ff-fb504251f646 · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:19:26.671414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:6628dd1b2f4fc1e7a9ad7efdd201ec4df7826d5a0e02015098155caa37af971d

Observation 11e3bc4f-1a24-46b6-bc62-fe1c63501e4a · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:05:06.043246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T21:57:18.752878Z digest=sha256:f79bb1c61f82ed5536566d7b5059bec126f4c27a7835c195d4ab0d93f35267c7

Observation d5bdb653-e02c-48f3-956e-f1f4ea54c001 · inbound

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models cites this paper.

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:25:23.812858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-25T06:22:09.828982Z digest=sha256:3da77af4fb2370af90f03c9073372d5aff8f654b7e97ac36bd4895b2c564eac0

Observation e3a9eb89-7f3b-4027-b856-c8cfc60c22d7 · inbound

Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection cites this paper.

Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T05:36:39.347016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-25T05:35:04.688774Z digest=sha256:46f5656dce9b2b0a1bc162226dd2c1af573f8b0f2044eab26fd7d63ef2da99fb

Observation d127cb2b-5642-4137-bfe6-4e544d8c4e5e · inbound

Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs cites this paper.

Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:52:35.500614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T18:49:56.917505Z digest=sha256:a301d97f2ecc200f9567c30e3f2b2ec7168a314d36799a89f2ae267f67fda21b

Observation f33bc7f5-1f00-42d3-bfaf-b0c109102fe9 · inbound

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning cites this paper.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.052467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:23e8bc31bf10a995ebe7a8e0a5b6757bde234335e76870fcc187e8e9ca83001f

Observation fcf235cf-9bf0-4e30-9026-1235a3239dac · inbound

Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models cites this paper.

Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:15.583725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:31:15.583725Z digest=sha256:d7b79ffc57b6388c1904dfd94828d900bd576a9821988816f14c2cbddb30ecc3

Observation dbf03d6d-aaf7-4f97-bf81-bbf121d14c5c · inbound

Measuring Semantic Abstractness of SAE Features via Nonlocality cites this paper.

Measuring Semantic Abstractness of SAE Features via Nonlocality Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:22:47.712287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:22:47.712287Z digest=sha256:46985956f5cd5e297c66de8c93eebbeb30d7c8f496ed6764fc105eab4cfa1efe