Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:34:26.980516Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 13 inbound Pith citation observations for arXiv:2501.09929.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:34:26.980516Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:35:04.713574Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
18 of 18 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation d87ac576-f7f9-4db6-9b55-bd6313842b14 · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Improving Steering Vectors by Targeting Sparse Autoencoder Features
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dae8fa78-5f68-49cc-b9f6-fde2532d1909 · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fb11239-5157-4fb7-8f87-384c9bf0e60c · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions <bos>I think this is a photo of a giant squid attacking a Russian submarine, and it is one of the most Incredible Aliens captured in Antarctica! These mind
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c6660f73-7181-4e0c-b75d-3f93070becb9 · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Kiho Park, Yo Joong Choe, and Victor Veitch
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9604e859-cb4e-4ed3-8088-fca4d0d0d53b · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions 10 Published at Building Trust Workshop at ICLR 2025 Gonçalo Paulo, Alex Mallen, Caden Juang, and Nora Belrose
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bf2712e3-c2a5-4c81-826a-ad5abe3ff685 · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Automatically Interpreting Millions of Features in Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f99a6f-72eb-473b-b2d3-27b888fcaef2 · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Gemma 2: Improving Open Language Models at a Practical Size
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2217090-5ff5-4b2b-80c3-0152c913626c · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Gemma 2: Improving Open Language Models at a Practical Size
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48140458-dea3-4ee5-ad4e-29366c4d5a3c · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Polysemanticity and Capacity in Neural Networks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173e66f3-3c14-4d4d-b6b7-bb39e76ef4fd · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Steering Language Models With Activation Engineering
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22ad1bad-899b-4aa7-a43a-2947b904b3c0 · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a22c022-9f4a-4eff-b2c3-42d62027012c · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 847f23d4-b30a-45cc-b955-88105f8c1a17 · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions The government is hiding the truth about alien contact
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9dbcc039-e3b8-4a68-938a-3451a94a6799 · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions me" in different contexts -1.039 2605 References to presence or absence of evidence Rollouts at Scale = 80 (Optimal Scale): 17 Published at Building Trust Workshop at ICLR 2025
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6dda3c0d-9531-41ce-b537-007f64dda116 · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Evaluate the text based on the criterion
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3a57c897-f4e2-4c75-b725-34e073feaf9b · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Aaron Gokaslan and Vanya Cohen
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5773f753-99df-4e71-819e-1462c9cd80d5 · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Sergey Chalnev, Michael Siu, and Alexander Conmy
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 54764d6d-e786-41f5-ac49-97dc4ff1bbaf · outbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f8c0b27d-97a8-4cc4-89bc-2fe2adc6e762 · inbound
EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9594690c-bd84-4876-9545-2840a176f7c3 · inbound
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8c32c189-ad64-4b1d-b692-2cb0ceaa4b47 · inbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 279
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b0e6d66a-db95-425d-95a8-a6c9b7bcc935 · inbound
On Emotion-Sensitive Decision Making of Small Language Model Agents Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0aa614dc-950f-4a7e-8cec-070917eb79ce · inbound
Continuous Interpretive Steering for Scalar Diversity Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5107464f-c7c5-472f-93ff-fb504251f646 · inbound
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 11e3bc4f-1a24-46b6-bc62-fe1c63501e4a · inbound
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d5bdb653-e02c-48f3-956e-f1f4ea54c001 · inbound
Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e3a9eb89-7f3b-4027-b856-c8cfc60c22d7 · inbound
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d127cb2b-5642-4137-bfe6-4e544d8c4e5e · inbound
Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f33bc7f5-1f00-42d3-bfaf-b0c109102fe9 · inbound
Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fcf235cf-9bf0-4e30-9026-1235a3239dac · inbound
Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf03d6d-aaf7-4f97-bf81-bbf121d14c5c · inbound
Measuring Semantic Abstractness of SAE Features via Nonlocality Interpretable Steering of Large Language Models with Feature Guided Activation Additions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.