Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:43:14.801525Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2501.15054.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:43:14.801525Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:59:30.298880Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T18:59:33.160545Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c1f030de-aa82-4686-8c78-12fe3ae90242 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Neel Nanda
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 268e257b-7345-4887-b610-8562bc4a4508 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models A review of taxonomies of explainable artificial intelligence (xai) methods
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6729c272-5017-4a3e-9ebd-bec9fcd2f384 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Mechanistic Interpretability for AI Safety -- A Review
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d45590e-84b4-42de-b444-628aa8166c6b · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac73ebc-e5f0-4645-b3e4-e0a97897b17c · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 810515a5-3cf0-4547-ab2a-799b453335be · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Localizing Model Behavior with Path Patching
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f2b69b-d54c-4439-b631-24a4305e28c3 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fc23120f-5abd-491a-84ea-8dcc3276034f · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ba1d4b-b0f9-4cae-a9da-d4d1ec3d2dab · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7759ce2-4653-44f3-960e-f793910915cc · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Emergent Abilities of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 708a0b8d-41c3-4d26-896f-e0fd0c4471ad · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Progress measures for grokking via mechanistic interpretability
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4382c841-0db5-4bfa-9031-dfca2a9091ca · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a802a3-b2a4-4638-a89a-146c3af9cfb7 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Locating and Editing Factual Associations in GPT
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 860676fe-f804-4626-9fbf-632dacd16e3a · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models A Survey of Machine Unlearning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f97076-180f-4218-8226-8aef73db19aa · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Circumventing interpretability: How to defeat mind-readers
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fee2dd8-deeb-431c-ac2d-e6ecb3d8d85e · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 599be8c9-332f-4cc9-b99f-d2ca329015c0 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Representation Engineering: A Top-Down Approach to AI Transparency
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c92d42-2d8c-43e8-9bf8-053b56e0a0dc · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models AI Deception: A Survey of Examples, Risks, and Potential Solutions
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d41b6da4-6911-4cac-9ea4-44a829bda27e · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models An overview of 11 proposals for building safe advanced AI
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cefe1f41-dc8d-4d9e-adfa-8be2419a5485 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Neel Nanda
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f5ceadc-785c-4c56-9ca3-dcfbfba038f4 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2552a8a-796f-4dd6-bef8-6f489b7c0018 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Lee Sharkey, Sid Black, and beren
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9015422d-b7ad-4ce4-8101-07aa05a192a1 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Understanding intermediate layers using linear classifier probes
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d45d8511-7f54-48d9-be16-a96a4dc080de · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models X-Risk Analysis for AI Research
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa06af59-535d-40a7-a8f9-58b83639b142 · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models An Overview of Catastrophic AI Risks
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d51a07f8-32d9-4b56-a357-a4e17aba6e6f · outbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models AI Alignment: A Comprehensive Survey
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43b03796-5d94-4077-addf-bed3abaabb52 · inbound
On the Effect of Uncertainty on Layer-wise Inference Dynamics Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.