Pith. sign in

Paper Citation Record · LEDGER

On the Limitations of Steering in Language Model Alignment

As of 22 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2505.01162.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01162 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:28:15.771544Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c05723da-31b9-471a-98b9-1656e057ca42 · outbound

This paper cites doi: 10.1162/coli a 00524.

On the Limitations of Steering in Language Model Alignment doi: 10.1162/coli a 00524

Reference 2

Resolution
malformed identifier
no resolver link, observed 2026-08-16T04:28:15.711620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.711620Z digest=sha256:9e9499d1c5d84d697e8fea253b6c9b95df0c81022dc8e7d9ec9aa5fa033882f0

Observation 4515535e-a1c2-4bb0-b119-0f5fa459dc6c · outbound

This paper cites Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.

On the Limitations of Steering in Language Model Alignment Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.717431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.717431Z digest=sha256:6854d89f9d0a990f53ccfdcd6bab1661343a5e229cc400887bec684967a3f116

Observation 1192d930-f012-4348-9a68-32f66b94a243 · outbound

This paper cites GPT-4 Technical Report.

On the Limitations of Steering in Language Model Alignment GPT-4 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.733955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.733955Z digest=sha256:5678c0dce9c4861f3ceaae09094891c83c8b89367fb7e0d034eb8d9ea667d60a

Observation 398dc27f-9d53-4db4-8c3a-33d8e5365e88 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

On the Limitations of Steering in Language Model Alignment Steering Llama 2 via Contrastive Activation Addition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.740063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.740063Z digest=sha256:0b6951d75f25cdcd0f2b5856512b4852e8824d380fe1f3a05e9f7a9a36c6223f

Observation 241c311f-3d57-4722-8623-ca7a38771e41 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

On the Limitations of Steering in Language Model Alignment Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.746137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.746137Z digest=sha256:412114686b1ff8d3cf87298b8238ba31d184821d3b707ecda36148ee0b220725

Observation 8b7f8c90-dd6e-4271-94a0-89c99fd1d291 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

On the Limitations of Steering in Language Model Alignment Linear Representations of Sentiment in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.759209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.759209Z digest=sha256:4c992b57b2213f9429d888739c20ad48f03235451f64a6f611e131381cfa8433

Observation f22ec94c-9dbf-4981-86c2-00301e31649c · outbound

This paper cites Steering Language Models With Activation Engineering.

On the Limitations of Steering in Language Model Alignment Steering Language Models With Activation Engineering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.765005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.765005Z digest=sha256:139409c120d5e3a2966ac5025a27037801ee931df346737714d963be2d1de92a

Observation e1567efb-18eb-43b9-914e-45dd97395275 · outbound

This paper cites Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability.

On the Limitations of Steering in Language Model Alignment Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.771544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.771544Z digest=sha256:83e279f398637f1da7b852fa48643b3fe4967c8f2d7dfadb1dd12966a568a088

Observation 3da1400d-4fe1-460d-b28c-821007d36f2a · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

On the Limitations of Steering in Language Model Alignment Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.727690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.727690Z digest=sha256:81aa04d310c9056d470a49d3fafa3e990ba1de0004438a21779aed470c9acced

Observation 26caa330-1b9d-4c81-bdb1-de15e44dad8a · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

On the Limitations of Steering in Language Model Alignment Refusal in Language Models Is Mediated by a Single Direction

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.705041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.705041Z digest=sha256:50bc11e067be57dd46a7eca0ff8144864e156eeb6b1118e91df0911b6b75d943

Observation 7e38e289-cb8c-403c-bb84-12372feb1afe · outbound

This paper cites Analyzing the Generalization and Reliability of Steering Vectors.

On the Limitations of Steering in Language Model Alignment Analyzing the Generalization and Reliability of Steering Vectors

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.752786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.752786Z digest=sha256:c8c27e4a9154cac1a4c4499add3761b964f9d86bc48e00d13ce31118491e3b70

Pith citing papers

No inbound Pith citation observations are available.