Pith. sign in

Paper Citation Record · LEDGER

Scaling laws for activation steering with Llama 2 models and refusal mechanisms

As of 17 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2507.11771.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11771 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:05:04.820314Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:25:43.377432Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:25:43.489612Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4047946d-31b8-4f67-8717-90c8410077c6 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Mechanistic Interpretability for AI Safety -- A Review

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:03.622889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:03.622889Z digest=sha256:25fff2bcff6c14c9d9528aedc81195f3cb0ce78ae7f453c03a45e102d631990e

Observation 62cb6a61-d972-474a-b3f4-cee12f636adc · outbound

This paper cites A toy model of universality: Reverse engineering how networks learn group operations.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms A toy model of universality: Reverse engineering how networks learn group operations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:05.434920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:05:03.666165Z digest=sha256:8a87bc5cdb34cf808daf55bb0a8445940e8c19eb5199eb0d4ce9aaaf9da98578

Observation 601b57e4-96ce-48b6-a9da-67fd7168c73d · outbound

This paper cites Towards automated circuit discovery for mechanistic interpretability.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Towards automated circuit discovery for mechanistic interpretability

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:03.761144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:03.761144Z digest=sha256:0ecd14e8c6e492b90eff6b0b49a4acd39b009c43730ba8fe60dfc1c4e4250a2e

Observation 90722fc7-0fda-4edf-a2bb-16fcfada84b8 · outbound

This paper cites Toy Models of Superposition.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Toy Models of Superposition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:03.845054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:03.845054Z digest=sha256:589d1e841be08714837141631c8d9097b0b3114d6679411854cf07c25a76aa74

Observation f7a901bc-7b60-4059-94c4-2b8109e29806 · outbound

This paper cites Finding Neurons in a Haystack: Case Studies with Sparse Probing.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Finding Neurons in a Haystack: Case Studies with Sparse Probing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:03.916915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:03.916915Z digest=sha256:70762caa8eea2b2ea06c77029398d75128ec5dbf9987e67a9d691d4d3d95853d

Observation 14e88341-cb6e-43bd-92b4-c920351d5f2d · outbound

This paper cites Datamodels: Predicting Predictions from Training Data.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Datamodels: Predicting Predictions from Training Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:03.982552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:03.982552Z digest=sha256:f59e3e5c1d084dfa87ff7f1fbaa1d496e981873380e54e6f330955effe0521d1

Observation df3b9283-6c2b-4e5d-8145-24e472089997 · outbound

This paper cites TRAK: Attributing Model Behavior at Scale.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms TRAK: Attributing Model Behavior at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:04.041203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:04.041203Z digest=sha256:f2cec7b42fce1b6b89fea333ced3b82e4afb804a2cdf587a979cbd75267679a2

Observation 5f63daaa-e723-46f2-9e5f-6a78cfe93ced · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Steering Llama 2 via Contrastive Activation Addition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:04.150422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:04.150422Z digest=sha256:a02f7dc7bb13a1bbb6befd5b5b7207e9aec2fac7ede69c6f71d11f6055db5e3b

Observation 253994b4-5436-405c-87b1-8615443815ef · outbound

This paper cites M., Ilyas, A., and Madry, A.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms M., Ilyas, A., and Madry, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:05.297143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:05:04.258595Z digest=sha256:6b85a5b6941fc0f9f48b3a13eee9667649cac91a845ebaf03f47cb5376501e4b

Observation 7b4fcd37-a3f9-40dd-880e-4f48ea5a59ac · outbound

This paper cites Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:05.162627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:05:04.338235Z digest=sha256:82cac0f7b081c29e63d6146ef1ffdd2ecee5fe3f4f0197e74a94611b1dd05324

Observation d045c2ef-959a-4761-9493-bba4b61e0933 · outbound

This paper cites Steering Language Models With Activation Engineering.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Steering Language Models With Activation Engineering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:04.423071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:04.423071Z digest=sha256:7f4b6ad1d521e17ea9a5f3ca039e55a440852b5e079680ea89b8c728700de04f

Observation 2201a38b-4fcd-4c41-8296-90544649bed1 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:04.509549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:04.509549Z digest=sha256:44c3a9ce5a2e22d57fbaffb65a604a971b5de1e231a4af9296bc65f075f4cf31

Observation acd26905-b1d2-4e65-a081-912f6b562194 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:04.599235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:04.599235Z digest=sha256:7d387483a073653dbf8a0fbe6a004e0ebf4dd3fcd7d954f039e9e9876e09a44e

Observation b9d183bf-d72a-4cdf-9fbd-c045b2f050d9 · outbound

This paper cites Knowledge Conflicts for LLMs: A Survey.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Knowledge Conflicts for LLMs: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:04.694957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:04.694957Z digest=sha256:931928717a31f83ae5a91b7e340a9275cdf9e9d60d3c804f2b63eccacc175305

Observation 182fc9f9-0577-4c67-9158-41413fef0fbf · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:04.749036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:04.749036Z digest=sha256:2e4af2f06f0fd199f41115a07051cc7bb00991abc38aee273adb0a421b64780d

Observation 2462df79-733e-4f0a-8d58-4bd941ba5c39 · outbound

This paper cites write newline.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms write newline

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:04.820314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:04.820314Z digest=sha256:77362ff98f6badbad040ab6f05de11ae1f5515e408bcd2c7865b0832caf14b1c

Pith citing papers

Observation 472cee5e-8c34-4763-958d-f25e41e2c27e · inbound

When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs cites this paper.

When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs Scaling laws for activation steering with Llama 2 models and refusal mechanisms

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:25:43.493544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-12T00:25:43.377432Z digest=sha256:9faecad3493a480aa4411581a6c59dfb18726f873c61b5381b5a7db78fd56b00