Pith. sign in

Paper Citation Record · LEDGER

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2502.05945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05945 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:23:02.436296Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcb25acb-4414-4901-a6da-fc94dc3c99f4 · outbound

This paper cites AI coordination.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models AI coordination

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:23:02.705955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:23:02.426157Z digest=sha256:e2e7e42a13904c6228a2da53212071a6d16553d79747dc2f872c30100197941c

Observation c79286d2-e8bf-4dca-bc5e-22518a51632a · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:23:02.390992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:23:02.390992Z digest=sha256:bb3ed99fd97e0b6aa5e699a0b2ab70d6b3cb0f06f504dc02ee28f3b246945013

Observation e6ab7f7d-2ddb-4d9c-ab30-0d9aad046df1 · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T17:23:02.396046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:23:02.396046Z digest=sha256:c6342f6f423cfd0cf45bbe711264d50ede9c6d6ec833e307d18bcd936476e689

Observation 49968655-2551-40e0-aa83-2bd47fff95a5 · outbound

This paper cites Accessed: 2024-05-22.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models Accessed: 2024-05-22

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:23:02.737789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:23:02.401637Z digest=sha256:5101cdcd4a3cc94666522417c5121b1983e0f148df2c5ef2f24a5cbf154ac6ae

Observation 9e07f4ef-4507-46dc-8776-473692323ce3 · outbound

This paper cites doi: 10.18653/v1/2024.acl-long.828.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models doi: 10.18653/v1/2024.acl-long.828

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T17:23:02.406476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:23:02.406476Z digest=sha256:51b1850d4f9b5df0148c27df21c4e7f1b63aa924006ae70fd50ed423bc1edaa0

Observation 356fc040-2a4f-4a07-97db-d08ac3b1d959 · outbound

This paper cites Adaptive activation steering: A tuning-free llm truthfulness improvement method for di- verse hallucinations categories.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models Adaptive activation steering: A tuning-free llm truthfulness improvement method for di- verse hallucinations categories

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T17:23:02.417268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:23:02.417268Z digest=sha256:29033b3146bef9cc77186a17b1ab863417509de524c2e2e32ab6a82c55b9ac84

Observation 1c56fcf8-ea38-46ce-8c3a-99424e69293d · outbound

This paper cites Answer: (A) / (B).

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models Answer: (A) / (B)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:23:02.722261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:23:02.421863Z digest=sha256:af0cb4be65bc5d3915aa13e23c6bd2d8ded29ccd1c8908eb51443d5e5cfddd96

Observation 0098616a-80b9-4103-8baa-38190c04deb2 · outbound

This paper cites AI coordination.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models AI coordination

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:23:02.674911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:23:02.436296Z digest=sha256:55bccd7303869730b386854a6e13a495f71e8cdfa3d892057a9f5d1b2fedb9a4

Observation ca98ba4e-69c9-41b5-ba0c-d65b8e1c262b · outbound

This paper cites AI coordination.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models AI coordination

Reference 125

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:23:02.690535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:23:02.431041Z digest=sha256:bbf1cc853f56dc88854e419cf80dc90f151dfc0e1ed02f021fe151ab35242627

Observation 361c525c-c4ee-4040-813c-ead88b017d5b · outbound

This paper cites LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T17:23:02.411676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:23:02.411676Z digest=sha256:528d46dcd331188ed3934742ad652f2095c0ca4ddac21bf48d80ed503da1e89a

Observation 4be989f5-6a69-4353-a3ae-5fde75fdffd4 · outbound

This paper cites Will releasing the weights of future large language models grant widespread access to pandemic agents?.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models Will releasing the weights of future large language models grant widespread access to pandemic agents?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T17:23:02.386267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:23:02.386267Z digest=sha256:3d3e9ea1cbfe6ccb67ecbacf5532db1611c7f41036cb77b5041b2c5b16f3d74f

Observation d763bcab-a3b5-4ba5-9bc2-04b545cc6f05 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T17:23:02.374370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:23:02.374370Z digest=sha256:952869984ff84e1841a5a8b8d6006f28875fc7198da75267e268dd497f788345

Pith citing papers

No inbound Pith citation observations are available.