Pith. sign in

REVIEW 1 cited by

Controlling Language and Diffusion Models by Transporting Activations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.23054 v2 pith:LTJG3RMF submitted 2024-10-30 cs.LG cs.AIcs.CLcs.CV

classification cs.LGcs.AIcs.CLcs.CV
keywords modelmodelsactivationscontrolconceptsdiffusioneffectivelyfine-grained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing capabilities of large generative models and their ever more widespread deployment have raised concerns about their reliability, safety, and potential misuse. To address these issues, recent works have proposed to control model generation by steering model activations in order to effectively induce or prevent the emergence of concepts or behaviors in the generated output. In this paper we introduce Activation Transport (AcT), a general framework to steer activations guided by optimal transport theory that generalizes many previous activation-steering works. AcT is modality-agnostic and provides fine-grained control over the model behavior with negligible computational overhead, while minimally impacting model abilities. We experimentally show the effectiveness and versatility of our approach by addressing key challenges in large language models (LLMs) and text-to-image diffusion models (T2Is). For LLMs, we show that AcT can effectively mitigate toxicity, induce arbitrary concepts, and increase their truthfulness. In T2Is, we show how AcT enables fine-grained style control and concept negation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment

    cs.LG 2026-04 accept novelty 6.0 of 10

    Learning an input-conditioned mapping from embeddings to the best single steering layer (W2S) consistently beats fixed-layer CAA and L2S on 13 behaviors for two LLMs, in- and out-of-distribution.

Pith tools