Pith. sign in

REVIEW 6 cited by

Controlling Language and Diffusion Models by Transporting Activations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.23054 v2 pith:LTJG3RMF submitted 2024-10-30 cs.LG cs.AIcs.CLcs.CV

Controlling Language and Diffusion Models by Transporting Activations

classification cs.LG cs.AIcs.CLcs.CV
keywords modelmodelsactivationscontrolconceptsdiffusioneffectivelyfine-grained
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The increasing capabilities of large generative models and their ever more widespread deployment have raised concerns about their reliability, safety, and potential misuse. To address these issues, recent works have proposed to control model generation by steering model activations in order to effectively induce or prevent the emergence of concepts or behaviors in the generated output. In this paper we introduce Activation Transport (AcT), a general framework to steer activations guided by optimal transport theory that generalizes many previous activation-steering works. AcT is modality-agnostic and provides fine-grained control over the model behavior with negligible computational overhead, while minimally impacting model abilities. We experimentally show the effectiveness and versatility of our approach by addressing key challenges in large language models (LLMs) and text-to-image diffusion models (T2Is). For LLMs, we show that AcT can effectively mitigate toxicity, induce arbitrary concepts, and increase their truthfulness. In T2Is, we show how AcT enables fine-grained style control and concept negation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control

    cs.LG 2026-06 unverdicted novelty 7.0

    LA-LQR applies latent-space linear-quadratic regulator control to steer text-to-video model activations toward desired features while penalizing excessive changes.

  2. Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control

    cs.LG 2026-04 conditional novelty 7.0

    Local linearity of LLM layers enables LQR-based closed-loop activation steering with theoretical tracking guarantees.

  3. Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution

    cs.LG 2025-02 unverdicted novelty 7.0

    Neurons exhibit concept-conditioned activation ranges forming Gaussian-like distributions with minimal overlap, and range-based interventions via NeuronLens outperform neuron-level masking in targeted manipulation wit...

  4. CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification

    cs.CL 2026-04 unverdicted novelty 6.0

    CausalDetox identifies minimal attention heads causally linked to toxicity via Probability of Necessity and Sufficiency, then applies targeted inference-time steering or fine-tuning to reduce toxic generation while pr...

  5. Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment

    cs.LG 2026-04 accept novelty 6.0

    Learning an input-conditioned mapping from embeddings to the best single steering layer (W2S) consistently beats fixed-layer CAA and L2S on 13 behaviors for two LLMs, in- and out-of-distribution.

  6. SHIFT: Steering Hidden Intermediates in Flow Transformers

    cs.CV 2026-04 unverdicted novelty 5.0

    SHIFT learns and applies steering vectors to selected layers and timesteps in DiT models to suppress concepts, shift styles, or bias objects while keeping image quality and prompt adherence intact.