REVIEW 3 cited by
Controlling Large Language Models Through Concept Activation Vectors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As large language models (LLMs) are widely deployed across various domains, the ability to control their generated outputs has become more critical. This control involves aligning LLMs outputs with human values and ethical principles or customizing LLMs on specific topics or styles for individual users. Existing controlled generation methods either require significant computational resources and extensive trial-and-error or provide coarse-grained control. In this paper, we propose Generation with Concept Activation Vector (GCAV), a lightweight model control framework that ensures accurate control without requiring resource-extensive fine-tuning. Specifically, GCAV first trains a concept activation vector for specified concepts to be controlled, such as toxicity. During inference, GCAV steers the concept vector in LLMs, for example, by removing the toxicity concept vector from the activation layers. Control experiments from different perspectives, including toxicity reduction, sentiment control, linguistic style, and topic control, demonstrate that our framework achieves state-of-the-art performance with granular control, allowing for fine-grained adjustments of both the steering layers and the steering magnitudes for individual samples.
Forward citations
Cited by 3 Pith papers
-
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
A single adversarially optimized image can reproduce activation-steering behavior in multiple VLMs and partially transfer to unseen models.
-
Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
ConVA identifies value-specific directions in an LLM's internal activations from context-matched GPT-4o-generated examples and gates minimal activation steering to control outputs across ten Schwartz values.
-
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
BiasLens uses concept activation vectors and sparse autoencoders to estimate LLM bias from internal representations, reporting moderate to strong agreement with behavioral bias metrics in a small evaluation.
Discussion (0). Continue with ORCID to comment.