REVIEW 3 cited by
Automatic Discovery of Visual Circuits
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
To date, most discoveries of network subcomponents that implement human-interpretable computations in deep vision models have involved close study of single units and large amounts of human labor. We explore scalable methods for extracting the subgraph of a vision model's computational graph that underlies recognition of a specific visual concept. We introduce a new method for identifying these subgraphs: specifying a visual concept using a few examples, and then tracing the interdependence of neuron activations across layers, or their functional connectivity. We find that our approach extracts circuits that causally affect model output, and that editing these circuits can defend large pretrained models from adversarial attacks.
Forward citations
Cited by 3 Pith papers
-
Certified Circuits: Stability Guarantees for Mechanistic Circuits
Certified Circuits uses deletion-based randomized smoothing to guarantee that circuit components stay included or excluded under bounded edits to the concept dataset, yielding more compact and more accurate circuits.
-
Multimodal Function Vectors for Visual Relations
Multimodal function vectors extracted from a handful of attention heads in OpenFlamingo-4B encode spatial relations and can be steered, fine-tuned, and composed to improve zero-shot relational reasoning.
-
Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept Representations
GCC discovers multiple concept-specific neuron circuits per query by combining first-order ablation sensitivity with top-k activation overlap.
Discussion (0). Continue with ORCID to comment.