MechaRule localizes sparse agonist neurons via contrastive hierarchical ablation and adaptive group testing to ground rule extraction, recalling 97% of high-effect activations at 2.14% cost while enabling near-total elimination of target behaviors.
McCormick, and David Madigan
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
verdicts
UNVERDICTED 2representative citing papers
Explores reference document choices for applying DeepSHAP to neural retrieval models and reports that its explanations differ substantially from those of LIME.
citing papers explorer
-
Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation
MechaRule localizes sparse agonist neurons via contrastive hierarchical ablation and adaptive group testing to ground rule extraction, recalling 97% of high-effect activations at 2.14% cost while enabling near-total elimination of target behaviors.
-
A study on the Interpretability of Neural Retrieval Models using DeepSHAP
Explores reference document choices for applying DeepSHAP to neural retrieval models and reports that its explanations differ substantially from those of LIME.