Malware-detection behavior in three instruction-tuned LLMs is concentrated in model-specific FFN layer bands and can be causally boosted or collapsed by scaling the attributed facilitating and inhibiting neurons.
19 Mukund Sundararajan, Ankur Taly, and Qiqi Yan
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Which Neurons Detect Malicious Code? A Probing Study of LLM Security Knowledge
Malware-detection behavior in three instruction-tuned LLMs is concentrated in model-specific FFN layer bands and can be causally boosted or collapsed by scaling the attributed facilitating and inhibiting neurons.