{"id":"6fcac14f-6837-4a6e-a2e9-bf06e56dc5ac","arxiv_id":"2505.01372","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper introduces an Explanatory Virtues Framework and argues, via a qualitative rubric, that Compact Proofs are the most promising method for mechanistic interpretability.","lead":"This paper proposes a framework for judging which explanations of neural networks are better, drawing on virtues from the philosophy of science. It applies the framework to four interpretability methods and concludes that Compact Proofs satisfy the most virtues.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3's virtue definitions require P(x|E), but MI explanations are not specified as probabilistic models; Table 2's rubric replaces the promised canonical computation with subjective judgments, so the claimed systematic theory-choice procedure is not established.","rationale":"Reader's weakest assumption is truth-conduciveness. I agree that premise is asserted rather than argued, but the more decisive fragility is upstream: the formal apparatus is not connected to MI explanations. The abstract's 'systematic' claim presupposes a procedure for evaluating any explanation; the paper's own definitions require P(x|E), a probabilistic model. Since no such model is specified for circuits, SAEs, or Compact Proofs, the framework cannot actually be applied to the motivating disagreement. The paper's own rubric invites readers to 'check their intuitive understanding', signalling that subjective judgments drive Table 1. This is consistent with a CONDITIONAL verdict: the framework may become systematic once P(x|E) is operationalized and the rubric is derived from the formal definitions. It does not change the reader's conditional verdict.","tokens_in":22310,"tokens_out":5664,"duration_ms":63144,"concrete_test":"Take the two competing explanations of group multiplication from Wu et al. (2024) and ask the authors to (a) define P(x|E) for each explanation using only the machinery in the paper, (b) compute all §3 virtues numerically, and (c) have three independent raters apply Table 2 to the same explanations. If (a) requires additional assumptions not stated, or if (b) and (c) disagree on the ranking, the framework is not yet a systematic selection procedure. A weaker but still decisive check: state one general rule that maps any circuit or SAE explanation to a distribution over model inputs; if no such rule is supplied, the central claim should be downgraded to a heuristic taxonomy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the framework gives systematic, objective criteria for theory choice—fails at the point where the formal definitions meet the case studies. Every Bayesian virtue in §3.1 is defined through probabilities P(x|E) over data: Accuracy, Precision, Descriptiveness, Co-Explanation, Power, and Unification. Hard-to-Varyness in §3.3 uses log(Acc(E)). This presupposes that an MI explanation E is a probabilistic model. Yet none of the four methods in §4 is given as a distribution: a circuit is a structural causal model, an SAE explanation is a dictionary plus feature descriptions, and a Compact Proof is a verifier program. The paper never specifies how to obtain P(x|E) from these artifacts, so the promised \"canonical way to compute each virtue\" (p. 4) is not defined for any real MI explanation. The Table 2 rubric then substitutes subjective ✓/●/✗ judgments. As a symptom, Co-Explanation and Unification are the same formal quantity—log(P(xT|E)/∏P(x_i|E)) and its expectation—but Table 1 rates Compact Proofs ✓ on the former and ✗ on the latter, which is only possible because the rubric is not derived from §3. Thus, even granting truth-conduciveness (itself supported only by Schindler 2018), the framework does not yet provide a computable method for choosing between Chughtai et al. and Stander et al.; it provides a vocabulary for discussing explanations.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces an Explanatory Virtues Framework for evaluating explanations in mechanistic interpretability (MI), drawing on Bayesian, Kuhnian, Deutschian, and Nomological accounts from the philosophy of science. The authors define a set of virtues with mathematical formulas, apply the framework to four MI methods (Clustering, Sparse Autoencoders, Causal Circuits, and Compact Proofs), and conclude that Compact Proofs are a promising approach because they exhibit many virtues. The paper also proposes research directions emphasizing simplicity, unification, and nomological principles.","tokens_in":22676,"tokens_out":5507,"duration_ms":46908,"significance":"If the framework delivered what it promises—a canonical, computable way to compare competing MI explanations—it would be a valuable contribution to interpretability research, which currently lacks principled theory-choice criteria. The paper surveys relevant philosophy of science literature, and its taxonomy of virtues is genuinely useful as a vocabulary for discussing explanations. However, the paper does not provide machine-checked proofs, reproducible code, or quantitative data; the central evaluation in Table 1 is based on the authors' qualitative judgments. The paper's main value is therefore as a conceptual proposal, not as an established evaluation methodology.","major_comments":[{"comment":"The definitions of Accuracy, Precision, Descriptiveness, Co-Explanation, Power, and Unification all presuppose that an explanation E is associated with a conditional probability P(x|E) over observational data. Yet none of the four methods analyzed in Section 4 is specified as a probabilistic model: a circuit is a structural causal model, an SAE is a dictionary plus feature descriptions, and a Compact Proof is a verifier program. The paper never states how to derive P(x|E) from these artifacts, so the promised 'consistent and canonical way to compute each virtue' (Section 3, p. 4) is not defined for the very explanations the framework is meant to compare.","section":"Section 3.1, glossary of Bayesian virtues"},{"comment":"The framework's formal definitions imply that Co-Explanation and Unification are the same quantity—CoEx(E) = log(Acc(E)) − Desc(E) and Unif(E) = Prec(E) − Power(E) = E_{xT∼X}[log(P(xT|E)/∏P(xT,i|E))], which are identical up to the expectation. Yet Table 1 rates Compact Proofs as '✓' on Co-explanation and '✗' on Unification. This inconsistency is only possible because the Table 2 rubric is not derived from the formal definitions; it shows that the case-study ratings are subjective judgments rather than outcomes of the framework.","section":"Table 1 vs. Section 3.1"},{"comment":"The central conclusion that 'Compact Proofs consider many explanatory virtues and are hence a promising approach' (Abstract) is based entirely on the authors' own qualitative ratings in Table 1, not on any measured data or inter-rater agreement. The paper provides no evidence that another researcher applying the Table 2 rubric would reach the same ratings, so the claimed systematic comparison between Chughtai et al. (2023) and Stander et al. (2024) is not actually demonstrated.","section":"Section 4.2 and Abstract"},{"comment":"The definition of valid MI explanations (Model-level, Ontic, Causal-Mechanistic, Falsifiable) is imported from the authors' own forthcoming paper (Ayonrinde & Jaburi, 2025), which is not available for verification. Because this definition is the foundation upon which the Explanatory Virtues Framework is built, the framework's epistemic grounding is self-referential; the paper should either provide the argument for this definition in the present manuscript or cite a published, accessible source.","section":"Section 2.2"},{"comment":"The truth-conduciveness of the listed virtues is asserted with only a reference to Schindler (2018). This premise is load-bearing: if virtues are not truth-conducive, the framework becomes a subjective preference list. The paper should give at least one concrete argument or empirical example for why these particular virtues indicate truth, or explicitly reframe the contribution as a descriptive account of explanatory values.","section":"Section 3, 'Explanatory Virtues are properties that are reliable indicators of truth'"},{"comment":"The formalization of hard-to-varyness as a local maximum of hv(E) = log(Acc(E)) − k(E) is under-specified because 'local' is only informally defined as 'a small number of edit operations apart' (footnote 9); without a metric on the space of explanations, the definition cannot be applied to decide whether a given explanation is hard-to-vary, which is exactly what Table 1 claims.","section":"Section 3.3, Hard-to-Varyness definition"}],"minor_comments":[{"comment":"The subscript formatting is inconsistent: the text uses 'x T', 'x I', and 'xT,i' in different places, and the glossary uses 'x T∼X' without clearly defining the distribution over inference-time data; please standardize the notation.","section":"Section 3.1, notation"},{"comment":"There is a typo: 'Sparse Auutoencoder' should be 'Sparse Autoencoder'.","section":"Section 3.2, Pragmatic Utility paragraph"},{"comment":"The FCM criteria use expressions F(C\\K) and F(M\\K) without defining the function F; presumably F is the model output function on a data distribution, but this should be stated explicitly.","section":"Section 4.1.3, FCM criteria"},{"comment":"The caption for Figure 1 references colors, bold arrows, and dashed arrows, but the figure is not included in the manuscript text; please ensure the figure is present or describe the relationships in the text.","section":"Section 3.5 and Figure 1"},{"comment":"Section 4.2 describes Compact Proofs as a method for evaluating other explanations, yet Table 1 treats Compact Proofs as an explanation method comparable to Clustering, SAEs, and Circuits; please clarify this distinction or reclassify the comparison.","section":"Section 4.2 and Table 1"},{"comment":"The claim that 'the MI community has sought to understand the universality ... with mixed results' lacks a citation supporting the 'mixed results' assertion; please add a reference or soften the claim.","section":"Section 5, paragraph on universality"},{"comment":"Several references are to unpublished or forthcoming works (Ayonrinde & Jaburi 2025; Jaburi et al. 2025; Ayonrinde 2025), making it difficult for readers to verify the cited claims; consider providing preprints or including the relevant content in an appendix.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's reliance on the authors' own forthcoming work (Ayonrinde & Jaburi 2025; Jaburi et al. 2025) for foundational definitions and for the evaluation of Compact Proofs—a method co-authored by one of the present authors—creates a perception of self-referential validation that the authors should address explicitly (e.g., by including the relevant definitions in the appendix or by having independent researchers apply the rubric). The paper's scope is primarily philosophical; the editors may wish to consider whether the contribution's fit with a CS/LG venue is best served by a framing as a position paper rather than as a systematic evaluation framework."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a synthesis paper, not a measurement paper. It borrows from Wojtowicz and DeDeo, Kuhn, Deutsch, and Hempel to build a list of explanatory virtues and applies that list to four MI methods. The goal is to give the field a common way to say why one mechanistic explanation is better than another. That is a real need.\n\nWhat is genuinely new is the application, not the pieces. Putting Accuracy, Precision, Simplicity, Unification, Hard-to-varyness, Nomologicity into one glossary, and sketching how Clustering, SAEs, Circuits, and Compact Proofs score on them, is a useful organizing act. The distinction between explanatory virtues (truth-conducive, normative) and explanatory values (what practitioners actually prize) is a helpful frame. The MDL-SAEs discussion, where Parsimony and Conciseness pull apart, is a good example.\n\nThe soft spots are load-bearing. The Section 3 definitions are all in terms of P(x|E), but no MI explanation in Section 4 is given as a probabilistic model. A circuit is a structural causal model, an SAE is a dictionary, a Compact Proof is a verifier program. The paper never says where P(x|E) comes from. So the formal definitions cannot be computed for any actual method. The Table 2 rubric steps in with qualitative judgments, but those judgments are not derived from the definitions. As the stress-test notes, Co-Explanation and Unification are the same formal object—one is the expectation of the other—yet Table 1 gives Compact Proofs different marks on them. (The stress-test says ✓ for Co-Explanation; the table actually shows ●, but the point stands.) The 'systematic' theory-choice procedure the abstract promises is not actually operational.\n\nThe Compact Proofs conclusion is therefore under-supported. The table's ratings are the authors' own, and the prose gives reasons that are mostly assertions about what methods 'can' do. Co-author Jaburi has a Compact Proofs paper, so the direction aligns with the authors' program; that is not disqualifying, but it means the evaluation needs to be independently auditable. The validity conditions for MI come from the authors' own forthcoming Part I.i, which makes the foundation self-referential until that paper is public.\n\nThe truth-conduciveness premise is asserted and supported only by a citation to Schindler. That is a legitimate philosophical position, but it is doing a lot of work; if the virtues are not truth-conducive, the framework becomes a preference list. The paper should at least say how contested that premise is.\n\nWho this is for: people who want a shared vocabulary for MI evaluation, especially in workshops and position papers. It is not yet a tool a practitioner can run to choose between Chughtai and Stander. Does it deserve a serious referee? Yes. The paper is clear, cites the relevant literature, and the synthesis could be a foundation for later operationalization. A referee should push hard on the gap between the math and the rubric, and ask the authors to either make the framework computable or honestly present it as a philosophical vocabulary.","headline":"A useful philosophical vocabulary for MI evaluation, but not yet the systematic framework it claims to be; the Compact Proofs verdict rests on a subjective rubric.","tokens_in":23139,"tokens_out":5935,"would_cite":false,"duration_ms":54142,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that a pluralist Explanatory Virtues Framework gives systematic criteria for choosing between competing mechanistic interpretability explanations, and that Compact Proofs, which embody many of these virtues, are a…","keywords":["mechanistic interpretability","explanatory virtues","theory choice","philosophy of science","explanation evaluation","compact proofs","sparse autoencoders","causal circuits"],"falsifier":"Take one model, such as a small transformer trained to compose group operations, generate two valid MI explanations of the same behaviour that score very differently on the virtues, then run a battery of held-out intervention experiments that the two explanations predict differently; if the lower-scoring explanation predicts the held-out behaviour at least as well as the high-scoring one, the claim that the virtues are truth-conducive is falsified.","tokens_in":22129,"feed_emoji":"🧠","tokens_out":14483,"duration_ms":124882,"temperature":0.7,"pith_summary":"The paper seeks to establish that mechanistic interpretability can settle disputes between competing explanations of the same neural network by importing criteria for theory choice from the philosophy of science. It proposes a pluralist Explanatory Virtues Framework: a battery of formalisable properties—accuracy, precision, simplicity, unification, fruitfulness, hard-to-varyness, nomologicity, and others—that are claimed to be reliable indicators of whether an explanation is true. Applied to four common methods, the framework finds that clustering explanations, sparse autoencoder decompositions, and causal-circuit analyses each neglect one or more virtues, while Compact Proofs consider many and are therefore singled out as promising. If the framework is right, interpretability researchers gain an epistemic basis for preferring one explanation over another, and three research directions follow: clarifying what simplicity means, prioritising unifying and co-explanatory accounts, and seeking universal principles about neural networks.","feed_headline":"Compact proofs outrank rivals on AI explanation scorecard","feed_subtitle":"A philosophy-of-science framework gives criteria for choosing between competing accounts of neural-network behavior.","key_machinery":"The central object is the Explanatory Virtues Framework itself, built from the Bayesian, Kuhnian, Deutschian, and Nomological accounts of explanation: a directed graph of virtues, each with a mathematical definition, that turns theory choice into a comparison problem. Its machinery includes Bayesian decompositions—$\\mathrm{Acc}(E)=P(x_T\\mid E)$ for accuracy, precision as expected log-likelihood, descriptiveness and co-explanation as additive and joint components of log-likelihood, power and unification as their theoretical counterparts—and a hard-to-varyness condition: an explanation is hard-to-vary if it sits at a local maximum of $\\log(\\mathrm{Acc}(E))-k(E)$, where $k$ is a complexity measure and a modification is a sequence of insertions, deletions, substitutions, or transpositions of symbols. The other load-bearing piece is the Compact Proofs evaluator, in which an explanation is converted into a verifier $V(\\theta,E)$ that returns a worst-case performance bound; the tightness of the bound is accuracy, and the computational cost of checking the proof is simplicity, so a good explanation pushes out the (tightness, compactness) Pareto frontier. Validity conditions—model-level, ontic, causal-mechanistic, falsifiable—must be met before the virtues are compared.","core_discovery":"The central claim is that 'what makes a good explanation?' can be answered by a pluralist set of explanatory virtues, and that these virtues—not subjective preference—should drive theory choice in mechanistic interpretability. To be valid, an MI explanation must be model-level, ontic, causal-mechanistic, and falsifiable. Among valid explanations, the paper maintains that the better explanation is the one embodying more of the virtues: empirical virtues such as accuracy, descriptiveness, co-explanation, and fruitfulness, and theoretical virtues such as precision, power, unification, consistency, simplicity, hard-to-varyness, and nomologicity. The framework attaches each virtue to one of four philosophy-of-science accounts and gives it a mathematical definition; for example, accuracy is a likelihood and hard-to-varyness is being a local maximum of log-accuracy minus complexity. When the rubric is applied to clustering, sparse autoencoders, causal-circuit analysis, and Compact Proofs, the paper finds that Compact Proofs are the method that considers many virtues and that current methods systematically neglect simplicity, unification, co-explanation, and nomological principles.","pith_inferences":["A natural test of the framework is to instantiate each virtue as a computable proxy—likelihood for accuracy, description length for simplicity, edit distance to a local maximum for hard-to-varyness—and score the same set of explanations with and without the proxies; convergence would suggest the rubric is measuring something real.","The truth-conduciveness claim is an empirical hypothesis: run a prediction tournament where rival explanations of the same model are scored on virtues and then probed by novel interventions, and see whether the virtue leader continues to predict behaviour outside its training distribution.","The framework's emphasis on unification suggests that shared, reused substructures should be privileged across tasks; if the same units keep predicting behaviour in new tasks, that would strengthen the link between unification and truth.","If truth-conduciveness fails, the framework still documents how interpretability researchers actually trade off values, but it would lose its normative force as a guide to which explanation is correct."],"forward_implications":["MI researchers can replace subjective intuition with a shared rubric when two explanations of the same model conflict, giving explicit epistemic reasons for theory choice.","Methods can be improved on the axes the framework flags as neglected: simplicity, unification, co-explanation, and nomological principles.","Compact Proofs offer a concrete way to turn the accuracy–simplicity trade-off into a measurable Pareto frontier, with faithful explanations yielding tighter bounds at lower verification cost.","A clearer definition of explanatory simplicity would let different explanation methods be compared on a single accuracy–simplicity curve.","Seeking universal principles and reused building blocks would move interpretability from cataloguing individual cases toward nomological, predictive explanations."],"supporting_citations":[{"why":"Documents two mutually inconsistent yet seemingly valid MI explanations of the same group-operation task, the theory-choice problem the framework is built to solve.","marker":"Wu et al. (2024)"},{"why":"Cited for the truth-conduciveness of explanatory virtues and for the adhocness test; the framework's epistemic force rests on this source.","marker":"Schindler (2018)"},{"why":"Supplies the Bayesian account of explanatory values behind the log-likelihood definitions of accuracy, precision, descriptiveness, co-explanation, power, and unification.","marker":"Wojtowicz & DeDeo (2020)"},{"why":"Lists the classic theory-choice virtues—accuracy, consistency, scope, simplicity, fruitfulness—that form the Kuhnian pillar of the framework.","marker":"Kuhn (1981)"},{"why":"Provides the deductive-nomological model of explanation that underlies the nomologicity virtue.","marker":"Hempel & Oppenheim (1948)"},{"why":"Supplies the hard-to-varyness criterion, formalised as a local maximum of log-accuracy minus complexity.","marker":"Deutsch (2011)"},{"why":"Defines Compact Proofs and shows faithful explanations yield tighter and more compact performance bounds, the basis for the paper's favourable evaluation.","marker":"Gross et al. (2024)"},{"why":"Provides the validity conditions for MI explanations (model-level, ontic, causal-mechanistic, falsifiable) and the compression view of explanation used throughout.","marker":"Ayonrinde & Jaburi (2025)"},{"why":"The MDL-SAE case study showing that choosing description length over parsimony as the simplicity measure resolves SAE pathologies; supports the simplicity research direction.","marker":"Ayonrinde et al. (2024)"}],"fun_headline_variants":["Philosophy of science offers new scorecard for AI explanations","Compact proofs top new AI explanation virtues framework","Why AI explanations fail: missing simplicity and unification","A virtue-based rubric for judging neural net explanations","Four philosophical views sharpen AI explanation quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the listed explanatory virtues are truth-conducive—that an explanation embodying more of them is genuinely more likely to be correct—because without that link the framework becomes a subjective preference list rather than a guide to which explanation of a neural network is actually true.","fun_headline_variants_meta":{"raw":{"variants":["Philosophy of science offers new scorecard for AI explanations","Compact proofs top new AI explanation virtues framework","Why AI explanations fail: missing simplicity and unification","A virtue-based rubric for judging neural net explanations","Four philosophical views sharpen AI explanation quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000256,"raw_usage":{"total_tokens":1558,"prompt_tokens":911,"completion_tokens":647,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":577}},"tokens_in":527,"tokens_out":647,"duration_ms":6657,"temperature":1.0,"reasoning_tokens":577,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:18:55.190426+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one model, such as a small transformer trained to compose group operations, generate two valid MI explanations of the same behaviour that score very differently on the virtues, then run a battery of held-out intervention experiments that the two explanations predict differently; if the lower-scoring explanation predicts the held-out behaviour at least as well as the high-scoring one, the claim that the virtues are truth-conducive is falsified.","supporting_citations":[{"cited_title":"Theoretical Virtues in Science: Uncovering Reality Through Theory","cited_arxiv_id":null,"evidence_quote":"Cited for the truth-conduciveness of explanatory virtues and for the adhocness test; the framework's epistemic force rests on this source."},{"cited_title":"From probability to consilience: How explanatory values implement bayesian reasoning","cited_arxiv_id":null,"evidence_quote":"Supplies the Bayesian account of explanatory values behind the log-likelihood definitions of accuracy, precision, descriptiveness, co-explanation, power, and unification."}],"review_version":1}