{"id":"19792509-afef-4b42-b9d2-faaccf826902","arxiv_id":"2511.07002","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":11,"one_line_summary":"Probe prompting groups attribution-graph features by cross-prompt activation signatures into concept-aligned supernodes, but the abstract's claim of validated steering in 45,596 interventions is unsupported by the body's 5-prompt descriptive evaluation.","lead":"This paper introduces probe prompting, a rule-based pipeline that groups internal features of a language model into concept-aligned 'supernodes' based on how they respond to a small set of carefully chosen probe prompts. If it works, circuit interpretation could drop from hours per prompt to minutes, but the paper's body supports only a small descriptive demonstration, not the abstract's grander claims.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline claim of 45,596 entity-swap steering validations is absent from the body; the paper's actual evidence is 5 prompts, 39 features, and no causal steering. The central claim as advertised is therefore unsupported.","rationale":"The reader's verdict of REJECT is correct, but the single most load-bearing concern is not the weakest_assumption they list (CLT linearization faithfulness). The CLT concern is genuine and acknowledged in §5.2, but even if replacement models were perfectly faithful, the paper still lacks the 45,596-intervention steering study promised in the abstract. The reader's strongest_claim correctly identifies the abstract's universal steering validation, and their rationale mentions the abstract–body inconsistency, so there is partial agreement. However, because the formal weakest_assumption field focuses on CLT linearization rather than the missing evidence, I mark agreement as partial. A concrete internal check—searching for the claimed count and any steering/ablation results—would settle whether the concern lands. Until such evidence appears, the central claim as stated cannot be accepted; the appropriate disposition is to keep the REJECT verdict, or to require major revision that either supplies the missing experiment or rewrites the abstract to match the body's proof-of-concept scope.","tokens_in":15779,"tokens_out":3705,"duration_ms":37447,"concrete_test":"Search the manuscript, appendices, and linked repository for the string '45,596' and for any steering or ablation protocol with outcome tables (e.g., activation patching or feature steering results). If the count and protocol are absent, the abstract's universal steering claim should be withdrawn or the missing experiments supplied. A minimal analytic check: enumerate all entity-swap interventions derivable from §4 and §A.3 (five prompts, 39 features, one swap pair); if the maximum possible number is far below 45,596, the headline claim is unsupported as written.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing concern is not primarily the CLT replacement-model linearization acknowledged in §5.2, but a direct contradiction between the abstract and the reported experiments. The abstract states: 'Across four factual domains, on Gemma-2-2B with a public CLT dictionary and 45,596 entity-swap interventions, we find that the labeled supernodes have the predicted steering behavior in every one of them.' The full text never reports such a study. §4 lists five prompts (Austin, Oakland, Michael Jordan, small–opposite, muscle–diaphragm); Table 3's transfer analysis uses 39 features and a single entity swap (Dallas→Oakland); no table, appendix, or repository artifact describes 45,596 interventions. No arithmetic in the methods (five prompts, 39 features, 5–10 probes) yields that count. Instead, §2 and §5.2 explicitly state that causal ablations and steering are 'planned for future work.' Thus the abstract's central causal claim is unverifiable from the manuscript as written. The body's narrower claims—compression (Completeness 0.83), behavioral coherence (2.3×/5.8×), and early-vs-late transfer (64%, 10.1-layer gap)—may be legitimate proof-of-concept results, but they do not support 'predicted steering behavior in every one of 45,596 interventions.' This is an internal inconsistency, not a disagreement with external consensus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 'probe prompting,' a rule-based pipeline that groups features of an attribution graph into concept-aligned supernodes (Semantic, Relationship, Say-X) based on cross-prompt activation signatures (CPAS). The method is evaluated on five prompts in four factual domains using Gemma-2-2B with a public CLT dictionary. The body reports three descriptive findings: (C1) concept-aligned grouping yields higher behavioral coherence than cosine or layer-adjacency baselines on the Michael Jordan circuit; (C2) subgraphs preserve 83% Completeness while compressing graph size; (C3) an entity-swap study (Dallas→Oakland) suggests early-layer features transfer more robustly than late-layer Say-X features. The abstract, however, claims validation of 'predicted steering behavior' across 45,596 entity-swap interventions, which does not appear anywhere in the body. The manuscript is transparent about its small scale and correlational nature in Section 5.2 and Appendix A.4, but the abstract's causal claim is unsupported.","tokens_in":16215,"tokens_out":3223,"duration_ms":35977,"significance":"If the body's modest claims were confirmed with proper held-out evaluation, the method could meaningfully reduce the manual cost of circuit interpretation and provide a reusable, transparent taxonomy. The paper ships a public codebase, interactive demo, deterministic pipeline, and clearly documents many limitations. However, the abstract advertises a large-scale causal validation that the manuscript does not contain, and the evaluation of the main behavioral-coherence claim is circular because the same metrics used to define categories are used as success criteria. The likely contribution is a proof-of-concept tool, not the strong empirical result claimed in the abstract.","major_comments":[{"comment":"The abstract states: 'Across four factual domains, on Gemma-2-2B with a public CLT dictionary and 45,596 entity-swap interventions, we find that the labeled supernodes have the predicted steering behavior in every one of them.' The body never reports this study. Section 4 lists five prompts, Table 3 reports 39 features and a single entity swap (Dallas→Oakland), and Section 5.2 says causal ablations and steering are 'planned for future work.' No arithmetic from the described setup yields 45,596. This is a direct internal contradiction between the advertised central claim and the reported evidence.","section":"Abstract vs. Section 4/5.2"},{"comment":"The evaluation of C1 (behavioral coherence) is circular. The classification rules in Section 3.2 define Semantic (Dictionary) by peak consistency ≥0.80 and Say-X by functional dominance and functional consistency, while Table 2 compares concept-aligned grouping to baselines on 'Peak Token Consistency' and 'Sparsity Consistency' — the very quantities the rule-based classifier is designed to optimize. The paper itself concedes in Appendix A.4 that 'No held-out test set: All tuning was performed on circuits included in evaluation.' Without a held-out split or evaluation on metrics not used in the decision rules, the 2.3×/5.8× improvements in Table 2 are expected by construction, not evidence of superior grouping.","section":"Section 3.2 and Table 2"},{"comment":"The transfer hierarchy claim (C3) rests on a single entity-swap pair and an arbitrary threshold (cosine > 0.80) for defining transfer. The 10.1-layer difference is reported without significance tests or confidence intervals, and the authors acknowledge the small sample in the text. Given that the abstract elevates this to a claim about 'predicted steering behavior' across tens of thousands of interventions, the single-pair descriptive finding cannot support the advertised claim. Replication across multiple state-capital pairs and a proper test of the layer-difference hypothesis are needed before this can be called a 'backbone-and-specialization' result.","section":"Table 3 and Section 4"}],"minor_comments":[{"comment":"Typo: 'michael hordan plays' appears in the example-data directory listing; should be 'michael jordan plays.'","section":"Appendix A.3"},{"comment":"In the target-token mapping rules, 'Texas’scapital' is missing a space; also the possessive marker example uses a curly apostrophe inconsistently with the rest of the text.","section":"Appendix A.7"},{"comment":"The sentence 'the current prototype remains unreliable in multilingual contexts without explicit cross-lingual adaptation' is repeated nearly verbatim in Section 3.2 and Appendix A.7; consider consolidating.","section":"Section 5.2"},{"comment":"The caption says 'same pinned nodes' and then notes that Replacement/Completeness are near-invariant; it would be clearer to state explicitly that these metrics are not discriminative for the comparison, as the text does.","section":"Table 2 caption"},{"comment":"The set of prompts is described as 'four factual domains' but the list includes five prompts; clarify whether the antonymy and anatomy pairs count as factual domains or as separate categories.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The abstract's 45,596-intervention steering claim is a serious integrity issue that must be corrected before publication. Even after rewriting the abstract, the circular evaluation in C1 needs a proper held-out design. The authors are transparent about limitations in the body, which is commendable, but the abstract currently misrepresents the work. I would not reject outright because the tool and its proof-of-concept results have some value, but the revision must align the abstract with the body and address the circularity with new experiments or clear reframing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know up front: the abstract claims “45,596 entity-swap interventions” with “predicted steering behavior in every one of them,” but the body never reports that study. The actual evidence is five prompts, 39 features in the transfer analysis, one entity-swap pair, and no causal steering or ablation—the text explicitly says those are future work. That is a serious abstract–body inconsistency, and it is the paper’s main problem.\n\nWhat is genuinely new: the move from geometric similarity to behavior-grounded grouping via probe prompts. The CPAS (Cross-Prompt Activation Signature) idea—summarizing a feature’s response across systematically varied probes, then grouping by consistency and peak behavior—is a legitimate new technique. The rule-based taxonomy (Semantic/Relationship/Say-X) is transparent and cheap to apply. The authors also ship code, an interactive demo, and example data, which is real credit.\n\nThe body is also refreshingly honest: it labels results as descriptive, acknowledges no significance tests, flags the CLT attention-freezing limitation, and even documents a French-language failure. So the soft spots are not hidden in the text—they are just contradicted by the abstract. Beyond that mismatch, the main weaknesses are small scale (n=5 circuits, n=39 features), evaluation metrics that overlap with the classification thresholds (e.g., peak consistency appears both as a rule and an outcome), and a baseline comparison on only one circuit. None of these are fatal for a proof-of-concept, but they make the headline claims untenable.\n\nI agree with the stress-test note: the load-bearing issue is internal inconsistency, not disagreement with external consensus. The fix is straightforward: rewrite the abstract to match the body, present this as a descriptive feasibility study, and either supply the 45,596-intervention steering validation or remove that claim entirely.\n\nWho should read this? People doing circuit interpretation who want a fast first-pass grouping tool, and anyone studying how interpretability claims get oversold. Worth a serious referee after the authors revise—the method deserves engagement, and the transparency in the body is a good precedent.\n\nMy recommendation: send it to peer review, but with a clear request to correct the abstract and reframe the contributions. If the authors can align the claims with the evidence, this becomes a useful, citable method paper.","headline":"The abstract promises 45,596 causal steering validations; the body delivers 5 prompts, 39 features, and explicitly defers causal tests to future work—fix that mismatch and this is a worthwhile proof-of-concept, not a bomb.","tokens_in":16671,"tokens_out":1888,"would_cite":true,"duration_ms":23389,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Probe prompting automatically groups transformer circuit features into interpretable supernodes from their activation patterns on concept-targeted probe prompts, and the paper reports descriptive evidence that these supernodes are behaviora","keywords":["probe prompting","attribution graphs","circuit interpretation","cross-layer transcoders","supernodes","behavioral signatures","mechanistic interpretability","layerwise hierarchy"],"falsifier":"Run the published pipeline on a new entity-swap pair with identical syntax (e.g., 'The capital of France is' vs 'The capital of England is') and check whether transfer is again layer-dependent (early features ~64%, late Say-X ~36%). Additionally, ablate the labeled Say-X supernodes and measure the target logit: if removal does not lower the logit of the correct answer, the Say-X labels are not causally implicated in output promotion. A negative result on either test would undermine the central claims.","tokens_in":15621,"feed_emoji":"🧠","tokens_out":8018,"duration_ms":75769,"temperature":0.7,"pith_summary":"This paper tries to establish that the thousands of sparse features in a transformer's attribution graph can be automatically grouped into a small number of interpretable supernodes by measuring each feature's behavior on a handful of concept-targeted probe prompts, rather than by clustering activation vectors geometrically. The authors argue that these behavioral groupings are more coherent—features in one group tend to fire on the same tokens across contexts—and that the resulting subgraphs preserve about 83% of explanatory coverage while shrinking from thousands of nodes to tens. They also report evidence for a layerwise division of labor: features in early layers transfer across entity substitutions, while late-layer features specialize in promoting the specific output token. The payoff would be cutting circuit interpretation from hours to minutes and giving researchers a shared vocabulary for feature roles.","feed_headline":"Probe prompting keeps 83% of circuit coverage in readable subgraphs","feed_subtitle":"Automated grouping by behavior across varied prompts exposes a layerwise split: early transfer, late specialization.","key_machinery":"The central object is the Cross-Prompt Activation Signature (CPAS): for each feature, a behavioral profile computed from its activation on a set of probe prompts that preserve syntax but vary semantics. The profile records cosine similarity to the seed activation, robust z-scores, peak token identity and position, activation density, and sparsity. Aggregated across probes, these signatures feed transparent decision rules that assign each feature to one of three categories—Semantic (stable word or concept detector), Relationship (diffuse, sentence-spanning binder), or Say-X (output-promotion via functional tokens like 'is' or 'the')—and these categories become the supernodes. The work is carr","core_discovery":"The central claim is that a feature's role in a circuit is best read from its cross-prompt activation signature—how consistently and where it activates across systematically varied prompts—rather than from its raw activation vector or layer position. Using a rule-based classifier over these signatures, the paper labels features as semantic detectors, relationship binders, or output-promotion ('Say-X') nodes, and groups them into supernodes. On five factual-recall prompts in a 2-billion-parameter model equipped with a public cross-layer transcoder dictionary, these supernodes score higher on behavioral coherence than cosine or layer-adjacency baselines (2.3× peak-token consistency, 5.8× activ","pith_inferences":["The initial abstract's claim that 45,596 entity-swap interventions show predicted steering in every case is not supported by the body's reported experiments (five prompts, 39 features in the transfer study); the body itself labels its validation correlational. This discrepancy means the headline claim should be treated as a target, not as achieved evidence, until the large-scale data are released.","Because the attribution graphs come from replacement models that freeze attention patterns, the method may systematically miss features whose causal effect flows through attention changes; an attention-aware attribution variant could alter which features are labeled Say-X and could shift the layerwise hierarchy.","The rule thresholds (e.g., peak consistency ≥ 0.80, Say-X layer ≥ 7) were hand-tuned on the same circuits used for evaluation; a direct test would apply the pipeline to held-out prompt families and check whether default thresholds remain sensible without re-tuning.","A concrete way to stress-test the three-category taxonomy: generate probe prompts that include polysemous or multi-token entities and check whether the directionality rules ('is' → forward, 'of' → backward) still assign correct targets; the paper itself notes the 'of' directionality is currently implemented forward in code, a known issue."],"forward_implications":["If the method works as claimed, initial circuit analysis drops from roughly two hours per prompt to minutes, allowing broad surveys of circuit structure across many prompts and models.","A standardized Semantic/Relationship/Say-X vocabulary would make it possible to compare circuit motifs across tasks and to study whether the same relational features recur (e.g., 'capital-of' vs 'member-of').","The reported early-transfer/late-specialization split, if confirmed, implies that activation steering or weight edits aimed at transferring behavior should target early-layer backbone features, while late-layer output promoters are task-specific.","The subgraph compression (83% Completeness at ~54% Replacement) offers a legibility-vs-coverage trade-off that could be used as a default for human review, though the paper notes this applies to factual recall, not long reasoning."],"fun_headline_variants":["Probe prompting yields supernodes with 100% steering validation","Cross-prompt signatures make LLM circuits readable and causal","Grouping LLM features by probe responses beats cosine baselines","Probe prompting: readable subgraphs with 100% causal validation","From opaque circuits to supernodes via cross-prompt signatures"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the attribution graph—computed under a replacement model that freezes attention patterns and layer norms—faithfully reflects the model's causal computation; if attention routing carries the effects the graph is supposed to capture, the supernode labels and the early-vs-late hierarchy could describe a distorted mechanism.","fun_headline_variants_meta":{"raw":{"variants":["Probe prompting yields supernodes with 100% steering validation","Cross-prompt signatures make LLM circuits readable and causal","Grouping LLM features by probe responses beats cosine baselines","Probe prompting: readable subgraphs with 100% causal validation","From opaque circuits to supernodes via cross-prompt signatures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000434,"raw_usage":{"total_tokens":2043,"prompt_tokens":734,"completion_tokens":1309,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":1224}},"tokens_in":478,"tokens_out":1309,"duration_ms":12194,"temperature":1.0,"reasoning_tokens":1224,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T23:08:23.205399+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the published pipeline on a new entity-swap pair with identical syntax (e.g., 'The capital of France is' vs 'The capital of England is') and check whether transfer is again layer-dependent (early features ~64%, late Say-X ~36%). Additionally, ablate the labeled Say-X supernodes and measure the target logit: if removal does not lower the logit of the correct answer, the Say-X labels are not causally implicated in output promotion. A negative result on either test would undermine the central claims.","supporting_citations":[],"review_version":1}