{"id":"76c05c22-151c-4d3b-84cd-df141ed5486b","arxiv_id":"1908.08005","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Grammar-guided genetic programming with dimensional consistency constructs high-level physics features and improves classifier accuracy on two of three HEP datasets.","lead":"This paper presents a genetic programming method that builds new data features for physics classification while enforcing consistent physical units. It aims to give physicists automatically constructed, interpretable features that improve classifier accuracy on particle collision data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported p-values do not support the abstract's 'significant gain' claim: Table II's Higgs p-value is irreconcilable with the means/SDs, and no significant gain is shown on τμ3.","rationale":"The reader's conditional verdict is appropriate. My stress-test identifies the statistical significance reporting as the most load-bearing concern because the central claim explicitly asserts significant classification gains, and the numbers in Table II do not obviously support that assertion. The discrepancy between the reported p-value and the descriptive statistics is concrete and testable, and the same table suggests the τμ3 dataset shows no gain over simple GP, contradicting the abstract's phrasing. The reader's weakest_assumption focused on type assignments and grammar completeness; that is a genuine reproducibility and generality concern, but since the paper's headline is an empirical performance claim, a failure in the significance evidence would more directly undermine the stated contribution. I am not alleging misconduct; the issue could be a mislabeled p-value or a typo, but it must be resolved before the results are taken at face value. A requested data release or re-run with per-run gains would settle this concern cleanly. The verdict remains CONDITIONAL rather than REJECT because the underlying method may still work; the required corrections are reporting and verification oriented.","tokens_in":12818,"tokens_out":5596,"duration_ms":56573,"concrete_test":"Recover or regenerate the per-run gains for Table II (at least 20 runs per method per dataset, same settings) and recompute the Welch p-values for (i) best of GGGP/PGGGP vs simple GP on each dataset and (ii) each constrained method vs baseline. Then check: does the Higgs comparison reproduce p≈10^{-17}, and does the τμ3 comparison reach p<0.05? If either fails, the abstract and conclusion need to be narrowed to the datasets and comparisons that actually pass, or the significance claim must be retracted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim is that the constructed features bring a significant gain in classification accuracy. The evidence as printed does not support that claim. In Table II, the best constrained method on Higgs has mean gain 2.23% vs. simple GP's 1.92%, with standard deviations 0.68 and 0.11. A Welch t-test on those numbers with n=20 gives t≈2.0 and p≈0.06, not the reported p-value of 10^{-17}; to reach 10^{-17} from those means and SDs would require an implausibly large number of runs. If the reported p-value instead compares a constrained method against the baseline, it does not establish that the GP constraint is what helps. On τμ3, the table shows GGGP (0.43±0.28) and PGGGP (0.54±0.37) are not better than simple GP (0.58±0.36), yet the abstract says results on three datasets show significant gain and the conclusion says 'for the three datasets... significant improvement' after earlier saying 'two high-energy physics datasets.' The Discussion also promotes a single best Higgs run (+3.82%) while the table reports the mean is 2.23±0.68, which overstates expected performance. These issues attack the central claim directly: unless the significance evidence is corrected, the method's advantage over unconstrained GP is not established. The type-mapping/grammar incompleteness is a reproducibility concern, but the statistical mismatch is more load-bearing because it concerns the headline empirical conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a genetic programming (GP) approach to automatic feature construction for high-energy physics, in which a context-free grammar enforces dimensional consistency of the evolved expressions. Two variants are studied: GGGP, which uses only the grammar, and PGGGP, which additionally uses a hand-set transition probability matrix over grammar operators. The method is evaluated on three physics datasets (Higgs, DVCS, and τ→μμμ) with XGBoost as the main classifier, and the paper claims significant accuracy gains over simple unconstrained GP and over a PSO feature-construction baseline, while also claiming interpretability of the constructed features, validated by HEP experts.","tokens_in":13142,"tokens_out":3434,"duration_ms":35723,"significance":"If the claims are sustained, the work is a practical contribution to a real need: automatic construction of physically meaningful, dimensionally consistent observables for collider analyses, with interpretability for domain experts. The main strengths are the use of grammars to enforce dimensional consistency, the exploration of several fitness functions and classifiers, the comparison across three different experimental setups, and the explicit attempt to interpret the evolved formulas. The paper also makes a useful design point: the manually chosen transition matrix can trade off single-feature performance for interpretability. However, the headline empirical claim of a \"significant gain\" is not currently supported by the reported statistics, and the reproducibility of the search space is limited by the absence of the full grammar and type assignments.","major_comments":[{"comment":"The first p-value reported for the Higgs dataset (10^-17) is irreconcilable with the means and standard deviations shown in the same table. For a comparison between the best constrained method (GGGP, 2.23±0.68) and simple GP (1.92±0.11) with n=20, a Welch t-test gives t≈2.0 and p≈0.06, not 10^-17. Please report the exact test statistic, degrees of freedom, and the precise comparison performed, and re-examine the abstract's \"significant gain\" claim in light of the corrected test.","section":"IV-C, Table II"},{"comment":"The paper promotes a single best GGGP feature on the Higgs dataset with a +3.82% accuracy gain and compares it to the invariant-mass feature (+2.91%), concluding \"we overcome the invariant mass.\" This is a selected maximum over the runs, whereas Table II reports a mean of 2.23±0.68 for GGGP on the same task. The expected gain is therefore not +3.82%, and the comparison overstates the method's typical performance. Please report the mean and standard deviation of the gain over runs, or provide a proper paired comparison against the invariant-mass feature, and adjust the wording.","section":"IV-D"},{"comment":"The claim that the method brings a significant gain on \"three physics datasets\" is not supported by the results in Table II. On the τμ3 dataset, GGGP (0.43±0.28) and PGGGP (0.54±0.37) are not better than simple GP (0.58±0.36), and no p-value is reported comparing the constrained methods to simple GP for this dataset. The conclusion itself contains a contradiction: it first says the method significantly improves accuracy \"for two high-energy physics datasets,\" then says \"for the three datasets, our interpretable GGGP-based feature construction algorithm brings a significant improvement.\" The claim must be restricted to the datasets and comparisons where the statistical evidence actually supports it.","section":"Abstract and Conclusion"},{"comment":"The paper does not provide the physical type assignment for the 17, 36, and 46 input features of the three datasets, nor the complete grammar used in the experiments; Figure 1 is explicitly stated to be a simplified version. Since the method's search space is defined by these type tags and grammar productions, a mistyped feature or an incomplete type set could exclude the best possible feature, and the experiments cannot be reproduced without this information. Please publish the full grammar, the type map for each dataset, and the exact transition matrix used.","section":"III-A and dataset description"}],"minor_comments":[{"comment":"The text contains a typo: \"Welchs t-test\" should be \"Welch's t-test.\"","section":"II-C"},{"comment":"The caption says the first p-value compares the best of PGGGP and GGGP to simple GP, but the table only shows a first p-value for the Higgs row; the corresponding cells for DVCS and τμ3 are missing. Please either report those values or clarify that they were not computed.","section":"IV-C, Table II"},{"comment":"Equation (1) is difficult to read because of formatting artifacts such as \"missingtE\" and \"missingtE2\" and the placement of exponents and square roots. Please typeset the formula cleanly, and also check the similar garbled notation in Table I and around the discussion of the five-feature set.","section":"IV-D, Equation (1)"},{"comment":"The paper says results are means over \"at least 20 independent runs\" but does not state how the random seeds were generated or whether the same seeds were used across compared methods. This information would help assess the variance and the paired/unpaired nature of the significance tests.","section":"IV-B"},{"comment":"The interpretability discussion would benefit from a more precise statement of how the HEP expert validation was conducted, since the abstract mentions expert validation but the body only describes qualitative agreement with physics intuition.","section":"IV-D"}],"recommendation":"major_revision","confidential_remarks":"The Table II p-value inconsistency appears to be a genuine error, and the Section IV-D best-run comparison is misleading. These directly affect the central empirical claim, so I recommend requiring a corrected statistical analysis and full specification of the grammars/type assignments before reconsidering the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: the core idea is worth your time—using a context-free grammar to enforce dimensional consistency in genetic programming for feature construction is a natural fit for HEP, and the transition-matrix bias toward physics-like formulas is a small but sensible twist. The paper does a real service by showing the grammar can rediscover something like a transverse-energy-balance observable on the Higgs data. That part is new and plausible.\n\nThe problems are in the evidence, not the idea. Table II's Higgs p-value of 1e-17 is irreconcilable with the reported means and standard deviations: GGGP 2.23±0.68 vs simple GP 1.92±0.11 over 20 runs gives a Welch t around 2 and p around 0.06. Either the p-value is a typo, the runs are correlated in a way not described, or the number of runs is much larger than the stated \"at least 20.\" As printed, the headline significance claim is unsupported. On τμ3, the constrained methods actually do not beat simple GP (0.43±0.28 and 0.54±0.37 vs 0.58±0.36), yet the abstract says three datasets show significant gain, and the conclusion contradicts itself—first saying two datasets, then saying three. That overclaim matters.\n\nAlso soft, in decreasing order: the Discussion promotes a single best Higgs run (+3.82%) against invariant mass while the table's mean is 2.23±0.68, which overstates what a user should expect; the full grammar and per-feature type assignments for the 17/36/46 features are not published, so the search-space constraints are not reproducible; and no code is shipped. The paper does give credit where it is due—it cites the dimensional-consistency grammar of Ratle and Sebag, and it is honest about not doing a full expert validation study.\n\nWho is this for? People working on feature construction, especially with physical units, and applied GP folks who want to see a realistic use of grammar constraints. The method is likely useful, but the empirical case needs serious cleanup. I would send it to peer review rather than desk reject, with a strong request for corrected statistics, a fair baseline comparison (including reporting mean and spread for the invariant-mass comparison), and either code or a complete type/grammar specification. The central idea holds up; the reporting does not.","headline":"A genuinely useful grammar-guided GP idea for physics feature construction, but the paper's own statistics do not back the 'significant gain on three datasets' claim as printed.","tokens_in":13675,"tokens_out":1597,"would_cite":true,"duration_ms":17621,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Unit-aware genetic programming automatically constructs interpretable, dimensionally consistent features that improve classification accuracy on particle-collision data by up to 3.8 percentage points.","keywords":["feature construction","grammar-guided genetic programming","dimensional consistency","high-energy physics","interpretability","units of measurement","classification","genetic programming"],"falsifier":"Re-run the three experiments with the type tags deliberately corrupted—mark one Energy feature as Float, or collapse all types to a single Float type. If the accuracy gains over baseline survive, the grammar's dimensional constraint is not what produces them; if the best known physics feature becomes unreachable and accuracy drops, the constraint is load-bearing.","tokens_in":12579,"feed_emoji":"⚛️","tokens_out":11279,"duration_ms":100719,"temperature":0.7,"pith_summary":"Physicists routinely hand-build high-level variables, such as invariant masses, from measured momenta and angles to improve classifiers that separate signal from background. This paper claims that this search can be automated without losing physical meaning: a genetic programming algorithm, constrained by a grammar that only permits dimensionally consistent expressions, evolves new features that respect units and stay readable. On three simulated particle-collision datasets, adding the evolved features to the base feature set improves the accuracy of a gradient-boosted classifier, with the best single feature on the first dataset adding 3.82 percentage points over the baseline, compared with 2.91 points for the hand-built invariant-mass feature. The authors further report that this is the first feature-construction method for interpretable features that explicitly handles units of measurement, and that physics experts validated the constructed features.","feed_headline":"Unit-aware feature search lifts physics accuracy by 3.8 points","feed_subtitle":"Automatically built dimension-consistent features lift classifier accuracy on collision data by up to 3.8 points.","key_machinery":"The load-bearing object is a context-free grammar whose nonterminals are physical types—Energy, Angle, Float, squared Energy—and whose production rules permit only combinations that respect dimensional analysis: an Energy can be added to an Energy, divided by a Float, or obtained by taking the square root of a squared Energy, but an Energy and an Angle can never be added. Every candidate feature is a derivation from this grammar, so its expression is guaranteed to be independent of the system of measurement. A second component, a probability transition matrix over operators, biases derivations toward recurrent physics patterns, such as a square root applied to a sum of squares; the matrix is set by hand for the experiments and compared with uniform probabilities. Evolution uses standard genetic programming operators—initialization, mutation, crossover—restricted so that all offspring remain grammar-compliant, and fitness is the cross-validated accuracy of a gradient-boosted classifier trained on the constructed feature set.","core_discovery":"The central discovery is that enforcing dimensional consistency through a context-free grammar does more than filter out invalid expressions: it concentrates the search on features that are simultaneously physically meaningful and competitive for classification. On the first dataset, the grammar-guided method finds a single feature that outperforms the invariant mass; on the second dataset it delivers a statistically significant accuracy gain over unconstrained genetic programming; on the third dataset the gain is smaller, which the paper attributes to a lack of basic angular variables. The paper also shows that a hand-set probability transition matrix over grammar rules—for instance, favouring square roots of sums of squares—steers evolution toward formulas whose components physicists recognize, such as transverse-energy balances and cosine differences of angles, while keeping final accuracy comparable to the unguided grammar method. When several features are constructed jointly, the probability-guided variant matches the unguided one while remaining more interpretable.","pith_inferences":["A natural stress test would re-run the three datasets with deliberately mistyped features: if the gains survive when all types are collapsed to a single Float, the grammar is not the active ingredient; if they vanish, the type tags are load-bearing.","Because the transition matrix is hand-set, its entries could instead be estimated from a corpus of physics formulas; such a data-driven prior would make the method portable and could be compared directly against the hand-tuned matrix on the same benchmarks.","Unit-consistent constructed features are also natural inputs for tasks beyond classification, such as anomaly detection or regression in experimental physics, where readable observables matter equally."],"forward_implications":["Physicists gain an automated source of candidate observables whose expressions visibly carry their physical dimensions, so each candidate can be assessed by an expert before use.","Because the constructed features are independent of the measurement system, they can be reused across detector calibrations and simulation setups.","The transition-matrix guidance behaves as a performance-versus-interpretability dial: constructing several features at once lets the guided method match the unguided one while producing recognizable physics components.","Since the grammar is defined over physical types rather than dataset-specific variable names, the same machinery can be re-instantiated for other unit-carrying datasets.","The gains persist across different fitness functions—gradient-boosted trees, decision trees, and k-nearest neighbours—which indicates the constructed features themselves, not one particular wrapper, drive the improvement."],"supporting_citations":[{"why":"Supplies the genetic programming framework, including the ramped half-and-half initialization used to generate initial trees.","marker":"[12]"},{"why":"Supplies the grammar-guided dimensional-consistency approach that the paper adapts to high-energy physics.","marker":"[27]"},{"why":"Supplies strongly typed genetic programming, the typing discipline that the grammar enforces on expression trees.","marker":"[22]"},{"why":"Supplies the probabilistic model-building idea underlying the transition matrix that guides tree construction.","marker":"[33]"},{"why":"Supplies the particle-swarm feature-construction baseline that constrained genetic programming outperforms on the first dataset.","marker":"[10]"},{"why":"Supplies the first classification dataset on which the main accuracy gains are demonstrated.","marker":"[34]"},{"why":"Supplies the physics context and dataset for the third classification task.","marker":"[35]"}],"fun_headline_variants":["Unit-aware GP: interpretable features, 3.8-point accuracy gain","Grammar enforces dimensions, improves physics ML by 3.8 points","First interpretable features with units boost collision classification","Dimensional consistency guides GP to better, interpretable physics features"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every input feature must be tagged with the correct physical type (Energy, Angle, Float, squared Energy, etc.) and the type system must be complete enough to express the genuinely best feature; the paper does not publish its tag assignments and states the grammar figure is simplified.","fun_headline_variants_meta":{"raw":{"variants":["Unit-aware GP: interpretable features, 3.8-point accuracy gain","Grammar enforces dimensions, improves physics ML by 3.8 points","First interpretable features with units boost collision classification","Dimensional consistency guides GP to better, interpretable physics features"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000725,"raw_usage":{"total_tokens":3246,"prompt_tokens":934,"completion_tokens":2312,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":2239}},"tokens_in":550,"tokens_out":2312,"duration_ms":18392,"temperature":1.0,"reasoning_tokens":2239,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:50:07.226967+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the three experiments with the type tags deliberately corrupted—mark one Energy feature as Float, or collapse all types to a single Float type. If the accuracy gains over baseline survive, the grammar's dimensional constraint is not what produces them; if the best known physics feature becomes unreachable and accuracy drops, the constraint is load-bearing.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the genetic programming framework, including the ramped half-and-half initialization used to generate initial trees."},{"cited_title":"Grammar-guided genetic programming and dimensional consistency: application to non-parametric identiﬁcation in mechanics,","cited_arxiv_id":null,"evidence_quote":"Supplies the grammar-guided dimensional-consistency approach that the paper adapts to high-energy physics."},{"cited_title":"Strongly Typed Genetic Programming,","cited_arxiv_id":null,"evidence_quote":"Supplies strongly typed genetic programming, the typing discipline that the grammar enforces on expression trees."},{"cited_title":"Avoiding the Bloat with Stochastic Grammar-based Genetic Programming","cited_arxiv_id":"cs/0602022","evidence_quote":"Supplies the probabilistic model-building idea underlying the transition matrix that guides tree construction."},{"cited_title":"New Representations in PSO for Feature Construction in Classiﬁcation,","cited_arxiv_id":null,"evidence_quote":"Supplies the particle-swarm feature-construction baseline that constrained genetic programming outperforms on the first dataset."},{"cited_title":"Learning to discover: the Higgs boson machine learning challenge - Documentation,","cited_arxiv_id":null,"evidence_quote":"Supplies the first classification dataset on which the main accuracy gains are demonstrated."}],"review_version":1}