{"id":"09937c81-7517-4d24-b1a1-fad7724ef3a4","arxiv_id":"2507.05858","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Symbolic regression produces analytic, detector-level CP-odd observables for WBF Higgs production and an analytic reconstruction of the Collins-Soper angle in ttH that are competitive with black-box ML and classical methods.","lead":"This paper uses symbolic regression to extract simple analytic formulas for CP-violating Higgs observables at the LHC, replacing black-box neural networks. The formulas work at detector level, are interpretable, and match or beat standard machine learning and classical reconstruction in simulated LHC events.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The ttH performance numbers depend on perfect b-jet assignment and b/bbar discrimination; without an experimental assignment strategy the detector-level claim is not established.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the ttH results assume perfect b-jet assignment and b/bbar discrimination, and this assumption conditions the headline Delta chi squared values. I agree that this is the most load-bearing issue. The paper is transparent about the assumption, but transparency does not remove the fact that the learned formulas use individual b and bbar momenta with different coefficients, so they are not robust to the combinatorial and charge-identification ambiguities present in real LHC data. This does not undermine the methodological contribution or the WBF results, and it is not an internal inconsistency; it is an experimental idealization that should be quantified before the detector-level claim is accepted. The reader's CONDITIONAL verdict is therefore appropriate, and my analysis does not move it. Other limitations, such as approximate detector smearing rather than a full simulation and the acknowledged non-CP-odd SymbolNet formula in WBF, are real but secondary; the b-jet assignment assumption has the most direct impact on the strongest numerical claim.","tokens_in":26283,"tokens_out":13027,"duration_ms":164986,"concrete_test":"Re-run the scenario-6 pipeline without truth-level b/bbar labels: assign the two b-tagged jets to the leptonic and hadronic tops via a simple heuristic such as minimizing |m_{b,lepton}-m_top|, and replace b/bbar charge information with charge-symmetrized or unordered b-jet momenta. Retrain SymbolNet and PySR under this realistic input and recompute the Delta chi squared values in Fig. 12. If the advantage over classical reconstruction shrinks substantially or disappears, the headline ttH claim is conditional on an experimentally unavailable oracle.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is in Sec. 4.1 and App. A: the ttH regression uses the momenta of b and bbar as separate, correctly assigned inputs. The formulas in Appendix D, e.g. scenario 6, contain p_{z,b}, p_{z,bbar}, p_{z,q}, and p_{z,qbar} with different fitted coefficients, so the learned mapping is explicitly not symmetric under swapping the two b-jets or flipping b/bbar. At the LHC, b-jet charges are not tagged reliably and the assignment of b-jets to the leptonic versus hadronic top is combinatorial; this is exactly the difficulty a detector-level analysis must solve. The reported Delta chi squared = 7.628 (SymbolNet) and 7.491 (PySR) in Fig. 11 are therefore upper bounds conditioned on an oracle that labels b and bbar correctly. The paper acknowledges this ('we assume that the b-jets have been correctly assigned...', 'experimentally very difficult'), but does not quantify the degradation. Because the central performance claim is the detector-level advantage in the most realistic ttH scenario, this idealization is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes using symbolic regression to obtain analytic, interpretable observables for CP studies. Two SR implementations are used, PySR and an adapted SymbolNet. In WBF Higgs production (H->γγ + 2 jets, generated with MadGraph+Pythia+Delphes, SMEFT operator cHW~), both methods learn detector-level analytic CP-odd observables; their bin-wise asymmetries give significances comparable to or better than a BDT and can be checked analytically for CP parity. In ttH production (semi-leptonic decays, six increasingly realistic scenarios), the SR methods reconstruct the Collins-Soper angle from final-state momenta; in the most realistic scenario 6 (E_T^miss, extra jet, smearing) they recover about 80% of the parton-level Δχ² for α_t=45° versus about 60% for classical reconstruction. The paper emphasizes interpretability, data efficiency, and the complementarity of the two SR algorithms.","tokens_in":26547,"tokens_out":12474,"duration_ms":142490,"significance":"If the results hold, this is a useful contribution: explicit analytic formulas that can be checked for CP parity and used directly, with data-efficiency advantages and a fair comparison against BDT and classical reconstruction. The WBF study is particularly convincing because the CP-odd property of the learned formulas is verified analytically, and the comparison with the known parton-level observable is a good sanity check. The ttH study is well structured into six benchmark scenarios, with explicit formulas in Appendix D and repeated training runs. However, the quantitative detector-level claims are conditioned on idealized simulation and, in ttH, on perfect b-jet assignment and b/bbar discrimination; the absolute numbers should be read as upper bounds rather than as realistic experimental projections.","major_comments":[{"comment":"The headline ttH result is conditioned on an oracle that correctly assigns the two b-jets to the leptonic and hadronic top decays and distinguishes b from bbar. This assumption is stated in Sec. 4.1 and App. A, but it is not mitigated. The formulas in Appendix D (e.g., scenario 6) use p_z,b, p_z,bbar, p_z,q and p_z,qbar with independent fitted coefficients, so the learned mapping is explicitly not invariant under swapping the two b-jets or under b<->bbar interchange. At the LHC, b-jet charges are not tagged reliably and the assignment of the two b-jets is combinatorial; a detector-level analysis must solve this problem. The quoted Δχ² = 7.628 (SymbolNet) and 7.491 (PySR) in Fig. 11, and the advantage over classical reconstruction in Fig. 12, are therefore upper bounds under perfect assignment, and the paper does not quantify the degradation. Because the \"most realistic ttH scenario\" result is the central performance claim, this is load-bearing. Please add a misassignment/tagging robustness study or explicitly and consistently label the scenario-6 numbers as idealized upper bounds, in the abstract and conclusions as well as in the figure captions.","section":"Sec. 4.1, App. A, Fig. 11"},{"comment":"The quantitative claims are all made on Monte Carlo events from a single leading-order pipeline: MadGraph LO with a constant K-factor of 1.13 for ttH, Delphes fast simulation for WBF, and only simple smearing (no pileup, no jet clustering, no b-tagging, no lepton isolation) for ttH scenarios 5 and 6. No systematic uncertainties are included in any of the quoted significances, so the numbers in Table 2 and Figs. 7, 12, and 13 are statistical-only projections. The relative ranking of PySR/SymbolNet versus BDT or classical reconstruction may be robust because all methods face the same simplifications, but the abstract's phrase \"at the detector level\" and the conclusion's \"most realistic scenarios\" overstate the level of realism. Please either add at least a basic treatment of dominant systematics (jet energy scale, b-tagging/misassignment, PDF/scale uncertainties) or rephrase the claims as idealized, statistical-only benchmarks.","section":"Sec. 3.2, Sec. 4.1, Table 2, Figs. 12-13"},{"comment":"It is not stated which of the two SymbolNet formulas enters the significance comparison in Table 2: the CP-odd formula in Eq. (32) or the non-CP-odd formula in Eq. (33), which the text says discriminates the SM better but would not be a valid CP probe. Since the stated goal is to construct a CP-odd optimal observable and the paper itself warns that a non-CP-odd classifier can inflate significance, each quoted significance should be accompanied by an explicit CP-parity check or at least a statement of which formula was used. Without this, the reader cannot tell whether the slight SymbolNet advantage over PySR in Table 2 reflects genuine CP sensitivity or a CP-even contamination.","section":"Sec. 3.3, Table 2"}],"minor_comments":[{"comment":"The heading \"Collin-Soper angle\" should be \"Collins-Soper angle\".","section":"Section 4 heading"},{"comment":"The axis label in Fig. 5 uses pT,j0 pT,j1 while Eq. (31) defines pT,j1 pT,j2 sin Δφ_jj; the notation should be aligned.","section":"Fig. 5 and Eq. (31)"},{"comment":"In the sentence introducing Eq. (52), \"where the is obtained\" is missing the words \"standard deviation\"; please correct.","section":"App. B"},{"comment":"The data usage is described as \"250k events for training and testing as well as 100k events for validation\"; please clarify whether the test set is used only for final evaluation or also for model/formula selection, to rule out selection-on-the-test-set bias.","section":"Sec. 3.2"},{"comment":"The statement that extra checks are needed because operations can render 4-vectors unphysical is never specified; please state how negative Minkowski norms or non-timelike vectors are handled during training and evaluation.","section":"Sec. 2.2"},{"comment":"No indication is given that the training pipeline, the modified SymbolNet implementation, or the trained formulas will be released; for a methods paper with many decimal-coefficient formulas, code/data availability would substantially aid reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a solid methods paper; the main obstacle is not the SR methodology but the gap between the \"detector-level\" framing and the idealized simulation and perfect-assignment assumptions. If the authors add a b-misassignment robustness study and soften the claims, I would support publication; as written, the central ttH performance claim needs further work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a paper worth taking seriously, and it does more than its title promises. The two genuinely new pieces are the first systematic use of symbolic regression to build detector-level CP-odd observables in WBF Higgs production, and the first analytic reconstruction of the Collins-Soper angle in ttH from reco-level inputs. The vectorized SymbolNet is a real technical extension, and the comparison between PySR and SymbolNet is informative: PySR is more data-efficient and more robust when the CP-odd signal in training is small, while SymbolNet is more expressive when smearing and radiation are present.\n\nThe paper is honest where it could have cheated. In WBF they show an explicit SymbolNet formula from a second run that is not CP-odd, and they use this to explain why the BDT significance can be inflated by CP-even components. Their learned WBF formulas reduce to the known pT,j1 pT,j2 sin(Delta phi) structure, which is a good sanity check. In ttH they show in Appendix A that even an optimal CP-odd observable built from triple products gives only Delta chi2 ~ 0.5, so they don't overclaim the CP-odd route and focus on CP-sensitive CS angle reconstruction.\n\nThe soft spots are real but not disqualifying. The ttH performance numbers in scenario 6 rely on perfect b-jet assignment and b/bbar discrimination; the formulas are explicitly asymmetric in the b and bbar momenta, and the Delta chi2 = 7.6 vs parton-level 9.4 is an upper bound. The authors acknowledge this but do not quantify the degradation. That should be stated clearly in the paper, and some robustness test (e.g. random b-jet swapping) would strengthen it considerably. Also, the event generation is leading order, detector simulation is Delphes, and there are no systematic uncertainties. These are standard limitations for a proof-of-principle ML paper, but they mean the numbers should be read as method demonstration, not as a sensitivity projection.\n\nThe in-sample nature of the MC evaluation is the same caveat that applies to any classifier trained and tested on the same generator pipeline; the checks against parton-level truth and known analytic observables mitigate it, but a release of code and data would make the results reproducible and would settle the question.\n\nWho is this for? Anyone working in Higgs CP phenomenology or on interpretable ML for LHC analyses. It deserves a serious referee: the method is novel, the comparisons are fair, and the limitations are mostly acknowledged. My recommendation is to send it to review, with the request that the authors add a discussion (or better, a numerical study) of b-jet misassignment in ttH and commit to releasing code and data.","headline":"A solid, honest methods paper on symbolic regression for Higgs CP observables; referee it, but require a b-jet assignment robustness check and code/data release.","tokens_in":27035,"tokens_out":3158,"would_cite":true,"duration_ms":30642,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that symbolic regression can learn analytic, detector-level observables for Higgs CP-violation searches that match or beat black-box networks and classical reconstruction while remaining explicitly checkable for CP parity.","keywords":["CP violation","symbolic regression","Higgs boson","optimal observable","Collins-Soper angle","ttH production","vector boson fusion","LHC physics"],"falsifier":"Re-run the scenario-6 Collins-Soper reconstruction on events where the two b-jet assignments are deliberately swapped, or enumerated over all b-to-lepton and b-to-quark pairings: if the learned formulas' $\\Delta\\chi^2$ collapses toward the classical-reconstruction value, the reported advantage is an artifact of the perfect-assignment premise rather than of symbolic regression itself.","tokens_in":26118,"feed_emoji":"⚛️","tokens_out":17412,"duration_ms":167742,"temperature":0.7,"pith_summary":"Searching for CP violation in the Higgs sector is a fundamental-symmetry test, and machine-learning classifiers built to find it are effective but hard to control. This paper argues that symbolic regression, which learns closed-form equations instead of opaque networks, can produce detector-level CP-sensitive observables that are both competitive and explicitly checkable. For vector-boson-fusion Higgs production, the learned analytic optimal CP-odd observables match or exceed a boosted-decision-tree classifier, and each formula can be inspected to verify that it is genuinely CP-odd. For top-associated Higgs production, learned formulas reconstruct the parton-level Collins-Soper angle from reconstruction-level data, recovering up to about 80 percent of the parton-level CP information in the most realistic scenario, where a classical top reconstruction keeps roughly 60 percent. A reader should care because an unambiguous probe of a fundamental symmetry requires knowing exactly what the machine learned, and a formula can be checked; a network cannot.","feed_headline":"Learned formulas recover 80 percent of Higgs CP sensitivity","feed_subtitle":"Readable equations outperform black-box classifiers on Higgs CP tests and stay checkable by hand.","key_machinery":"The central mechanism is symbolic regression itself, in two complementary implementations. PySR is an evolutionary algorithm that mutates and recombines formula trees, growing simple expressions into complicated ones; here it is trained with an added loss term that penalizes non-CP-odd outputs, so the final equation is CP-odd by construction. SymbolNet starts from a densely connected neural network whose activation functions are mathematical operators and prunes it down to a sparse, extractable equation, extended here to a vectorized version whose symbolic layers act on 4-vectors and can apply Lorentz boosts. The training targets are set by the physics: in WBF production the target is the Neyman-Pearson optimal CP-odd observable $\\omega_{CP\\text{-odd}} = p_o/p_e$, recovered by training a classifier to separate events with positive and negative $c_{H\\widetilde W}$; in $t\\bar t H$ production the target is the parton-level Collins-Soper angle $\\cos\\theta^* = \\vec p_t\\cdot\\vec n\\,/\\,(|\\vec p_t||\\vec n|)$, the CP-sensitive angle between the $t\\bar t$ system and the beam axis, which must be reconstructed from semileptonic decay products without full neutrino information.","core_discovery":"On the paper's own terms, it establishes that two complementary symbolic-regression algorithms, PySR and an extended vectorized SymbolNet, learn analytic observables for CP searches directly at the detector level. In WBF Higgs production with $H\\to\\gamma\\gamma$, the learned CP-odd observables reach significances around $7\\sigma$ for $c_{H\\widetilde W}=1$ versus the SM at 300 fb$^{-1}$, slightly above the boosted-decision-tree classifier and the classic parton-level observable $p_{T,j_1}p_{T,j_2}\\sin\\Delta\\phi_{jj}$. In $t\\bar t H$ production, the learned expressions for the Collins-Soper angle preserve the CP sensitivity of the parton-level variable: in the most realistic scenario, with an extra jet and detector smearing, SymbolNet and PySR reach $\\Delta\\chi^2 = 7.628$ and $7.491$ for excluding $\\alpha_t = 45^\\circ$, against $9.385$ at parton level, while the classical reconstruction captures only about 60 percent of the CP information. The paper further shows the two methods are complementary, with PySR the more data-efficient and stable of the pair and SymbolNet the more accurate when enough data are available, and that the learned formulas retain recognizable parton-level structures.","pith_inferences":["Because each learned observable is an entire event-level function, the same formula can be re-evaluated for any future value of the CP-violating coefficient without retraining; a natural extension the paper does not carry out is to apply it to the two companion CP-odd operators listed in Eq. (19) of the paper.","The same recipe should transfer to other latent-variable reconstructions at the LHC, wherever a parton-level CP-sensitive quantity needs an analytic, human-checkable proxy built from detector-level inputs, for instance in other $t\\bar t H$ decay channels.","The gap between the learned formulas and classical reconstruction is the quantity most likely to shrink in a real experimental setting: once combinatorial b-jet assignment is folded in, the 80-percent-versus-60-percent comparison becomes an upper bound that any experimental analysis would have to defend."],"forward_implications":["In WBF Higgs production, the learned analytic observables attain about $7\\sigma$ significance for $c_{H\\widetilde W}=1$ versus the SM at 300 fb$^{-1}$, matching or slightly exceeding both the BDT and the classic $p_{T,j_1}p_{T,j_2}\\sin\\Delta\\phi_{jj}$ baseline.","Because a learned formula is one fast-to-evaluate equation with explicitly checkable CP parity, an observed asymmetry based on it can be certified as genuine CP violation rather than a classifier artifact.","PySR's data efficiency means useful CP-odd observables can be learned from as few as 1000 training events, or from training samples with only a small CP-odd component ($c_{H\\widetilde W}=\\pm 0.1$), in regimes where the BDT and SymbolNet degrade.","In $t\\bar t H$ production, the learned Collins-Soper angle reconstructions keep roughly 80 percent of the parton-level CP information in the most realistic scenario, versus about 60 percent for classical top reconstruction, giving $\\Delta\\chi^2 \\approx 7.5\\text{--}7.6$ against the parton-level value of 9.385.","The learned formulas retain recognizable parton-level structures, PySR's around $\\sin(\\sum_i a_i p_{z,i}/\\sum_i b_i E_i)$ and SymbolNet's around a boosted ratio, which the paper reads as evidence that the same analytic skeleton carries the CP information at detector level."],"supporting_citations":[{"why":"Supplies the PySR evolutionary symbolic-regression engine used for all PySR results in the paper.","marker":"[28]"},{"why":"Supplies the SymbolNet architecture that the paper vectorizes and extends with a three-stage training procedure.","marker":"[29]"},{"why":"Defines the optimal-observable construction $\\omega_{\\text{CP-odd}} = p_o/p_e$ that the WBF classifiers are trained to approximate.","marker":"[7]"},{"why":"Establishes $p_{T,j_1}p_{T,j_2}\\sin\\Delta\\phi_{jj}$ as the parton-level optimal CP-odd observable in WBF that serves as the performance baseline.","marker":"[11]"},{"why":"Provides the detector-level WBF formula and the modified simulated-annealing acceptance used in the PySR training.","marker":"[18]"},{"why":"Supplies the boosted-decision-tree classifier used as the numerical baseline the learned WBF observables must match or beat.","marker":"[44]"},{"why":"Supplies the classical top-quark reconstruction with the $W$-mass constraint that the learned Collins-Soper angle reconstructions are compared against.","marker":"[52]"},{"why":"Provides the expected $t\\bar t H$ event yields at 300 fb$^{-1}$ used to convert reconstructed $\\cos\\theta^*$ distributions into $\\Delta\\chi^2$ CP-sensitivity values.","marker":"[56]"},{"why":"Parameterizes the top-Yukawa coupling with $c_t$ and $\\tilde c_t$, defining the CP phase $\\alpha_t$ that the $t\\bar t H$ analysis targets.","marker":"[45]"}],"fun_headline_variants":["Symbolic regression yields readable Higgs CP probes","Analytic formulas rival black-box CP classifiers at LHC","Learn CP-odd observables with equations, not black boxes","PySR and SymbolNet craft checkable Higgs CP tests","Symbolic regression recovers CP sensitivity in Higgs data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that in the $t\\bar t H$ events the two b-jets have already been correctly assigned to the lepton and the light quarks, and that $b$ and $\\bar b$ can be told apart; the learned formulas take the individual $b$ and $\\bar b$ momenta as inputs, so the reported $\\Delta\\chi^2$ values assume this ordering is known.","fun_headline_variants_meta":{"raw":{"variants":["Symbolic regression yields readable Higgs CP probes","Analytic formulas rival black-box CP classifiers at LHC","Learn CP-odd observables with equations, not black boxes","PySR and SymbolNet craft checkable Higgs CP tests","Symbolic regression recovers CP sensitivity in Higgs data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00016,"raw_usage":{"total_tokens":1209,"prompt_tokens":899,"completion_tokens":310,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":232}},"tokens_in":515,"tokens_out":310,"duration_ms":3918,"temperature":1.0,"reasoning_tokens":232,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:17:14.454851+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the scenario-6 Collins-Soper reconstruction on events where the two b-jet assignments are deliberately swapped, or enumerated over all b-to-lepton and b-to-quark pairings: if the learned formulas' $\\Delta\\chi^2$ collapses toward the classical-reconstruction value, the reported advantage is an artifact of the perfect-assignment premise rather than of symbolic regression itself.","supporting_citations":[{"cited_title":"SymbolNet: Neural Symbolic Regression with Adaptive Dynamic Pruning for Compression","cited_arxiv_id":"2401.09949","evidence_quote":"Supplies the SymbolNet architecture that the paper vectorizes and extends with a three-stage training procedure."},{"cited_title":"Davier, L","cited_arxiv_id":null,"evidence_quote":"Defines the optimal-observable construction $\\omega_{\\text{CP-odd}} = p_o/p_e$ that the WBF classifiers are trained to approximate."},{"cited_title":"Chen and C","cited_arxiv_id":null,"evidence_quote":"Supplies the boosted-decision-tree classifier used as the numerical baseline the learned WBF observables must match or beat."},{"cited_title":"$\\cal{CP}$-sensitive simplified template cross-sections for $t\\bar t H$","cited_arxiv_id":"2406.03950","evidence_quote":"Provides the expected $t\\bar t H$ event yields at 300 fb$^{-1}$ used to convert reconstructed $\\cos\\theta^*$ distributions into $\\Delta\\chi^2$ CP-sensitivity values."}],"review_version":1}