{"id":"d3aa8388-4ba2-4b21-b61b-ad7cca4d8edf","arxiv_id":"2607.18222","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"HyperIso is a new modular software tool that computes flavour observables and BSM Wilson coefficients, validated against SuperIso and MARTY, with a copula-based uncertainty engine.","lead":"This paper presents HyperIso, a new open-source software package that computes flavour-physics observables (rare B, K, D decays and muon g−2) in the Standard Model and many beyond-Standard-Model theories. It aims to make such calculations extensible and reproducible for anyone scanning new physics models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MARTY Wilson coefficients for scalar operators are off by 22–26% and the paper concedes this is not full matching; this undermines the 'general BSM calculator' claim for user-defined models.","rationale":"The paper's central claim is that HyperIso is a general BSM calculator for flavour observables, with the extension to user-defined models enabled by the MARTY interface. The native SM/THDM/SUSY implementations are validated against SuperIso at the per-mille level, and the reproducibility suite is well designed. However, the generic-model capability - the key advertised differentiator - is precisely where the weakest assumption resides. The published MARTY comparison shows discrepancies of 22-26% in the scalar Wilson coefficients C_Q1 and C_Q2, and the paper itself states that the direct projection is not equivalent to a dedicated EFT matching. Scalar operators are not a marginal sector: they enter Bs->mu+mu-, B->D(*)tau nu, and other observables that HyperIso claims to support. Therefore, if a user defines a new BSM model, the leading BSM contribution to these key observables can be wrong by tens of percent. The concern is not about disagreement with external consensus but about internal consistency: the paper's own validation demonstrates that the advertised 'general BSM calculator' capability is not yet reliable for these operator sectors. This is exactly the load-bearing assumption identified by the reader. The suggested concrete test - comparing all scalar and primed coefficients against an independent matching calculation - would settle the issue. Since the paper is honest about the limitation and the native modes work well, a conditional verdict remains appropriate rather than outright rejection.","tokens_in":31339,"tokens_out":3926,"duration_ms":34109,"concrete_test":"Recompute the MARTY-route Wilson coefficients for the released THDM benchmark and compare the complete coefficient list (C_Q1, C_Q2, C'_Q1, C'_Q2 for e, mu, tau and the B-sector C'_7,8,9,10) against the native HyperIso THDM matching at LO at the same scale and scheme. Then run the native observable pipeline with the two coefficient sets for Bs->mu+mu-, B->K*mu+mu-, and B->D tau nu; if any branching ratio or angular observable shifts by more than its quoted theoretical uncertainty, the MARTY general-model route is not yet reliable for phenomenology. An independent cross-check with MatchMakerEFT or a manual one-loop matching for C_Q1, C_Q2 in the THDM would settle whether the discrepancy is a missing operator reduction or a convention issue.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Direct projection onto the HyperIso basis is the mechanism by which user-defined BSM models are turned into Wilson coefficients (§7, §8.2). The published validation shows this route is not accurate for the scalar sector: in Table 34, for the type-II THDM BSM contribution, C_mu_Q1 and C_mu_Q2 differ from the native calculation by 22.6% and 26.1%, respectively, and acquire spurious imaginary parts (e.g., C10 gets +0.0886i in the SM block, Table 33). The text explicitly states the projection 'is not yet equivalent to a dedicated EFT matching calculation' and notes differences 'especially in sectors where the projection is numerically sensitive or where scalar and primed structures are involved.' Since scalar/pseudoscalar operators are exactly those relevant for Bs->mu+mu-, B->D(*)tau nu, and other flagship observables supported by HyperIso, a user applying the advertised MARTY workflow to a new model can obtain BSM scalar Wilson coefficients with 20-30% errors. That is not a harmless normalization detail; it will shift predicted branching ratios by more than the quoted SM uncertainties. The printed tables also omit the electron/tau scalar and C'_Q entries, so the full extent of the discrepancy is not documented. Thus the abstract's unqualified claim of a 'general BSM calculator' - including user-defined models - rests on an assumption the paper itself shows to be violated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents HyperIso, a new C++/Python/CLI/GUI software package for computing flavour observables in the SM, THDM, SUSY, and, via the MARTY framework, user-defined BSM models. The native implementations are validated against SuperIso for Wilson coefficients at the matching scale and for angular observables, showing sub-per-mille agreement in most B-sector coefficients and good agreement for F_L and P_5'. The statistical engine introduces copula-based non-Gaussian uncertainty propagation and profiling/marginalisation methods. The MARTY interface is described as generating LO Wilson coefficients by direct projection onto the HyperIso operator basis; the validation in §8.2 shows this route agrees well for C7, C8, C9 but deviates by 22–26% for the muon scalar coefficients C_Q1/C_Q2 in the type-II THDM BSM contribution, in addition to spurious imaginary parts. The paper includes a reproducible CLI test suite and frozen reference outputs.","tokens_in":31618,"tokens_out":4555,"duration_ms":42528,"significance":"If the caveats are properly addressed, HyperIso would fill a practical niche: a single, modular, well-documented code that reproduces SuperIso's physics while adding a user-facing statistical engine and a path to automated BSM matching. The native SuperIso validation is strong and the reproducibility suite (frozen references, SHA-256 metadata, fixed-seed Monte Carlo) is a genuine strength that should be highlighted. However, the advertised 'general BSM calculator' capability for user-defined models rests on the MARTY route, and the paper's own validation demonstrates that this route is not reliable in the scalar-operator sector—precisely the sector relevant for B_s→μ^+μ^- and B→D^{(*)}τν. The paper is honest about the limitation in §8.2, but the abstract and conclusion make unqualified claims that overstate the current version's capabilities. With appropriate revisions that either correct the MARTY matching or clearly scope the claims, the paper could be acceptable; as it stands, the central claim needs qualification.","major_comments":[{"comment":"The MARTY route fails precisely where it matters most. For the type-II THDM BSM contribution, C^μ_Q1 and C^μ_Q2 differ from the native calculation by 22.6% and 26.1% respectively, and acquire imaginary parts (e.g., −0.000322i and +0.000315i). These scalar operators are not peripheral: they enter B_s→μ^+μ^- and B→D^{(*)}τν, two flagship observables listed in Table 4. A user applying the advertised MARTY workflow to a new model will therefore obtain scalar Wilson coefficients with ~25% error, which will strongly bias the predicted branching ratios. Since §8.2 itself states that the projection 'is not yet equivalent to a dedicated EFT matching calculation', the abstract's unqualified claim of a 'general BSM calculator' for user-defined models is not supported. The paper should either implement a proper matching/projection for the scalar and primed sectors or restrict the claim to the valida","section":"§8.2, Tables 6, 33, 34"},{"comment":"The complete validation tables omit the electron and tau scalar entries and the full C'_Q family, with the explanation that they are omitted 'for compactness'. However, because the discrepancy is concentrated in scalar/primed structures, omitting these entries hides the full extent of the problem. The claim in the text that the complete set is 'reported in D' is inaccurate. At minimum, the paper should state explicitly whether the observed 22–26% discrepancies extend to the omitted lepton flavours and primed scalar coefficients, and provide the machine-readable files so readers can check. Without this, the validation is incomplete in exactly the sectors where the MARTY route is least reliable.","section":"Tables 33–34 and Appendix D"},{"comment":"The abstract says HyperIso 'interfaces with the MARTY framework to compute automatically BSM Wilson coefficients at leading order' and the conclusion says the MARTY route provides a 'complementary validation' of the generic-model interface. These statements do not reflect the numerical reality in Table 34: the MARTY route at LO is not a validated replacement for the native matching in the scalar sector. The 'general BSM calculator' phrasing in the title and abstract should be qualified, e.g., 'for vector/axial-vector operators; scalar/primed matching under development'. This is not a mere wording issue: it affects how a user will trust the output for a new model.","section":"Abstract and §11 Conclusion"}],"minor_comments":[{"comment":"The native validation is performed exclusively against SuperIso, which is the same group's earlier code. While this is appropriate for a reimplementation, an independent cross-check against flavio or EOS for at least a few observables would substantially strengthen the claim of correctness. If such a comparison is not feasible, a sentence explicitly acknowledging the lack of third-party validation would be helpful.","section":"§8.1"},{"comment":"The relative-difference definition uses |C^HI_i| in the denominator; for coefficients that vanish in the native calculation (e.g., C1, C2 in THDM), the comparison is undefined. The paper uses a dash for such entries, which is fine, but it would be clearer to also state this in the text.","section":"Eq. (13)"},{"comment":"Minor formatting issues: some entries contain an extra digit (e.g., '0.158 396 345 3' vs '0.158 396 345 1') and the column header alignment is inconsistent. These are cosmetic but should be cleaned up.","section":"Table 5"},{"comment":"The text says 'the Wilson coefficients are first calculated at a scale μ_W~M_W' but later the code example uses qmatch=81 GeV. The numerical value used in the validation tables should be stated explicitly (e.g., μ_W = 81 GeV or M_W = 80.36 GeV) to avoid ambiguity.","section":"§7.3"},{"comment":"The paper relies heavily on the SuperIso documentation for theoretical formulas. While this is acceptable for a software paper, it would help the reader to have at least one equation for the effective Hamiltonian used for the scalar operators, rather than referencing [17] only. This is a readability issue, not a correctness issue.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The MARTY validation issue is the gatekeeper. The authors are clearly aware of the limitation (they acknowledge it in §8.2), so the main task is to align the claims with the actual capability. If they choose to keep the 'general BSM calculator' branding, they need to fix the scalar matching; otherwise, they should substantially soften the language. I would also encourage them to release the omitted scalar/primed validation data, even if only as supplemental material. The native part and the reproducibility suite are strong and deserve credit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: HyperIso is a clean, well-engineered flavour package that reproduces SuperIso's native SM/THDM/SUSY modes to per-mille accuracy, and it ships a real reproducibility suite. The advertised MARTY route for user-defined models, however, has documented 20–30% errors in scalar Wilson coefficients, and the paper itself concedes the projection is not full matching. The abstract overstates that capability. Read it for the native modes and the new copula-based statistics; don't yet trust the 'general BSM calculator' claim for arbitrary models.\n\nThe genuinely new parts are the software architecture (hexagonal, C++ core with Python/CLI/GUI) and the statistics engine: copulas for non-Gaussian nuisances, analytic profiling of well-behaved nuisances, and three contour-projection methods. That is a real step beyond the legacy SuperIso Gaussian and first-order propagation. The validation against SuperIso is careful: Wilson coefficients agree to <0.08% across the three benchmark models, and the angular observables F_L, P5' match as shown in Fig. 5. The paper is also honest about what is not validated: the MARTY comparison in Tables 33–34 shows relative differences of 22.6% and 26.1% for C_mu_Q1 and C_mu_Q2 in the THDM BSM block, plus a spurious imaginary part in C10 from the SM block, and the text explicitly says the direct projection is 'not yet equivalent to a dedicated EFT matching calculation'. The printed tables omit electron/tau scalar and primed C'_Q entries, so the full extent of the discrepancy isn't documented. That is a real soft spot, and it is the right one to worry about: scalar operators are exactly what drive Bs->mu+mu- and B->D(*)tau nu.\n\nThe circularity concern about validation depending on SuperIso itself is minor here: the point of the native comparison is to prove the reimplementation is faithful, and the code is archived. The fit reproduction relies on the authors' own prior fit, but that's a reproducibility check, not a physics claim.\n\nVerdict: this paper deserves a serious referee. The native mode is solid and the statistics engine is a genuine contribution. The MARTY route needs either better matching or a much tempered abstract before it can be used for phenomenology in user-defined models. I'd recommend acceptance after a revision that clarifies the scope and adds the missing scalar columns.","headline":"Solid, honest re-implementation of SuperIso with a genuinely useful statistics engine; the MARTY 'general BSM' route is not yet accurate enough for scalar operators.","tokens_in":32183,"tokens_out":2156,"would_cite":true,"duration_ms":18220,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HyperIso is a new standalone program that computes flavour observables for SM, THDM, SUSY and user-defined BSM models through a single shared backend.","keywords":["flavour physics","Wilson coefficients","beyond Standard Model","effective field theory","B meson decays","muon g-2","uncertainty propagation","statistical fits"],"falsifier":"Run a dedicated one-loop EFT matching calculation for the same type-II THDM benchmark and compare its muon scalar Wilson coefficients C_Q1 and C_Q2 to the automated route: if the ~0.23–0.26 relative differences survive the comparison, the general-BSM claim for those sectors is unsupported; if they shrink to the per-mille level, the projection is validated.","tokens_in":31151,"feed_emoji":"⚛️","tokens_out":6932,"duration_ms":61866,"temperature":0.7,"pith_summary":"HyperIso is a new standalone program for computing flavour-physics observables. It inherits the physics of an earlier code but is rebuilt as a modular C++ package with common input, calculation, and statistical layers. The paper's central claim is that this one program covers the Standard Model, the two-Higgs-doublet model, supersymmetry, and, through an automated symbolic-calculus route, user-defined BSM models, returning Wilson coefficients, observables, uncertainties and fits from a single backend. This matters because it removes the usual need to re-engineer observable and statistics code for every new model. The paper validates the native calculations against the legacy code at the sub-per-mille level and documents the remaining discrepancies of the automated route in scalar-operator sectors.","feed_headline":"Unified flavour calculator covers SM, SUSY, THDM and custom BSM","feed_subtitle":"One backend runs Wilson coefficients, uncertainties and fits for standard and user-defined BSM models.","key_machinery":"The pivotal mechanism is the separation of a Wilson-coefficient pipeline from decay-specific calculators, all sharing one runtime parameter cache. A user-defined model enters as a spectrum file; the program either uses native analytical Wilson coefficients (SM, THDM, SUSY) or sends the model to an automated symbolic-calculus engine that generates leading-order amplitudes and projects them onto the effective-operator basis. The BSM contribution is isolated diagrammatically—any diagram containing at least one non-SM particle—and added to the SM contribution, avoiding the cancellation errors of a subtraction method. Around this core, a copula-based statistical engine lets non-Gaussian uncertain","core_discovery":"The central discovery is that a single program can take a user-supplied LHA-style spectrum for a BSM model and produce the corresponding flavour observables, complete with QCD running, uncertainty propagation, and maximum-likelihood fits, because the Wilson-coefficient layer is decoupled from the observable layer. For established models the coefficients are computed analytically at orders up to NNLO; for new models, an automated symbolic engine computes the leading-order BSM contribution directly in the program's operator basis, and the SM contribution is added without relying on a subtraction of two large numbers. Validation shows agreement with the parent code for native SM, THDM and SUSY","pith_inferences":["The documented ~23–26% relative differences in the muon scalar coefficients of the THDM benchmark suggest that, before the promised dedicated matching layer arrives, the generic-model route should be cross-checked against a dedicated EFT calculation whenever scalar operators matter for the observable.","If the direct-projection limitation is fixed, the architecture could become a de facto standard for fast BSM flavour scans, since the observable and statistics layers are already model independent.","A natural stress test is to feed a model with an analytically known one-loop matching, such as a simple leptoquark, and compare the automated scalar Wilson coefficients; this would quantify the projection error in sectors beyond the B-sector coefficients displayed in the paper.","The copula-based likelihood construction could be reused as a stand-alone fitting recipe for any collider observable with non-Gaussian systematics, not just flavour decays."],"forward_implications":["For a user-defined BSM model, a single spectrum file is enough to obtain Wilson coefficients, observables, uncertainty bands and best-fit contours through identical C++, Python, CLI and GUI entry points.","Existing flavour analyses that used the legacy code can be ported to the new program with per-mille-level changes in the native SM/THDM/SUSY results.","Because correlations are modelled with copulas, fits no longer need to force experimental uncertainties to be Gaussian; asymmetric or flat nuisance distributions can be used without losing the correlation structure.","The heavy multi-bin angular decays dominate runtime; the paper's parallel cache-filling and Monte-Carlo outer-loop scheduling bring typical calls from seconds to well under a second on consumer hardware.","The frozen reproducibility suite makes release-to-release comparisons of Wilson coefficients, observables and seeded Monte-Carlo results possible."],"fun_headline_variants":["HyperIso: one tool for flavour physics across all BSM models","Flavour observables for SM, THDM, SUSY, or your own BSM","HyperIso auto-computes Wilson coefficients for new BSM models","General flavour calculator: SM to custom BSM in one backend","HyperIso: modular BSM flavour calculator with auto Wilson coefficients"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The program's promise of a general BSM calculator rests on the automatically generated leading-order Wilson coefficients being accurate enough for phenomenology; the paper itself states in §8.2 that the direct basis projection is not yet a dedicated EFT matching calculation and reports relative differences of roughly 23–26% for the muon scalar coefficients in the THDM benchmark.","fun_headline_variants_meta":{"raw":{"variants":["HyperIso: one tool for flavour physics across all BSM models","Flavour observables for SM, THDM, SUSY, or your own BSM","HyperIso auto-computes Wilson coefficients for new BSM models","General flavour calculator: SM to custom BSM in one backend","HyperIso: modular BSM flavour calculator with auto Wilson coefficients"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000383,"raw_usage":{"total_tokens":1883,"prompt_tokens":776,"completion_tokens":1107,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":1010}},"tokens_in":520,"tokens_out":1107,"duration_ms":9823,"temperature":1.0,"reasoning_tokens":1010,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T15:37:27.716072+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a dedicated one-loop EFT matching calculation for the same type-II THDM benchmark and compare its muon scalar Wilson coefficients C_Q1 and C_Q2 to the automated route: if the ~0.23–0.26 relative differences survive the comparison, the general-BSM claim for those sectors is unsupported; if they shrink to the per-mille level, the projection is validated.","supporting_citations":[],"review_version":1}