{"id":"08ca053b-c293-4190-8fef-9b8a265dc216","arxiv_id":"1908.10811","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Autotunes splits event generator parameter spaces into correlated subspaces and assigns automatic observable weights, enabling iterative Professor-based tuning in higher dimensions than the standard approach.","lead":"This paper introduces Autotunes, an algorithm that splits high-dimensional Monte Carlo event generator tuning into smaller parameter subspaces and automatically weights observables. It tests the method on toy polynomials, Pythia 8 pseudo-data, and LEP tunes of Herwig 7 and Pythia 8.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The core chunking measure M (Eq. 4) is not stable under user-chosen parameter ranges; Appendix A shows Setup 3 flips the grouping, so the final tune may inherit an arbitrary range dependence.","rationale":"The reader's weakest assumption points at the same issue, and I agree. What makes it load-bearing is the causal chain: M's argmax selects the sub-tunes; the selected sub-tunes determine which parameters are fixed first and which bin weights enter Eq. 5; the bin weights change the objective function of Professor. Appendix A's flip is therefore not a cosmetic detail of the weight plots. The paper has real supporting evidence — the polynomial ideal test recovers the planted groupings, the Pythia pseudo-data comparison favors Autotunes over random and physically motivated groupings, and the LEP tunes produce parameter values that are physically interpretable. Those tests support the method in fixed user-chosen settings, which is why I would not reject the paper. But they are all conditional on the initial ranges. The range-dependence experiment proposed above (Setup 1 vs Setup 3 ranges on the same LEP tune, comparing final parameters and χ²) directly tests whether the instability matters. If the final tunes agree within the quoted 80% run-combination bands, the concern is resolved; if not, the verdict should remain conditional, with the condition being that users must validate range robustness before trusting a tune. I also note the paper explicitly declines to treat over-represented data (Section VI), but that is a stated limitation of scope, not an error in the central argument. There is no internal inconsistency in the equations; the issue is an unvalidated sensitivity of the central heuristic. Hence I do not change the reader's CONDITIONAL verdict.","tokens_in":15200,"tokens_out":11482,"duration_ms":127489,"concrete_test":"Run the full Autotunes pipeline on the same Herwig 7 cluster-model LEP data twice, with all settings identical except the initial ranges: once with the Appendix A Setup 1 ranges and once with the Setup 3 ranges, which the paper shows produce different parameter groupings. Then evaluate the final tuned parameter vector from each run at the same LEP observables using Eq. 1 with w_i=1, and compare the resulting χ² and parameter shifts against the 80% run-combination stability ranges quoted in Table III. If the two final tunes lie within those ranges of each other, the grouping flip does not propagate; if they differ beyond them, the central claim is range-dependent and the verdict should remain conditional at best.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the chunking measure M(J)=Σ_i(N_i·J)^2 (Eq. 4). The argmax of M fixes both the subspace decomposition and the bin weights (Eq. 5) used in every Professor sub-tune, so any instability in M propagates through the whole pipeline. Appendix A demonstrates exactly such an instability: with the same Herwig 7 setup, narrowing the pTmin range in Setup 3 changes the optimal grouping from (ClMaxLight,pTmin)+(αS,gCM) to (ClMaxLight,gCM)+(αS,pTmin). The authors acknowledge the flip but only claim that the weight distributions are 'fairly stable if the same parameters are found to be correlated'; that conditional does not cover the case where the parameters are not the same. Since N^i_j=S^i_j/S^N_j is a range-normalized linear-regression slope, its value and argmax depend structurally on the initial hyper-rectangle that the user must supply. The paper never tests whether later iterations erase this dependence, and no pseudo-data recovery is run with two different initial ranges that yield different groupings. If the final tune values inherit the range dependence, the central claim that Autotunes 'algorithmically' identifies the subspaces that should be tuned together is not established; at best it identifies subspaces for one particular user-chosen range.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Autotunes, an iterative algorithm for tuning high-dimensional Monte Carlo event generator parameter spaces. The method first splits the full parameter space into lower-dimensional subspaces by maximizing a measure M built from range-normalized linear-regression slopes (Section III.A, Eq. 4), assigns per-observable weights from the same slope vectors (Section III.B, Eq. 5), runs Professor on each sub-tune, and then uses the spread of the best 80% of Professor run-combinations to define updated parameter ranges for the next iteration (Sections III.C and III.D). The algorithm is tested on polynomial pseudo-data, on Pythia 8 pseudo-data where it is compared with random and physics-motivated subgroupings, and on LEP data for Pythia 8, Herwig 7 with cluster hadronisation, and Herwig 7 with Pythia 8 string hadronisation. The paper concludes that the method enables semi-automatic tuning in spaces of 18 to 22 parameters and recovers pseudo-data parameters better than the two baselines.","tokens_in":15510,"tokens_out":2030,"duration_ms":24376,"significance":"If the central claim holds, the paper makes a useful practical contribution: event generator tuning is computationally expensive, and an automatic, reproducible way to split high-dimensional parameter spaces and iterate Professor-style tunes would benefit both generator development and phenomenological studies. The strengths of the paper are its concrete algorithmic formulation, the pseudo-data recovery tests with two baseline comparisons, and the honest discussion of range dependence in Appendix A. The method is not presented with formal guarantees, but the empirical comparison on Pythia 8 pseudo-data (Figure 4) is a meaningful step beyond a single idealized example. The main open question is whether the user-chosen initial ranges, through the chunking measure M, materially affect the final tuned values; the paper does not yet close this gap for the real-data tunes.","major_comments":[{"comment":"The load-bearing chunking measure M(J)=sum_i (N_i dot J)^2 (Eq. 4) depends on the user-supplied initial hyper-rectangle through the range-normalized slopes N_i (Eq. 2). Appendix A demonstrates that changing the initial range for one parameter (Setup 3) flips the optimal grouping from (ClMaxLight, pTmin) plus (alphaS, gCM) to (ClMaxLight, gCM) plus (alphaS, pTmin). Since the subspace decomposition fixes both which parameters are tuned together and the weights in Eq. 5, this range dependence propagates through every sub-tune. The authors state that the weights are 'fairly stable if the same parameters are found to be correlated', but that conditional does not apply to the Setup 3 case, where the parameters found are not the same. No pseudo-data test is run with two different initial ranges that yield different groupings, so the paper does not establish whether later iterations erase the range dependence of the final tune. I would like the authors to either provide such a test, or explicitly qualify the claim that the method 'algorithmically' identifies the subspaces that should be tuned together.","section":"Section III.A and Appendix A"},{"comment":"The real LEP tunes report only the tuned parameter values and their stability ranges; no chi-squared values, observable-level comparisons to the default tunes, or comparisons between the Autotunes result and a conventional Professor tune are given. Without any goodness-of-fit or distribution-level evidence, the statement that the method 'produces plausible LEP tunes' (Tables II and III) is not quantitatively supported. Please add, at minimum, the final chi-squared for each tune and a comparison of key observables (for example thrust, aplanarity, and charged multiplicity) against the defaults used as starting points.","section":"Section V and Tables II-III"},{"comment":"The pseudo-data comparison in Figure 4(d) is the strongest evidence for the method, but the statement that 'the iterated Autotunes method improves the agreement' is only shown for a single random pseudo-data realization and three runs per baseline. The summed deviation metric is also not normalized per parameter, so the plot conflates the number of badly recovered parameters with the magnitude of deviations. I recommend reporting per-parameter deviations or a normalized chi-squared-like quantity, and, if computationally feasible, repeating the pseudo-data recovery for at least one additional random true parameter point to confirm that the advantage over the physically motivated baseline is not specific to one point.","section":"Section IV.B and Figure 4(d)"},{"comment":"The weight definition w_i = (N_i dot J_Step)^2 / sum_j N_j^i is central to the observable weighting, but it is not derived or justified beyond the heuristic sentence describing the numerator and denominator. In particular, the denominator sum_j N_j^i is not defined with explicit index ranges, and the reader cannot tell whether the sum runs over parameters of the current sub-tune or over all parameters. Please state the index ranges explicitly and give at least a short argument for why this particular normalization is preferable to alternatives such as normalizing the numerator by its maximum over bins.","section":"Section III.B, Eq. (5)"}],"minor_comments":[{"comment":"The paper would benefit from a notation table or glossary; symbols such as Nsearch, the sub-tune dimension n, and the iteration count are introduced in the text but never collected in one place.","section":"General"},{"comment":"In the polynomial test, the correlation matrices C_a are diagonal with entries k>1, so the 'correlated parameters' are actually parameters that independently enhance the same set of bins. Please clarify in the text that this is intended as a simplified model of correlation on the observable level and not a test of true multi-parameter correlations beyond the linear-regression sense used in the algorithm.","section":"Section IV.A"},{"comment":"The caption of Figure 5 says 'The dashed lines correspond to Setup 2, which gives a same grouping of parameters as Setup 1', but Setup 2 also uses narrower ranges for two parameters; the text says Setups 1 and 2 give the same pairing, yet the figure shows Setup 1 and Setup 2 as dashed while Setup 3 is solid. Please make the caption and main text consistent about which setup is shown dashed and why.","section":"Appendix A, Figure 5"},{"comment":"There are several typographical and formatting issues: 'ﬁrst order approximation' in Section II.B should be 'first-order approximation'; 'runcombinations' is used without a hyphen or space in several places; the footnote about the implementation URL (gitlab.com/Autotunes) is not referenced in the text; and the sentence in Section IV.A 'As the result of each full tune serving as input to a next iteration' is grammatically incomplete.","section":"Minor typographical issues"}],"recommendation":"major_revision","confidential_remarks":"The paper is a competent methods contribution with a useful empirical demonstration on pseudo-data, but the range-dependence issue in Appendix A is not a side remark: Eq. (4) is the mechanism by which the algorithm claims to be automatic, and the Setup 3 flip shows that the grouping is not stable under a user-controlled choice. Because the real-data section lacks any goodness-of-fit validation, the manuscript currently supports the claim 'the method can recover pseudo-data parameters better than random/physics-motivated splits' but not fully 'the method algorithmically identifies the correct subspaces for real generator tuning'. I think this is fixable within the scope of a revision: add a range-sensitivity pseudo-data test, report chi-squared and observable comparisons for the LEP tunes, and clarify Eq. (5). The citation list is adequate for a hep-ph methods paper, and the authors are appropriately candid about the limitations in Appendix A."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my take on Bellm & Gellersen. The paper delivers a genuinely new algorithm for high-dimensional MC generator tuning: it splits the parameter space into subspaces by maximizing a slope-based measure M, assigns bin-wise weights automatically, and iterates with refined ranges. That is a real step beyond Professor's manual grouping and weighting. The pseudo-data tests are the strongest part: for an 18-parameter Pythia 8 setup, the Autotunes method recovers the injected parameters with clearly lower summed squared deviation than either random or physics-motivated sub-tunings. The polynomial toy test also picks out the built-in correlated blocks. For practitioners, the LEP tunes for Pythia 8, Herwig 7, and the Herwig+Pythia string hybrid are a useful resource, even without quantitative validation.\n\nThe soft spots are proportionate. The biggest one is the range dependence of the chunking step. Appendix A shows that narrowing the pTmin range flips the optimal grouping from (ClMaxLight,pTmin)+(αS,gCM) to (ClMaxLight,gCM)+(αS,pTmin). Since M's argmax fixes both the subspaces and the weights in Eq. (5), that flip can propagate into the final tune. The authors acknowledge the flip but don't test whether the iteration procedure erases it, and no pseudo-data recovery is run with two different initial ranges that give different groupings. That's a genuine gap, not a fatal one: the pseudo-data evidence shows the method works for a fixed range, and the range dependence is a feature of any local-gradient heuristic. Still, the phrase \"algorithmically splits the parameter space\" oversells it a bit; it's algorithmic given a user-chosen hyper-rectangle.\n\nThe second gap is evaluation of the real tunes. The paper reports tuned parameter values and stability ranges but no χ² comparison to the generator defaults. Without that, a reader can't tell whether the LEP tunes actually improve data description. This is addressable in revision. Also, the promised git repository is not yet available, so reproduction is limited; that matters in this field.\n\nOverall, the central idea is sensible and the evidence is honest, including the self-critical appendix. The paper deserves a serious peer review. I'd recommend sending it with the request that the authors provide code, a robustness test over initial ranges, and quantitative comparisons of the new tunes to the defaults. If those land, this will be a useful citable method.","headline":"A useful and honest extension of Professor tuning with a real algorithmic contribution, but the range-dependence of the chunking measure deserves a robustness check before the 'automatic' claim fully holds.","tokens_in":15978,"tokens_out":3391,"would_cite":true,"duration_ms":35935,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Autotunes algorithm makes it practical to tune Monte Carlo event generators with 18–22 parameters by splitting the parameter space into correlated subspaces, weighting observables per sub-tune, and iterating the Professor method.","keywords":["Monte Carlo event generators","parameter tuning","Professor method","high-dimensional optimization","Pythia 8","Herwig 7","Lund string model","LEP observables"],"falsifier":"Run the full Autotunes pipeline on the same pseudo-data with two different but reasonable initial ranges and compare the recovered parameters; if the grouping flip shown in Appendix A alters the final tuned values by more than the quoted stability bands, the claim that the algorithm reliably identifies the right subspaces is falsified. The missing observation is whether range-induced grouping changes propagate into the final tune.","tokens_in":14992,"feed_emoji":"⚛️","tokens_out":8577,"duration_ms":82495,"temperature":0.7,"pith_summary":"The paper presents Autotunes, an algorithm for tuning Monte Carlo event generators when the number of parameters is too large for the standard Professor method to handle in one pass. It splits the parameter space into lower-dimensional subspaces by measuring, from random samples, how strongly each parameter variation affects each observable bin; it weights bins so each sub-tune focuses on the observables its parameters actually influence, and it iterates the tuning with narrowed ranges. In idealized tests on polynomial pseudo-data the algorithm recovers the correlated parameter groupings it was built from, and in a pseudo-data test with Pythia 8 it recovers true parameter values better than either random or physically motivated sub-tunings. The real-life demonstrations are LEP tunes of Pythia 8, of Herwig 7 with its default cluster hadronisation, and of Herwig 7 showers combined with Pythia 8's Lund string model. This matters because event-generator tuning is a bottleneck for precision collider phenomenology: a semi-automatic retuning loop lets model improvements be tested against data quickly.","feed_headline":"New algorithm tunes event generators in 22 dimensions at once","feed_subtitle":"Clusters correlated parameters and weights key observables to retune Pythia 8 and Herwig 7.","key_machinery":"The central object is the normalized slope vector $\\vec{N}_i$, one per observable bin $i$, obtained by linear regression of the normalized generator response on normalized parameters and then divided element-wise by the total slope vector $\\vec{S}_N$. Its components encode which parameters influence which observables, and the shared normalization makes parameters that act on the same bins look correlated. $\\vec{N}_i$ drives the whole pipeline: the chunking measure $M(\\vec{J}) = \\sum_i (\\vec{N}_i\\cdot\\vec{J})^2$ picks the subspace $\\vec{J}$ that maximizes the squared projection of the slope vectors, and the same vectors set the bin weights $w_i$ that enter the Professor $\\chi^2$. The second essential piece is the iteration: the best 80% of the Professor runcombination fits define new parameter ranges, with a 20% margin, so each pass narrows the hyper-rectangle and improves the validity of the polynomial interpolation.","core_discovery":"On its own terms, the central claim is that high-dimensional event-generator tuning can be reduced to a sequence of low-dimensional Professor tunes by an algorithmic pre-processing step. Autotunes samples the parameter hyper-rectangle, normalizes parameter and observable ranges, and performs linear regressions to obtain a slope vector $\\vec{S}_i$ for each observable bin $i$; each slope vector is divided component-wise by the summed slope vector $\\vec{S}_N$ to give $\\vec{N}_i$. The chunking measure $M(\\vec{J})=\\sum_i (\\vec{N}_i\\cdot\\vec{J})^2$ selects, for a prescribed subspace dimension $n$, the set of parameters that should be tuned together, and the same $\\vec{N}_i$ enter the observable weights $w_i=(\\vec{N}_i\\cdot\\vec{J}_{\\text{step}})^2/\\sum_j N_i^j$ for each sub-tune. Tuning proceeds subspace by subspace with Professor's runcombination method, parameters from completed steps are held fixed, and the spread of the best fits defines narrowed ranges for the next iteration. In ideal polynomial tests the intended groupings are recovered; in the Pythia 8 pseudo-data test the iterated Autotunes method gives the smallest summed squared deviation from the true parameters; and in LEP tunes of 18 or 22 parameters the method yields stable values for the strong coupling $\\alpha_S$ while flagging parameters the data barely constrain, such as the strange-quark fragmentation parameter $\\mathtt{aExtraSQuark}$ and the shower cutoff $\\mathtt{pTmin}$.","pith_inferences":["The paper does not test whether the range-induced grouping flip documented in Appendix A changes the final tuned values; if it does, the algorithm's subspace split is not stable in the one place a user has control.","The same slope-vector machinery could be repurposed to cluster observables and down-weight over-represented data, an extension the paper explicitly postpones; one could test whether that changes tunes on large LEP data sets.","Because the generator is treated as a black box probed by random sampling, the method should carry over to other expensive simulators with many parameters, provided the polynomial approximation holds after range narrowing.","A concrete next test is to weight merged higher-order samples as the paper suggests, suppressing observables dominated by perturbative uncertainties, and check whether the tuned strong coupling $\\alpha_S$ shifts."],"forward_implications":["Retuning after a model change becomes a scripted pipeline rather than a manual expert exercise, because the subspace split and observable weights are produced algorithmically.","Professor-based tuning extends from roughly ten parameters to the 18- and 22-parameter examples demonstrated here.","Different hadronisation models can be tuned to the same data with the same algorithmic bias, making model comparisons less dependent on human tuning choices.","A small number of iterations is enough: in the pseudo-data tests the second iteration visibly improves parameter recovery, with later iterations giving only minor changes.","The per-parameter stability ranges produced by the runcombination spread identify which parameters the data actually constrain, such as the strong coupling, and which it leaves loose."],"supporting_citations":[{"why":"Supplies the Professor polynomial-interpolation and runcombination machinery that every sub-tune uses.","marker":"[21]"},{"why":"Define the Pythia 8 generator and the parameters whose values are tuned.","marker":"[7, 8]"},{"why":"Define the Herwig 7 generator and the two shower implementations used in the LEP retunes.","marker":"[2–5]"},{"why":"Supply the Lund string model used both in the Pythia 8 tune and in the Herwig 7 plus strings tune.","marker":"[33, 34]"},{"why":"Provides the physically motivated Monash tune parameter grouping used as the comparison baseline for pseudo-data recovery.","marker":"[22]"},{"why":"Provides the Rivet toolkit through which generator output is compared with reference data in the tuning workflow.","marker":"[32]"}],"fun_headline_variants":["High-dimensional tuning of event generators via subspace chunking","Autotunes: splitting parameters to tune event generators in 22D","Subspace-based method tunes Monte Carlo event generators in high dimensions","Retune Pythia and Herwig with subspace-chunked parameters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Changing the user-chosen initial ranges can flip which parameters are grouped together, even though the underlying generator is unchanged (the paper's Appendix A shows such a flip); the whole method therefore rests on the assumption that the initial ranges produce slope vectors that correctly identify which parameters belong together.","fun_headline_variants_meta":{"raw":{"variants":["High-dimensional tuning of event generators via subspace chunking","Autotunes: splitting parameters to tune event generators in 22D","Subspace-based method tunes Monte Carlo event generators in high dimensions","Retune Pythia and Herwig with subspace-chunked parameters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000769,"raw_usage":{"total_tokens":3437,"prompt_tokens":1005,"completion_tokens":2432,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":2359}},"tokens_in":621,"tokens_out":2432,"duration_ms":19691,"temperature":1.0,"reasoning_tokens":2359,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:33:31.297250+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the full Autotunes pipeline on the same pseudo-data with two different but reasonable initial ranges and compare the recovered parameters; if the grouping flip shown in Appendix A alters the final tuned values by more than the quoted stability bands, the claim that the algorithm reliably identifies the right subspaces is falsified. The missing observation is whether range-induced grouping changes propagate into the final tune.","supporting_citations":[{"cited_title":"Abreu et al","cited_arxiv_id":null,"evidence_quote":"Supplies the Professor polynomial-interpolation and runcombination machinery that every sub-tune uses."},{"cited_title":"The Perugia Tunes","cited_arxiv_id":"0905.3418","evidence_quote":"Provides the physically motivated Monash tune parameter grouping used as the comparison baseline for pseudo-data recovery."}],"review_version":1}