{"id":"4e4ca430-86a1-4c18-bb07-aa9969102ce5","arxiv_id":"2501.13544","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A global utility function for multi-purpose experiments is feasible and is already implicitly encoded in trigger decisions, run schedules, and optimization weights.","lead":"This paper argues that the goals of multi-purpose physics experiments can be captured in a single utility function, which is needed for AI-guided design. It supports the argument with trigger and run-schedule examples from CDF and BES-III and with an idealized gamma-ray array optimization study.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SWGO idealizations are the sole quantitative bridge from 'implicit utility' to design-space optimization; the paper does not show they preserve real layout trade-offs.","rationale":"The reader's verdict is CONDITIONAL with moderate confidence, and the weakest assumption is the SWGO idealization. I agree that this is the most load-bearing quantitative vulnerability. The central claim has two linked parts: (1) utility functions are feasible because collaborations already make quantitative trade-offs, and (2) therefore such utility functions do not hinder global design optimization. The evidence for (1) includes the BES-III run-schedule table and CDF trigger discussions, which are real but concern discrete operational allocations rather than differentiable design-space utilities. The only demonstration of (2) is the SWGO gradient-based optimization, which depends on idealization assumptions that are asserted, not validated, in the preprint. Thus if the SWGO idealizations do not preserve the layout trade-offs of a physical water Cherenkov array, the paper does not establish that a consensus utility can be effectively used in high-dimensional design optimization, and the general conclusion would need to be weakened. I did not find an internal inconsistency or a flaw in the conceptual argument; the concern is specifically about the strength of the empirical bridge, which matches the reader's call for external validation and reproducibility. Therefore I keep the reader's CONDITIONAL verdict unchanged rather than moving to ACCEPT or REJECT. The proposed test directly checks whether the idealization is innocuous by replacing it with a realistic detector response and comparing outcomes in the figures and table that carry the quantitative weight of the paper.","tokens_in":22455,"tokens_out":6285,"duration_ms":59533,"concrete_test":"Re-run the Sec. 3.2 SWGO optimization (S = +1.0 and S = -1.0, 1000-epoch gradient descent, 330 macro-tanks with 120-degree symmetry) after replacing the 100% efficiency / perfect discrimination assumptions with a realistic water-Cherenkov response, e.g., a detection efficiency that falls with distance and a finite e/γ-vs-muon misidentification rate calibrated from SWGO simulations [65]. Compare the optimized layouts and the UGF/UIR/UPR values to those in Fig. 3 and Table 4. If the utility differences between layouts optimized for different S change by more than the quoted ~10% or the qualitative layout features (hexagonal outer tanks, radial extent) are not preserved, the assertion in Sec. 3.1 that the simplifications do not affect the conclusions fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Sec. 4) is that a global utility function is 'not only possible, but implicitly done already' and hence 'not a significant hindrance.' The strongest quantitative support for the design-optimization link is the SWGO study (Sec. 3), which uses a closed-form radial shower parametrization from the authors' own [66], 100% detection efficiency, and perfect hadron discrimination (Sec. 3, first paragraph). The paper concedes this is 'a considerable simplification' (Sec. 3.1) and asserts that conclusions are unaffected, but gives no demonstration. The CDF trigger example (Sec. 2.2) and the BES-III run-schedule table (Table 3) show that collaborations make quantitative operational decisions, but these are discrete allocations (trigger prescales, run days), not differentiable scalar utilities over the O(10^3-10^4)-dimensional design space that the paper motivates in Sec. 1.1. The inference from 'collaborations can agree on run days' to 'a consensus utility over design parameters can be defined and optimized' thus rests entirely on the idealized SWGO example. If finite efficiency, imperfect e/mu discrimination, or the approximate shower model couple differently to detector layout than assumed, the demonstration does not transfer to real experiments, and the central claim lacks validated quantitative support.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that multi-purpose experiments in fundamental science can, and implicitly already do, define a global utility function that weights their different scientific goals. The argument is developed through three case studies: the CDF trigger system, where bandwidth allocation forced quantitative trade-offs among physics goals; the BES-III run schedule, which explicitly allocates run days and luminosities to different center-of-mass energies; and a SWGO-like water Cherenkov array, where the authors use their own idealized shower model and gradient-based optimization to show that different energy spectra lead to different optimized layouts. The paper concludes that the definition of a global utility function is not only possible but already implicit in many experimental situations, and therefore should not be a significant hindrance to global optimization programs.","tokens_in":22702,"tokens_out":2487,"duration_ms":24906,"significance":"If the central claim holds, the paper provides a useful conceptual bridge between the multi-objective nature of large experiments and the gradient-based, AI-assisted design-space exploration advocated in Sec. 1. The historical examples (CDF trigger meetings, BES-III run-schedule table) are concrete and externally grounded, and the SWGO demonstration, despite its simplifications, is an actual working example of optimizing a multi-component utility over a continuous layout space. The paper also honestly acknowledges several limitations, including the ill-defined nature of the citation-based proxies in Sec. 2.3 and the 'considerable simplification' of the SWGO model in Sec. 3.1. The main weakness is that the only quantitative demonstration of the full design-optimization link rests on an idealized detector model whose robustness to realistic detector effects is asserted but not tested.","major_comments":[{"comment":"The quantitative demonstration that a multi-target utility can drive useful design optimization relies entirely on an idealized detector model: 100% detection efficiency, perfect discrimination between the soft and hard shower components, and a closed-form radial shower parametrization from the authors' prior work [66]. The paper states in Sec. 3.1 that 'the conclusions we drew above on the complexity of the task are not affected by those simplifications,' but no sensitivity study or argument is provided to show that realistic efficiency, imperfect particle identification, or the approximate shower model preserve the layout trade-offs. This is load-bearing because the central conclusion in Sec. 4 extrapolates from this idealized SWGO example to real experiments; without a robustness check or a clearly scoped claim, the transfer is not established.","section":"Secs. 3 and 3.1"},{"comment":"The CDF citation-fraction analysis is explicitly acknowledged to carry 'a much more significant additional error from their ill-defined nature,' and the actual trigger bandwidth allocations are admitted to be unknown ('those numbers are not easy to determine with precision'). Consequently, the derived relative values of 57.0%, 18.2%, and 24.7% for the three trigger classes are at best qualitative. The paper itself concedes this in Sec. 2.3, but the conclusion in Sec. 4 still leans on these examples as demonstrations that a global utility is 'implicitly done already.' The qualitative lesson is sound, but the quantitative framing should either be removed or replaced with a clearly labeled illustration.","section":"Sec. 2.3 and Table 1"},{"comment":"The utility function in Eq. (3) depends on free coefficients λ_GF, λ_IR, λ_PR, and on the gradient-scaling factors w_GF, w_IR, w_PR, with no procedure given for how a collaboration would arrive at these numbers from its stated scientific priorities. The BES-III run-schedule example (Table 3) shows that collaborations can agree on discrete allocations of run days, but it does not demonstrate agreement on the differential weights needed for a differentiable scalar utility over a continuous O(10^3)-dimensional design space. The paper's central inference from 'collaborations can agree on run days' to 'a consensus utility over design parameters can be defined and optimized' therefore needs either a concrete elicitation procedure or a more modest conclusion.","section":"Secs. 1.2, 3.1, and Eq. (3)"}],"minor_comments":[{"comment":"Typo: 'impedence' should be 'impedance'.","section":"Sec. 2.2"},{"comment":"Typo in the caption: 'citiations' should be 'citations'.","section":"Table 1"},{"comment":"Two typos: 'whooping' should be 'whopping' in the discussion of QCD triggers, and 'affering' should be 'belonging' (or similar) in the sentence about the three publications in the broad QCD category.","section":"Sec. 2.3"},{"comment":"Reference [4] appears in the reference list without a title or source; either complete it or remove the citation.","section":"References"},{"comment":"The text says 'electro-positron collisions'; the standard term is 'electron-positron collisions.'","section":"Sec. 2.4"},{"comment":"The sentence 'by only examining the utility values of those eight solutions one might be led to believe that they lay close to the Pareto front' is clear, but Figure 3 would be easier to interpret if the axes were labeled and the benchmark points were identified by name in the figure itself.","section":"Sec. 3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper's historical examples are engaging and the BES-III run-schedule evidence is genuinely external and concrete. The main risk is that the concluding claim is broader than the validated evidence: the SWGO demonstration is the only quantitative support for the design-optimization link, and it is built on idealizations that are acknowledged but not tested. A sensitivity analysis or a clear restriction of the scope of the claim would be sufficient to address this. The paper fits the journal's scope as a methods/vision contribution, but it currently overstates the strength of the quantitative support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First: the paper's central argument holds. Utility functions are not a show-stopper for AI-driven experiment design, because collaborations already make quantitative trade-offs between goals when they set trigger menus and allocate run days. The BES-III run schedule is a genuinely concrete example. Second: the quantitative SWGO part is weaker than the framing suggests and should not be the load-bearing wall. That said, the paper knows this—it explicitly admits the model is a considerable simplification—so the weakness is acknowledged, not hidden.\n\nWhat is actually new: the CDF citation analysis (57.0/18.2/24.7 across trigger families) is an original and honest attempt to reconstruct a posteriori scientific value; the authors even flag the error from its ill-defined nature. The SWGO energy-spectrum cross-comparison (Table 4) is also new. The discussion of flexible vs. fixed arrays for SWGO is useful. Credit where due: the paper is clearly written, the examples are well chosen, and it does not oversell the CDF numbers.\n\nSoft spots. The stress-test note is on target: the only quantitative link between 'implicit utilities' and differentiable layout optimization is the SWGO demonstration, which assumes 100% efficiency, perfect hadron discrimination, and reuses the authors' own shower parametrization from [66]. The paper asserts these simplifications don't change the conclusions but provides no sensitivity study. A referee should ask for that. The CDF citation grouping is subjective, though the paper says so. Minor: the text overuses the word 'ground-breaking' and there is some self-citation, but not beyond what a methods-first paper would do. No fatal flaw; the central conceptual claim survives.\n\nWho is this for: people working on AI-assisted detector co-design, experimental design methodology, and perhaps the MODE/EUCAIF community. It is not a paper that produces a new physics result, but it removes a conceptual obstacle and gives two worked historical examples. I would bring it to a reading group only if the group cares about the co-design agenda. I would probably not cite it myself, since I would not want to rely on the idealised SWGO numbers. But it deserves serious peer review: the discussion is sensible, the evidence is presented honestly, and the field needs this conversation. Recommendation: send to review, with a request for a sensitivity analysis of the SWGO idealisations and, if possible, a code/data release for the optimization runs.","headline":"Convincing conceptual argument that experiment-wide utility functions already exist in practice, but the only quantitative demo is idealised and needs sensitivity checks.","tokens_in":23292,"tokens_out":2832,"would_cite":false,"duration_ms":26679,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that the global utility functions needed for AI-driven experiment design are not only possible but already implicitly present in how multipurpose experiments allocate trigger bandwidth and running time.","keywords":["utility function","multipurpose experiments","detector design optimization","Pareto front","gradient-based optimization","trigger systems","experiment run scheduling","gamma-ray observatory layout"],"falsifier":"Re-run the same layout optimization with a realistic detector response—finite efficiency, imperfect gamma/hadron separation, trigger thresholds, and mis-reconstructed background events—and compare the optimized array's utility against the benchmark layouts; if the advantage shrinks or reverses, the demonstration that utility-weighted gradients are sufficient for global design optimization would collapse.","tokens_in":22236,"feed_emoji":"🔭","tokens_out":10559,"duration_ms":92876,"temperature":0.7,"pith_summary":"The paper is trying to establish that the main conceptual obstacle to AI-driven design of large physics experiments—agreeing on the relative scientific worth of different goals—is not actually an obstacle. It argues that multipurpose experiments already perform this weighting implicitly whenever they decide which events to store, which trigger thresholds to raise, or how many days to spend at each beam energy. The paper makes the weighting explicit as a sum of utility terms and shows that gradient-based optimization of a gamma-ray observatory layout can beat expert benchmark arrays on targeted performance measures. Its conclusion is that defining a global utility function is possible, is in many cases already done, and should not block automated exploration of design spaces.","feed_headline":"Experiments already define global utility functions","feed_subtitle":"Case studies and a gamma-ray array demo show weighted scientific goals can drive gradient-based optimization.","key_machinery":"The central object is the additive utility function $U = \\lambda_1 U_1 + \\lambda_2 U_2 + \\cdots$, where each $U_i$ measures one scientific goal and each $\\lambda_i$ states the experiment's relative priority. The numerical demonstration uses $U_1 = \\lambda_{GF}U_{GF} + \\lambda_{IR}U_{IR} + \\lambda_{PR}U_{PR}$ for flux precision, energy resolution, and pointing resolution, moved by gradient-ascent updates $x_i \\to x_i + \\eta_i\\, dU/dx_i$. Dynamic scaling factors $w_i$ balance the disparate gradient magnitudes so that no single term such as flux dominates the layout updates. A closed-form parametrization of the ground density of shower particles makes all gradients computable, which is what allows the authors to search the layout space continuously instead of sampling a handful of benchmark designs.","core_discovery":"The central claim is that an experiment-wide utility function can be written down as a weighted sum of scientific goals, and that large collaborations already make exactly such weighted choices in practice. The paper reconstructs the implicit utility of a hadron-collider experiment from the citation counts of its publications, finding that the trigger streams that produced the highest-impact results were not the ones consuming the most bandwidth, and it quotes the published run plan of an electron-positron collider experiment as a case where the collaboration tabulates precise days and luminosities per energy point according to physics motivation. In the numerical part, the paper optimizes the layout of a high-altitude water Cherenkov gamma-ray array by gradient ascent on $U = \\lambda_{GF}U_{GF} + \\lambda_{IR}U_{IR} + \\lambda_{PR}U_{PR}$, against the benchmark layouts, and finds gains of roughly $+19\\%$ on integrated energy resolution and $+35\\%$ on pointing resolution when optimizing a two-term utility. The same machinery shows that the optimal layout depends on the assumed energy spectrum of the source, with cross-spectrum utility losses of about 8–10%. The paper concludes that a global utility function is not only possible but already implicit in many experimental situations, and therefore should not be considered a significant hindrance to global optimization programs.","pith_inferences":["Implicit extension: because the utility machinery is layout-agnostic, the same weighted-sum-plus-gradient approach could be transferred to other sparse arrays (radio telescopes, cosmic-ray networks, seismic or gravitational-wave configurations) where the cost of moving elements is high.","Implicit in the citation analysis: the utility encoded in trigger menus can drift from the collaboration's stated priorities, so periodically re-weighting the $\\lambda$ terms as science evolves may matter as much as the initial specification.","Not pursued by the paper: an observatory facing unknown source spectra could hedge by optimizing over a mixture of plausible spectra, or by choosing a layout that minimizes the worst-case utility loss across spectra, before committing construction funds.","Testable outside the paper: a collaboration that publishes its provisional global utility and invites external gradient-based proposals could compare those proposals in full simulation against internal designs, providing a direct test of whether practical obstacles remain."],"forward_implications":["A collaboration can explicitly write down an experiment-wide utility function with weights, because an existing collider experiment's published run schedule already allocates data-taking days by scientific goal.","Gradient-based movement of detector units, driven by a weighted utility sum, can improve energy resolution and pointing resolution by roughly 19% and 35% over expert benchmark layouts without degrading the third component.","The optimal detector layout is tied to the scientific prior on the source energy spectrum; choosing the wrong prior costs about 8–10% of the array's utility on the actual spectrum.","Staged construction that locks in an early layout prevents the extended array from reaching the performance of a layout optimized from the start for the final science goals."],"supporting_citations":[{"why":"Shows that a simple geometric exploration of a single-purpose scattering experiment halves the quoted resolution uncertainty, providing the motivating example for automated design.","marker":"[14]"},{"why":"Supplies the closed-form shower model, the benchmark array layouts, and the optimization program on which the numerical demonstration depends.","marker":"[66]"},{"why":"Contains the tabulated future run plan used as evidence that a collaboration can explicitly allocate running time by scientific goal.","marker":"[47]"},{"why":"Fixes the high-altitude site and the broad science case that define the utility components for the gamma-ray array study.","marker":"[65]"},{"why":"Documents the intractability of likelihoods in stochastic data-generating processes, the obstacle the paper argues differentiable surrogates now bypass.","marker":"[10]"},{"why":"Describes the differentiable-programming pipeline for end-to-end detector optimization that a global utility function would enable.","marker":"[6]"}],"fun_headline_variants":["Experiments already have implicit global utility functions","Weighted scientific goals drive experiment design","Gamma-ray array optimization with explicit utility","Utility-based design improves gamma-ray resolution","AI-ready utility functions for multipurpose experiments"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the idealized numerical setup—perfect detector efficiency, perfect separation of gamma-ray events from hadronic background, and a closed-form shower model—preserves the layout trade-offs of a real water Cherenkov array; the paper asserts this but does not demonstrate it with a realistic simulation.","fun_headline_variants_meta":{"raw":{"variants":["Experiments already have implicit global utility functions","Weighted scientific goals drive experiment design","Gamma-ray array optimization with explicit utility","Utility-based design improves gamma-ray resolution","AI-ready utility functions for multipurpose experiments"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000714,"raw_usage":{"total_tokens":3222,"prompt_tokens":966,"completion_tokens":2256,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":2193}},"tokens_in":582,"tokens_out":2256,"duration_ms":17881,"temperature":1.0,"reasoning_tokens":2193,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:50:16.984027+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same layout optimization with a realistic detector response—finite efficiency, imperfect gamma/hadron separation, trigger thresholds, and mis-reconstructed background events—and compare the optimized array's utility against the benchmark layouts; if the advantage shrinks or reverses, the demonstration that utility-weighted gradients are sufficient for global design optimization would collapse.","supporting_citations":[{"cited_title":"Toward the End-To-End Optimization of the SWGO Array Layout","cited_arxiv_id":"2310.01857","evidence_quote":"Supplies the closed-form shower model, the benchmark array layouts, and the optimization program on which the numerical demonstration depends."},{"cited_title":"The Southern Wide-Field Gamma-ray Observatory (SWGO)","cited_arxiv_id":"2012.13740","evidence_quote":"Fixes the high-altitude site and the broad science case that define the utility components for the gamma-ray array study."}],"review_version":1}