Pith. sign in

REVIEW 2 major objections 4 minor 300 references

An agentic system with a physicist in the loop can add a new LHC measurement to a global SMEFT fit and tighten constraints on top-quark couplings.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 04:23 UTC pith:PSWVUUMX

load-bearing objection A genuinely useful agentic re-casting toolkit with an honest validation appendix, but the central SMEFT scan may be running in the exact silent-failure mode the paper itself documents. the 2 major comments →

arxiv 2607.22813 v1 pith:PSWVUUMX submitted 2026-07-24 hep-ph

Agentic Re-Casting using Agentic Re-Simulations

classification hep-ph
keywords agentic AIanalysis re-castingSMEFTSFittertop quark sectorttZ productionrepeatable benchmarksilent failure modes
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper shows that the months-long expert workflow of LHC analysis re-casting—extracting a published measurement, re-simulating it under a new theory hypothesis, and refitting a global likelihood—can be carried out by an LLM agent system supervised by a physicist. It demonstrates this with SFitter's global top-sector SMEFT fit: the agents select and extract an ATLAS ttZ differential measurement, re-simulate it at NLO, scan 21 dimension-six operators, assemble the datacard, and run the global analysis, which tightens the bounds on the top-electroweak Wilson coefficients by 34–42%. The paper also reports a blind benchmark in which six independent agent runs recover the Wilson coefficients of an injected coloron signal, consistent with the truth within statistical fluctuations. If this holds, published LHC measurements could be folded into global fits quickly, reproducibly, and with a physicist retaining control at every step.

Core claim

On its own terms, the paper establishes that SFitterAgents—built on the MadAgents.v3 consultant architecture—can perform the complete re-casting chain for a new measurement: it picks the ttZ measurement, extracts per-bin values and uncertainties, re-simulates the SM signal at parton and particle level, scans the SMEFT Wilson coefficients to extract the per-bin kappa parametrization, validates and assembles the SFitter datacard, and runs the exclusive-likelihood global fit. Adding the normalized pT(Z) spectrum tightens the profiled constraints on Cφt, C−φQ, CtZ and CtW by 34–42%, while m(ttZ) alone gives 19% on CtZ; the agent flags a 3σ underfluctuation in one pT(Z) bin that pulls Cφt to larg

What carries the argument

SFitterAgents, an agentic system built on MadAgents.v3, whose orchestrator routes each query to specialized consultant, worker, and reviewer subagents. Its operating principles—source grounding in the locally installed code, lasting memory records, completion-vs-correctness checks, recorded confidence, and adversarial review—are the mechanism that keeps silent simulation failures from corrupting the physics. The re-casting chain itself is carried by four steps: measurement extraction, SM re-simulation, SMEFT scan with per-bin kappa extraction (the linear and quadratic dependence of each bin on the Wilson coefficients), and SFitter likelihood construction with the physicist validating each st

Load-bearing premise

The re-casting demonstration assumes the SMEFTatNLO simulations for the new ttZ measurement correctly handle MadGraph's perturbative truncation of dimension-six operators and recompute the top width; the paper's own Appendix B silent-failure test shows this setup is answered correctly in only 0–2 of 10 runs, and no check confirms the actual scan avoided that failure mode.

What would settle it

Re-run the SMEFT scan for the pT(Z) and m(ttZ) observables using the two settings that pass the Appendix B silent-failure test (adjusted truncation and recomputed top width), extract the kappa parameters, and redo the global fit; if the profiled constraints in Figure 5 change materially—in particular the 34–42% tightening or the large negative Cφt shift—the agentic re-casting result as presented is not reproducible.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Adding the normalized pT(Z) ttZ spectrum to the top-sector global fit tightens the profiled constraints on the top-electroweak operators by 34–42%, and no existing bound is loosened.
  • The public likelihood with 276 nuisance parameters can be folded into the fit for the statistics-dominated ttZ measurement with results nearly identical to simpler per-bin uncertainty treatments; the same machinery will matter once systematics-dominated analyses are re-cast.
  • Six independent agent runs on blind coloron datasets reproduce the injected Wilson coefficients, with marginal likelihoods centered on the truth, establishing a repeatable benchmark for agentic re-casting.
  • The documented workflow structure allows an agent to reproduce a previous global analysis exactly, making agent-run fits auditable.
  • The four-step re-casting workflow and the agentic interface generalize beyond SFitter and beyond the top sector, applying to any simulation tool and any global analysis framework.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the workflow scales as claimed, the same pipeline could maintain a continuously updated global SMEFT fit throughout the HL-LHC run, folding in every new differential measurement with a public likelihood as it appears.
  • A natural next step is to instrument the SMEFT scan so that the silent-failure checks from Appendix B (dimension-six truncation handling and top-width recomputation) run automatically before the kappa parameters are extracted; the paper's own validation shows that setup is exactly the case its agents answer correctly only 0–2 times out of 10.
  • The benchmark's observed breakdown of the dimension-six description near the coloron pole suggests an extension: at high invariant mass the workflow could match to UV-complete models directly instead of SMEFT, using the same agentic re-simulation chain.
  • Agent-driven re-casting could also serve as an automated new-physics scanner: bins that pull Wilson coefficients far from the SM, like the flagged 3σ pT(Z) bin, are surfaced to the physicist as candidate signals rather than being averaged away.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper presents MadAgents.v3, a consultant-based agentic layer for MadGraph, and SFitterAgents, an agentic interface to the SFitter global-fitting framework. The central demonstration is a physicist-in-the-loop re-casting exercise: the agents select the ATLAS ttZ measurement (arXiv:2312.04450), re-simulate the SM signal at parton and particle level, run SMEFTatNLO simulations with 21 Wilson coefficients, build κ parameterizations for pT(Z) and m(ttZ), and add these to the global top-sector SMEFT analysis, reproducing previous constraints and tightening them. Validation consists of five 'silent failure' tests (App. B) and a repeatable coloron-injection benchmark with six pseudo-datasets (App. C).

Significance. The paper has real strengths: App. C is a partly independent validation (coloron UV model, Eq. 17), the matching is checked against full coloron samples in Fig. 8, and the reproducible documentation structure plus Table 8 are concrete assets. App. B is unusually candid about hard failures. However, the main re-casting claim is not yet fully supported: the SMEFTatNLO setup used in Sec. 4.3 is exactly the class for which Table 7 shows all agent configurations fail most of the time (SMEFT setup: 0-2/10, warm 0/10), and the paper does not show that the actual scan avoided the truncation and top-width failure modes. Because the κ parametrization feeds the global likelihood and drives Figs. 4-5, this is a load-bearing gap. The central claim is defensible and the gap appears fixable, but requires additional evidence.

major comments (2)
  1. [Sec. 4.3 and App. B, Table 7] Table 7 reports that the SMEFT setup question — correct perturbative truncation of SMEFTatNLO and recomputation of the top width for a dipole-modified decay — is answered correctly 0/10 times by the warm configuration and at most 2/10 by any configuration; the text states 'none of the configurations answers it reliably.' Section 4.3 then builds the full κ parametrization of the new ttZ measurement from SMEFTatNLO runs with 21 Wilson coefficients, and these κ shapes feed the global SFitter likelihood behind Figs. 4–5 and the paper's central claim. The paper does not show that the actual Sec. 4.3 campaign avoided the two failure modes (default tree-level truncation dropping the SM amplitude and dipole operator; inconsistent top width). A reviewer output, a Feynman-diagram check, an independent cross-check of one κ bin, or a statement from the human supervisor is required to establish that
  2. [App. B (validation procedure) and Sec. 5] App. B's validation is self-referential in two ways that matter for the headline claim that MadAgents.v3 prevents silent failures. The grading is performed by an LLM agent (Claude Opus 4.8), and the warm configuration is trained on the same kind of silent-failure lessons on which it is tested. Table 7 gives no information on the grader's false-positive/negative rate, and for the SMEFT row all three configurations score 0–2/10. The paper should (i) report a human re-scoring of at least the SMEFT-question runs and (ii) soften the Sec. 5 statement that the workflow can be expanded 'with no risk concerning the quality of the results', which Table 7 does not support.
minor comments (4)
  1. [Sec. 4.3, p. 13] The selected option 'NLO parton+reuse κ for particle plots' reuses parton-level κ shapes for particle-level plots. Since Sec. 4.4 uses parton-level data for the global fit, this does not affect the main result, but Fig. 3 should state this explicitly.
  2. [Sec. 2.2, Eq. (3)] The correlation matrix sets ρ_ij=0.99 for all systematics pairs. This is a regularized full-correlation approximation, not exact full correlation; a sentence explaining the choice and any sensitivity test would help.
  3. [App. C, Eq. (17)] The injected truth is defined by tree-level matching; the text already says higher-order matching would be more precise. Please add a sentence in the benchmark summary marking that the recovery test validates the tree-level matching value, not a full higher-order SMEFT prediction.
  4. [Sec. 4.3, user prompt] Typo: 'out global analysis' should be 'our global analysis'.

Circularity Check

0 steps flagged

No significant circularity: the SFitter/MadGraph chain and the external coloron closure test give the central claim independent content; the App. B SMEFT silent-failure gap is a robustness risk, not a circular reduction.

full rationale

The derivation chain is: (i) extract ATLAS ttZ data; (ii) re-simulate the SM signal with MadGraph at NLO; (iii) generate SMEFTatNLO scans for the chosen Wilson coefficients and extract per-bin kappa responses; (iv) assemble an SFitter datacard and run the global likelihood. No step defines a Wilson coefficient or kappa in terms of the global-fit output: the kappa shapes come from simulation, and the ATLAS data enter only as comparison data in the likelihood. The claimed improvements are therefore a genuine theory-vs-data update, not a tautology. The strongest independent check is App. C: a coloron model with fixed parameters (Mc=3.75 TeV, tan(theta)=2.1, Gamma_c=1.26 TeV) is integrated out to give the tree-level relation c8/Lambda^2 = -g_c^2/M_c^2 = -0.40/TeV^2 (Eq. 17), and six Poisson-bootstrapped pseudo-datasets are fitted with SFitterAgents, recovering the injected Wilson coefficients and AC=0. This closure test is external to the fitted values and gives the central agentic claim independent content. Self-citations such as MadAgents [20], the previous top-SMEFT analysis [13], and SFitter [36-38] document the tools and baselines being benchmarked; they are not invoked as an unverified uniqueness theorem, and the paper grounds claims in the locally installed MadGraph/SFitter code. App. B does state a serious validation gap: 'The SMEFT question is the sole exception... none of the configurations answers it reliably' (Table 7: warm 0/10), and Sec. 4.3 does not report a check that its SMEFTatNLO scan avoided MadGraph's tree-level truncation and top-width recomputation issue. That is a correctness/robustness risk, not a circularity: the paper never shows that the Sec. 4.3 kappa values were set equal to that failure mode, and the coloron benchmark provides an independent success case. I therefore find no significant circularity.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

No invented entities: the only BSM object, the coloron, is taken from prior literature (Ref. [112]) as a test signal, not proposed as new physics. The free parameters are benchmark design choices and an inherited likelihood regularization, not fit results.

free parameters (2)
  • coloron benchmark parameters (M_c, tan θ, Γ_c) = M_c=3.75 TeV, tan θ=2.1, Γ_c=1.26 TeV
    Chosen by hand in Eq. (15) so the coloron is detectable and the SMEFT approximation holds; these values set the injected-truth Wilson coefficient in Eq. (17), so the benchmark's expected outcome depends on them, but they are inputs, not fit outputs.
  • Correlation-matrix regularization ρ_ij=0.99 = 0.99
    Ad hoc constant in Eq. (3) ensuring invertibility of the systematic-correlation matrix; inherited from the SFitter framework, not fitted here, but a user choice affecting the likelihood.
axioms (5)
  • domain assumption Dimension-6 SMEFT truncation (with linear+quadratic terms) is the correct framework for interpreting LHC top-sector data; dimension-8 operators are not systematically included.
    Sec. 2.1 states dimension-8 operators would exceed current data sensitivity and challenge the likelihood construction; the central re-casting result is phrased entirely in this truncated EFT.
  • domain assumption U(2) flavor symmetry on first two generations and zero light-quark masses (Eq. 12) define the 22-operator top-sector basis.
    Sec. 2.3; the number of operators and their naming, which the agents are built around, follow from this symmetry.
  • domain assumption Gaussian approximation for the top-sector likelihood (Eqs. 9-10) is valid because signals are large and backgrounds negligible.
    Sec. 2.2; the entire top-sector fit in the paper uses this simplified profile likelihood.
  • ad hoc to paper Tree-level coloron-to-SMEFT matching with a single Wilson coefficient for six color-octet operators (Eq. 17) is the correct injected truth for the App. C benchmark.
    The paper states matching 'could be done more precisely at higher-order perturbation theory, but this approximate result is sufficient to check the numerical SFitter result.' The benchmark's passing or failing is judged against this self-defined truth.
  • domain assumption Public likelihood nuisance parameters (276 for the ttZ measurement) can be grouped into SFitter's correlation structure without loss.
    Sec. 4.5; the comparison of uncertainty treatments assumes the profiled-likelihood extraction preserves the relevant correlations.

pith-pipeline@v1.3.0-alltime-deepseek · 27757 in / 17300 out tokens · 166956 ms · 2026-08-01T04:23:57.113034+00:00 · methodology

0 comments
read the original abstract

Analysis re-casting at the LHC is highly standardized and nevertheless requires resources, time, and physics input. Building on the new MadAgents.v3, we show how a global SFitter analysis can be updated by an agentic system with a physicist in the loop. The agentic interface allows us to make the advanced SFitter methodology available to a wider audience. All physical and technical aspects of this agentic re-casting study can be trivially generalized beyond SFitter.

Figures

Figures reproduced from arXiv: 2607.22813 by Daniel Schiller, Nikita Schmal, Sascha Diefenbacher, Tilman Plehn.

Figure 1
Figure 1. Figure 1: SFITTER agents structure. 7 [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of simulation vs data at parton level (left) and particle level [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Effect of a selection of Wilson coefficients chosen to deviate by 3 [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Profiled constraints on Wilson coefficients from the original dataset and [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Profiled constraints before and after adding the [PITH_FULL_IMAGE:figures/full_fig_p015_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Effect of correlations between systematics on the profiled global analysis [PITH_FULL_IMAGE:figures/full_fig_p016_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Difference between implementing a single total uncertainty, estimating [PITH_FULL_IMAGE:figures/full_fig_p017_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Three-way SMEFT-matching validation in m(t¯t), comparing the Standard Model, SM + dimension-six SMEFT at linear order, SM + SMEFT at linear and quadratic order, and the full coloron sample, each with per-curve MC statistical band. 23 [PITH_FULL_IMAGE:figures/full_fig_p023_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Joint profile likelihoods of the six repeated global analyses, one curve per [PITH_FULL_IMAGE:figures/full_fig_p025_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Marginalized likelihoods of the six runs and the eight Wilson coefficients, [PITH_FULL_IMAGE:figures/full_fig_p026_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

300 extracted references · 253 linked inside Pith

  1. [1]

    The Flavor of UV Physics

    Bruggisser, Sebastian and Sch. The Flavor of UV Physics. JHEP. 2021. doi:10.1007/JHEP05(2021)257. arXiv:2101.07273

  2. [2]

    Renormalisation group evolution effects on global SMEFT analyses

    Bartocci, Riccardo and Biek. Renormalisation group evolution effects on global SMEFT analyses. JHEP. 2025. doi:10.1007/JHEP05(2025)203. arXiv:2412.09674

  3. [3]

    A global analysis of the SMEFT under the minimal MFV assumption

    Bartocci, Riccardo and Biek. A global analysis of the SMEFT under the minimal MFV assumption. JHEP. 2024. doi:10.1007/JHEP05(2024)074. arXiv:2311.04963

  4. [4]

    A global analysis of axion-like particle interactions using SMEFT fits

    Biek. A global analysis of axion-like particle interactions using SMEFT fits. JHEP. 2023. doi:10.1007/JHEP09(2023)120. arXiv:2307.10372

  5. [5]

    and Laplace, S

    Hocker, Andreas and Lacker, H. and Laplace, S. and Le Diberder, F. A New approach to a global fit of the CKM matrix. Eur. Phys. J. C. 2001. doi:10.1007/s100520100729. arXiv:hep-ph/0104062

  6. [6]

    and Madigan, Maeve and Mantani, Luca and Moore, James M

    Costantini, Mark N. and Madigan, Maeve and Mantani, Luca and Moore, James M. A critical study of the Monte Carlo replica method. JHEP. 2024. doi:10.1007/JHEP12(2024)064. arXiv:2404.10056

  7. [7]

    and others

    Dittmaier, S. and others. Handbook of LHC Higgs Cross Sections: 2. Differential Distributions. 2012. doi:10.5170/CERN-2012-002. arXiv:1201.3084

  8. [8]

    and others

    Dittmaier, S. and others. Handbook of LHC Higgs Cross Sections: 1. Inclusive Observables. 2011. doi:10.5170/CERN-2011-002. arXiv:1101.0593

  9. [9]

    SMEFT matching to Z' models at dimension eight

    Dawson, Sally and Forslund, Matthew and Schnubel, Marvin. SMEFT matching to Z' models at dimension eight. Phys. Rev. D. 2024. doi:10.1103/PhysRevD.110.015002. arXiv:2404.01375

  10. [10]

    Impact of dimension-eight SMEFT contributions: A case study

    Dawson, Sally and Homiller, Samuel and Sullivan, Matthew. Impact of dimension-eight SMEFT contributions: A case study. Phys. Rev. D. 2021. doi:10.1103/PhysRevD.104.115013. arXiv:2110.06929

  11. [11]

    Corbett, Tyler and Eboli, O. J. P. and Gonzalez-Fraile, J. and Gonzalez-Garcia, M. C. Robust Determination of the Higgs Couplings: Power to the Data. Phys. Rev. D. 2013. doi:10.1103/PhysRevD.87.015022. arXiv:1211.4580

  12. [12]

    Data Preservation in High Energy Physics

    Arbey, Alexandre and others. Data Preservation in High Energy Physics. 2025. arXiv:2503.23619

  13. [13]

    Resolving the flavor structure in the MFV-SMEFT

    Bruggisser, Sebastian and van Dyk, Danny and Westhoff, Susanne. Resolving the flavor structure in the MFV-SMEFT. JHEP. 2023. doi:10.1007/JHEP02(2023)225. arXiv:2212.02532

  14. [14]

    Top and Beauty synergies in SMEFT-fits at present and future colliders

    Bi. Top and Beauty synergies in SMEFT-fits at present and future colliders. JHEP. 2021. doi:10.1007/JHEP06(2021)010. arXiv:2012.10456

  15. [15]

    and Thomas, Marion O

    Celada, Eugenia and Giani, Tommaso and ter Hoeve, Jaco and Mantani, Luca and Rojo, Juan and Rossia, Alejo N. and Thomas, Marion O. A. and Vryonidou, Eleni. Mapping the SMEFT at high-energy colliders: from LEP and the (HL-)LHC to the FCC-ee. JHEP. 2024. doi:10.1007/JHEP09(2024)091. arXiv:2404.12809

  16. [16]

    and Magni, Giacomo and Maltoni, Fabio and Mantani, Luca and Nocera, Emanuele R

    Ethier, Jacob J. and Magni, Giacomo and Maltoni, Fabio and Mantani, Luca and Nocera, Emanuele R. and Rojo, Juan and Slade, Emma and Vryonidou, Eleni and Zhang, Cen. Combined SMEFT interpretation of Higgs, diboson, and top quark data from the LHC. JHEP. 2021. doi:10.1007/JHEP11(2021)089. arXiv:2105.00006

  17. [17]

    Top, Higgs, Diboson and Electroweak Fit to the Standard Model Effective Field Theory

    Ellis, John and Madigan, Maeve and Mimasu, Ken and Sanz, Veronica and You, Tevong. Top, Higgs, Diboson and Electroweak Fit to the Standard Model Effective Field Theory. JHEP. 2021. doi:10.1007/JHEP04(2021)279. arXiv:2012.02779

  18. [18]

    Complete SMEFT predictions for four top quark production at hadron colliders

    Aoude, Rafael and El Faham, Hesham and Maltoni, Fabio and Vryonidou, Eleni. Complete SMEFT predictions for four top quark production at hadron colliders. JHEP. 2022. doi:10.1007/JHEP10(2022)163. arXiv:2208.04962

  19. [19]

    and Maltoni, Fabio and Nocera, Emanuele R

    Hartland, Nathan P. and Maltoni, Fabio and Nocera, Emanuele R. and Rojo, Juan and Slade, Emma and Vryonidou, Eleni and Zhang, Cen. A Monte Carlo global analysis of the Standard Model Effective Field Theory: the top quark sector. JHEP. 2019. doi:10.1007/JHEP04(2019)100. arXiv:1901.05965

  20. [20]

    and Moore, Liam and Russell, Michael and White, Chris D

    Buckley, Andy and Englert, Christoph and Ferrando, James and Miller, David J. and Moore, Liam and Russell, Michael and White, Chris D. Constraining top quark effective theory in the LHC Run II era. JHEP. 2016. doi:10.1007/JHEP04(2016)015. arXiv:1512.03360

  21. [21]

    Almeida, Eduardo da Silva and Alves, Alexandre and \'E boli, Oscar J. P. and Gonzalez-Garcia, M. C. Electroweak legacy of the LHC run II. Phys. Rev. D. 2022. doi:10.1103/PhysRevD.105.013006. arXiv:2108.04828

  22. [22]

    Constraining new physics from Higgs measurements with Lilith: update to LHC Run 2 results

    Kraml, Sabine and Loc, Tran Quang and Nhung, Dao Thi and Ninh, Le Duc. Constraining new physics from Higgs measurements with Lilith: update to LHC Run 2 results. SciPost Phys. 2019. doi:10.21468/SciPostPhys.7.4.052. arXiv:1908.03952

  23. [23]

    and Sanz, Ver \'o nica and You, Tevong

    Ellis, John and Murphy, Christopher W. and Sanz, Ver \'o nica and You, Tevong. Updated Global SMEFT Fit to Higgs, Diboson and Electroweak Data. JHEP. 2018. doi:10.1007/JHEP06(2018)146. arXiv:1803.03252

  24. [24]

    A measurement of the high-mass production cross-section at s =13 TeV with the ATLAS detector and constraints on new particles and couplings

    Aad, Georges and others. A measurement of the high-mass production cross-section at s =13 TeV with the ATLAS detector and constraints on new particles and couplings. JHEP. 2025. doi:10.1007/JHEP10(2025)054. arXiv:2503.19836

  25. [25]

    Interpreting ''Interpretability'' and Explaining ''Explainability'' in Machine Learning in Physics

    Gambhir, Rikab and Lucie-Smith, Luisa and Thaler, Jesse. Interpreting ''Interpretability'' and Explaining ''Explainability'' in Machine Learning in Physics. 2026. arXiv:2606.26228

  26. [26]

    Large Language Model-Assisted Framework for BSM Model Building

    Saad, Shaikh. Large Language Model-Assisted Framework for BSM Model Building. 2026. arXiv:2606.21316

  27. [27]

    LeWRON: Agentic Analysis of Electroweak Phase Transitions

    Wang, Isaac R. LeWRON: Agentic Analysis of Electroweak Phase Transitions. 2026. arXiv:2606.19425

  28. [28]

    and Doglioni, Caterina and G

    Costa, Antonio J. and Doglioni, Caterina and G. AgentRivet: an automated system for producing Rivet routines from journal publications. 2026. arXiv:2606.13535

  29. [29]

    RooAgent: An LLM Agent for Root-Based High Energy Physics Analysis

    Desai, Aman. RooAgent: An LLM Agent for Root-Based High Energy Physics Analysis. 2026. arXiv:2605.17318

  30. [30]

    and Palacios Schweitzer, Sofia and Pang, Ian and Mishra-Sharma, Siddharth and Shih, David

    Faroughy, Darius A. and Palacios Schweitzer, Sofia and Pang, Ian and Mishra-Sharma, Siddharth and Shih, David. Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction. 2026. arXiv:2605.13950

  31. [31]

    and Trifinopoulos, Sokratis

    Niarchos, Vasilis and Papageorgakis, Constantinos and Stapleton, Alexander G. and Trifinopoulos, Sokratis. When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning. 2026. arXiv:2605.06772

  32. [32]

    A Scientific Human-Agent Reproduction Pipeline

    Birk, Joschka and Kasieczka, Gregor and Mishra-Sharma, Siddharth and Nachman, Benjamin and Noll, Dennis and Wamorkar, Tanvi. A Scientific Human-Agent Reproduction Pipeline. 2026. doi:10.5281/zenodo.21078068. arXiv:2604.18752

  33. [33]

    and Gleyzer, Sergei and Matchev, Konstantin T

    Menzo, Tony and Roman, Alexander and Fleming, George T. and Gleyzer, Sergei and Matchev, Konstantin T. and Mrenna, Stephen. Agentic Diagrammatica: Towards Autonomous Symbolic Computation in High Energy Physics. 2026. arXiv:2603.26990

  34. [34]

    The FERMIACC: Agents for Particle Theory

    Agrawal, Prateek and Craig, Nathaniel and Madden, Amalia and Lombera, I \ n igo Valenzuela. The FERMIACC: Agents for Particle Theory. 2026. arXiv:2603.22538

  35. [35]

    QiboAgent: a practitioner's guideline to open source assistants for Quantum Computing code development

    Esposito, Lorenzo and Papaluca, Andrea and Carrazza, Stefano. QiboAgent: a practitioner's guideline to open source assistants for Quantum Computing code development. 2026. arXiv:2603.15538

  36. [36]

    An End-to-end Architecture for Collider Physics and Beyond

    Qiu, Shi and Cai, Zeyu and Wei, Jiashen and Li, Zeyu and Yin, Yixuan and Cao, Qing-Hong and Liu, Chang and Luo, Ming-xing and Yuan, Xing-Bo and Zhu, Hua Xing. An End-to-end Architecture for Collider Physics and Beyond. 2026. arXiv:2603.14553

  37. [37]

    and Matcheva, Katia and Gleyzer, Sergei

    Knipfer, Marco and Roman, Alexander and Matchev, Konstantin T. and Matcheva, Katia and Gleyzer, Sergei. AI Agents for Variational Quantum Circuit Design. 2026. arXiv:2602.19387

  38. [38]

    and Hammad, A

    Esmail, W. and Hammad, A. and Nojiri, M. CoLLM: AI engineering toolbox for end-to-end deep learning in collider analyses. 2026. arXiv:2602.06496

  39. [39]

    Sampling NNLO QCD phase space with normalizing flows

    Jan en, Timo and Poncelet, Rene and Schumann, Steffen. Sampling NNLO QCD phase space with normalizing flows. JHEP. 2025. doi:10.1007/JHEP09(2025)194. arXiv:2505.13608

  40. [40]

    Accelerating multijet-merged event generation with neural network matrix element surrogates

    Herrmann, Tim and Jan en, Timo and Schenker, Mathis and Schumann, Steffen and Siegert, Frank. Accelerating multijet-merged event generation with neural network matrix element surrogates. 2025. arXiv:2506.06203

  41. [41]

    Integrating particle flavor into deep learning models for hadronization

    Chan, Jay and Ju, Xiangyang and Kania, Adam and Nachman, Benjamin and Sangli, Vishnu and Siodmok, Andrzej. Integrating particle flavor into deep learning models for hadronization. Phys. Rev. D. 2025. doi:10.1103/hgbg-k7js. arXiv:2312.08453

  42. [42]

    Resummation of the C-Parameter Sudakov Shoulder Using Effective Field Theory

    Schwartz, Matthew D. Resummation of the C-Parameter Sudakov Shoulder Using Effective Field Theory. 2026. arXiv:2601.02484

  43. [43]

    Herwig 7.3 release note

    Bewick, Gavin and others. Herwig 7.3 release note. Eur. Phys. J. C. 2024. doi:10.1140/epjc/s10052-024-13211-9. arXiv:2312.05175

  44. [44]

    Event Generation with Sherpa 2.2

    Bothmann, Enrico and others. Event Generation with Sherpa 2.2. SciPost Phys. 2019. doi:10.21468/SciPostPhys.7.3.034. arXiv:1905.09127

  45. [45]

    and Frederix, R

    Alwall, J. and Frederix, R. and Frixione, S. and Hirschi, V. and Maltoni, F. and Mattelaer, O. and Shao, H. -S. and Stelzer, T. and Torrielli, P. and Zaro, M. The automated computation of tree-level and next-to-leading order differential cross sections, and their matching to parton shower simulations. JHEP. 2014. doi:10.1007/JHEP07(2014)079. arXiv:1405.0301

  46. [46]

    An introduction to PYTHIA 8.2

    Sj. An introduction to PYTHIA 8.2. Comput. Phys. Commun. 2015. doi:10.1016/j.cpc.2015.01.024. arXiv:1410.3012

  47. [47]

    MadGraph 5 : Going Beyond

    Alwall, Johan and Herquet, Michel and Maltoni, Fabio and Mattelaer, Olivier and Stelzer, Tim. MadGraph 5 : Going Beyond. JHEP. 2011. doi:10.1007/JHEP06(2011)128. arXiv:1106.0522

  48. [48]

    MadEvent: Automatic event generation with MadGraph

    Maltoni, Fabio and Stelzer, Tim. MadEvent: Automatic event generation with MadGraph. JHEP. 2003. doi:10.1088/1126-6708/2003/02/027. arXiv:hep-ph/0208156

  49. [49]

    2024 , eprint=

    AstroMLab 1: Who Wins Astronomy Jeopardy!? , author=. 2024 , eprint=

  50. [50]

    AstroLLaMA-Chat: Scaling AstroLLaMA with Conversational and Diverse Datasets

    Perkowski, Ernest and others. AstroLLaMA-Chat: Scaling AstroLLaMA with Conversational and Diverse Datasets. Res. Notes AAS. 2024. doi:10.3847/2515-5172/ad1abe. arXiv:2401.01916

  51. [51]

    2021 , eprint=

    Building astroBERT, a language model for Astronomy & Astrophysics , author=. 2021 , eprint=

  52. [52]

    Menzo, Tony and Roman, Alexander and Gleyzer, Sergei and Matchev, Konstantin and Fleming, George T. and H. HEPTAPOD: Orchestrating High Energy Physics Workflows Towards Autonomous Agency. 2025. arXiv:2512.15867

  53. [53]

    The AI Cosmologist I: An Agentic System for Automated Data Analysis

    Moss, Adam. The AI Cosmologist I: An Agentic System for Automated Data Analysis. 2025. arXiv:2504.03424

  54. [54]

    2024 , eprint=

    SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning , author=. 2024 , eprint=

  55. [55]

    2025 , eprint=

    Interpreting Multi-band Galaxy Observations with Large Language Model-Based Agents , author=. 2025 , eprint=

  56. [56]

    cosmosage: A natural-language assistant for cosmology

    de Haan, Tijmen. cosmosage: A natural-language assistant for cosmology. Astron. Comput. 2025. doi:10.1016/j.ascom.2025.100934. arXiv:2407.04420

  57. [57]

    AstroLLaMA: Towards Specialized Foundation Models in Astronomy

    Nguyen, Tuan Dung and others. AstroLLaMA: Towards Specialized Foundation Models in Astronomy. 2023. arXiv:2309.06126

  58. [58]

    2024 , eprint=

    The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery , author=. 2024 , eprint=

  59. [59]

    Multi-Agent System for Cosmological Parameter Analysis

    Laverick, Andrew and Surrao, Kristen and Zubeldia, Inigo and Bolliet, Boris and Cranmer, Miles and Lewis, Antony and Sherwin, Blake and Lesgourgues, Julien. Multi-Agent System for Cosmological Parameter Analysis. 2024. arXiv:2412.00431

  60. [60]

    Zhang, Xiaowen and Bi, Zhenyu and Lachance, Patrick and Wang, Xuan and Di Matteo, Tiziana and Croft, Rupert A. C. Bridging Literature and the Universe Via A Multi-Agent Large Language Model System. 2025. arXiv:2507.08958

  61. [61]

    FeynTune: large language models for high-energy theory

    Richmond, Paul and Papageorgakis, Constantinos and Niarchos, Vasilis and Chowdhury, Borun and Agarwal, Prarit. FeynTune: large language models for high-energy theory. Mach. Learn. Sci. Tech. 2026. doi:10.1088/2632-2153/ae47bb. arXiv:2508.03716

  62. [62]

    Bakshi, S. D. and others. ArgoLOOM: agentic AI for fundamental physics from quarks to cosmos. 2025. arXiv:2510.02426

  63. [63]

    Automating High Energy Physics Data Analysis with LLM-Powered Agents

    Gendreau-Distler, Eli and Ho, Joshua and Kim, Dongwon and Le Pottier, Luc Tomas and Wang, Haichen and Yang, Chengxi. Automating High Energy Physics Data Analysis with LLM-Powered Agents. 39th Annual Conference on Neural Information Processing Systems : Includes Machine Learning and the Physical Sciences (ML4PS). 2025. arXiv:2512.07785

  64. [64]

    Agents of Discovery

    Diefenbacher, Sascha and Hallin, Anna and Kasieczka, Gregor and Kr. Agents of Discovery. 2025. arXiv:2509.08535

  65. [65]

    arXiv:2303.18223

    A survey of large language models , author=. arXiv:2303.18223

  66. [66]

    National Science Review , volume=

    A survey on multimodal large language models , author=. National Science Review , volume=. 2024 , publisher=. arXiv:2306.13549

  67. [67]

    arXiv:2305.13971

    Grammar-constrained decoding for structured NLP tasks without finetuning , author=. arXiv:2305.13971

  68. [68]

    arXiv:2205.12255

    Talm: Tool augmented language models , author=. arXiv:2205.12255

  69. [69]

    Advances in Neural Information Processing Systems , volume=

    Toolformer: Language models can teach themselves to use tools , author=. Advances in Neural Information Processing Systems , volume=. arXiv:2302.04761

  70. [70]

    The eleventh international conference on learning representations , year=

    React: Synergizing reasoning and acting in language models , author=. The eleventh international conference on learning representations , year=. arXiv:2210.03629

  71. [71]

    arXiv:2304.05128

    Teaching large language models to self-debug , author=. arXiv:2304.05128

  72. [72]

    arXiv:2305.16291

    Voyager: An open-ended embodied agent with large language models , author=. arXiv:2305.16291

  73. [73]

    arXiv:2305.17126

    Large language models as tool makers , author=. arXiv:2305.17126

  74. [74]

    Advances in Neural Information Processing Systems , volume=

    Swe-agent: Agent-computer interfaces enable automated software engineering , author=. Advances in Neural Information Processing Systems , volume=. arXiv:2405.15793

  75. [75]

    Advances in neural information processing systems , volume=

    Chain-of-thought prompting elicits reasoning in large language models , author=. Advances in neural information processing systems , volume=. arXiv:2201.11903

  76. [76]

    Forty-first International Conference on Machine Learning , year=

    Improving factuality and reasoning in language models through multiagent debate , author=. Forty-first International Conference on Machine Learning , year=. arXiv:2305.14325

  77. [77]

    arXiv:2310.02170

    Dynamic llm-agent network: An llm-agent collaboration framework with agent team optimization , author=. arXiv:2310.02170

  78. [78]

    The Twelfth International Conference on Learning Representations , year=

    MetaGPT: Meta programming for a multi-agent collaborative framework , author=. The Twelfth International Conference on Learning Representations , year=. arXiv:2308.00352

  79. [79]

    Frontiers of Computer Science , volume=

    A survey on large language model based autonomous agents , author=. Frontiers of Computer Science , volume=. 2024 , publisher=. arXiv:2308.11432

  80. [80]

    arXiv:2412.17481

    A survey on llm-based multi-agent system: Recent advances and new frontiers in application , author=. arXiv:2412.17481

Showing first 80 references.