Pith. sign in

REVIEW 4 major objections 5 minor 93 references

An autonomous LLM agent can run full relativistic hydrodynamic studies end to end, and its first findings trace viscous flow suppression to high temperatures and separate oxygen nuclear models by flow–size correlation.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 00:45 UTC pith:YXUWZOZD

load-bearing objection A useful proof-of-principle for agent-driven hydrodynamics: the scan-compose-validate skill is the real contribution, while the O+O physics separation is too fragile to carry weight alone. the 4 major comments →

arxiv 2607.27822 v1 pith:YXUWZOZD submitted 2026-07-30 nucl-th hep-ph

CLVisc Agent for autonomous relativistic hydrodynamics studies

classification nucl-th hep-ph
keywords autonomous LLM agentrelativistic hydrodynamicsCLViscshear viscosity temperature dependencequark-gluon plasmaoxygen-16 nuclear structureelliptic flow correlationsheavy-ion collisions
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper claims that a single large-language-model agent, given a meta skill that scans a codebase, composes an operational skill, and validates it by real runs, can autonomously carry out end-to-end hydrodynamic studies of the quark-gluon plasma: it designs parameter scans, edits configuration files, launches GPU-accelerated simulations, extracts observables, and interprets the results without per-event human steering. The method is exercised on the CLVisc code in two scenarios. In Pb+Pb collisions, the agent's paired-event scan shows that the high-temperature branch of the shear-viscosity-to-entropy ratio η/s controls most of the viscous suppression of anisotropic flow, while changing η/s only below the crossover temperature has little effect. In O+O collisions, the same pipeline finds that four ab initio oxygen-16 ground-state descriptions separate into three distinguishable groups, with the unclustered configuration singled out by an unusually weak correlation between elliptic flow and mean transverse momentum. A sympathetic reader would care because this turns labor-intensive simulation workflows into repeatable agent-driven scans and produces qualitative parameter–observable knowledge of a kind Bayesian fits tend to bury.

Core claim

The central discovery is that an LLM agent can graduate from code exploration to autonomous research operation: using the scan–compose–validate meta-skill, it builds a CLVisc-specific skill containing launch commands, parameter semantics, observable extractors, and expected physical tendencies, then executes complete scientific workflows. The agent's physics conclusions are stated as follows. First, in Pb+Pb 30–40% centrality at 5.02 TeV, the high-temperature branch of a piecewise-linear η/s(T) profile governs the viscous suppression of yields, mean transverse momentum, and elliptic/triangular flow, while the sub-Tc branch is nearly inert—the low-temperature-slope profile is statistically in

What carries the argument

The central device is the scan–compose–validate meta-skill, a generic engine that reads a code checkout, writes a version-specific operational skill (a skill document plus scripts and references), and validates it by actually running the code before production use. The scientific load-bearing tools are the paired-event comparison—evolving the same 200 initial conditions under every viscosity profile so final-state differences isolate viscosity effects—and a set of observable constructs: the hydro response efficiency κn = vn/εn, which separates viscous damping from initial geometry; the flow–size correlation ρ(v2²,[pT]); the compactness diagnostic darea; and the intrinsic quadrupole amplitude

Load-bearing premise

The O+O comparison assumes the four ab initio 16O configurations are faithfully sampled and that, with the hydrodynamic medium held fixed, final-state differences trace to nuclear structure rather than to the tuned K-factor, centrality window, or the missing hadronic afterburner—if those reshape the model ordering, the claimed separability collapses.

What would settle it

Evolve the same 200 TRENTo initial conditions through a hadronic afterburner and recompute v2, v3, and ρ(v2²,[pT]): if the low-temperature-slope profile moves away from the constant η/s = 0.08 curve by more than about one percent, or if PGCM-uniform loses its outlier status in ρ(v2²,[pT]), the paper's central claims are falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the high-temperature-branch dominance holds, future constraints on η/s(T) can concentrate scan effort on the early high-temperature stage; the low-temperature branch can be fixed with little cost to flow observables in this centrality class.
  • Because v3 is roughly twice as viscous-sensitive as v2, triangular flow becomes a recommended differential observable for viscosity-profile extraction.
  • O+O collisions can act as a selective probe of 16O ground-state geometry, separating at least three of four ab initio models at fixed multiplicity.
  • The meta-skill's version-agnostic, knowledge-pack design means the same agent pipeline can be extended to other simulation codes by swapping a thin per-model knowledge pack rather than rewriting the engine.
  • The standardized high-dimensional datasets produced by agent-run scans are well suited to machine-learning or LLM-guided searches for new discriminants, such as symmetry-plane correlations.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's own limitation statement implies the model orderings are untested against hadronic afterburner effects; a natural next experiment is to rerun the same four 16O ensembles with an afterburner and check whether PGCM-uniform remains the ρ(v2²,[pT]) outlier.
  • If the size–shape decoupling diagnostic generalizes, ρ(ε2²,darea) could become a standard initial-state fingerprint distinguishing clustered from unclustered ab initio structure across other small systems such as Ne+Ne or Ar+Sc, not just O+O.
  • The claim that low-temperature η/s is nearly invisible is made in 30–40% centrality with bulk viscosity switched off; a cautious extrapolation is that this hierarchy may shift at other centralities or when bulk viscosity is enabled, which the same agent pipeline could map directly.
  • The paper's stated limitation that physical interpretations still require expert verification suggests the right reading is 'human-guided autonomous execution' rather than fully unsupervised discovery, since the agent's physics expectations come from the knowledge pack it is given.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents an LLM-agent framework for autonomous relativistic hydrodynamics studies. A meta-skill, project-explorer-and-skill-creator, scans a CLVisc checkout, composes a version-specific SKILL.md with scripts and references, and validates it by real execution. The agent is then asked to run two physics scenarios. In Scenario I (Pb+Pb at 5.02 TeV, 30–40%), five temperature-dependent η/s parametrizations are compared; the agent uses event-averaged observables and a paired-event analysis to conclude that the high-temperature branch of η/s controls most of the viscous suppression of flow, while modifying η/s below Tc has little effect. In Scenario II (O+O at 5.36 TeV, 0–5%), four ab initio 16O structure inputs (NLEFT, PGCM-clustered, PGCM-uniform, VMC) are evolved with fixed hydrodynamics; the agent argues that PGCM-uniform is singled out by an anomalously low ρ(v2^2,[pT]) caused by an initial-state decoupling of ellipticity and transverse compactness, while VMC and PGCM-clustered remain degenerate. The paper explicitly labels the orderings qualitative and lists limitations including the absence of a hadronic afterburner and the fixed medium in Scenario II.

Significance. If the results hold, the paper would demonstrate a useful methodological advance: an LLM agent can automate the full hydrodynamics workflow—parameter-scan design, job execution, event-by-event analysis, and interpretation—and can propose new initial-state diagnostics. The scan–compose–validate architecture is modular and the paired-event analysis in Scenario I is well designed. The physics claims, however, are exploratory. Scenario I is reasonably supported by the paired-event comparison and is consistent with standard viscous-damping expectations. Scenario II is the load-bearing weakness: the claimed three-way separation of NLEFT, PGCM-uniform, and VMC/PGCM-clustered rests on a small event sample, a single correlation observable, and unreported centrality/normalization choices. The paper itself states that the O+O orderings are qualitative until tested against variations of transport coefficients, centrality definitions, and analysis cuts, and against fuller statistical uncertainties.

major comments (4)
  1. [§II.D, Fig. 6, Table II] The O+O discrimination is not yet robust because the centrality selection and normalization are not reported. The text states that the agent retains only events falling within the desired centrality window and scales each TRENTo entropy profile with a 'tuned K-factor' to reach dNch/dη≈130, but neither the entropy cut, the retained event fraction, nor the K-factor value is given. Since ρ(v2^2,[pT]) is sensitive to centrality and mean multiplicity, a per-model K-factor or a slightly different entropy window could move PGCM-uniform’s value relative to the other models. Please report these numbers and demonstrate that the PGCM-uniform ordering survives variations in the centrality cut and K-factor (or rescaling after centrality selection). The paper’s own limitation section acknowledges this, but the central Scenario II conclusion currently rests on it.
  2. [Table II] The key discriminator ρ(ε2^2,darea) is presented without uncertainties. The central claim is that PGCM-uniform is decoupled (−0.006) while the other models have +0.08 to +0.12; with ~1000 events and the jackknife procedure described in §II.D, these numbers must carry error bars. If the jackknife uncertainty is comparable to the spread, the 'decoupling' claim is unsupported. Please add uncertainties to all entries in Table II and to the v2{2} and ρ(v2^2,[pT]) values in Fig. 6.
  3. [§III.B, Eq. (10)] The causal chain connecting darea to the final-state correlation is asserted rather than demonstrated. darea is introduced after the final-state pattern is seen, and the paper states that 'the hydrodynamic evolution carries ε2 into v2 and transverse compactness into [pT]' without an event-by-event validation. Please show that ρ(ε2^2,darea) and ρ(v2^2,[pT]) are related event-by-event, or provide a direct demonstration that the initial-state size–shape decoupling propagates to the final-state correlation, before claiming that the missing size–shape coupling accounts for the PGCM-uniform suppression.
  4. [§II.A] The methodological claim of autonomous skill creation is difficult to evaluate because the knowledge pack, SKILL.md, and agent logs are not included or referenced. The agent’s choices are steered by 'default expectations' encoded in the knowledge pack, and the paper’s evidence for autonomy depends on this unpublished material. Please provide the skill files or a repository as supplementary material, and report the minimal information needed to reproduce the agent’s planning decisions. This is load-bearing for the paper’s central methodological claim, even though it does not affect the hydrodynamic results themselves.
minor comments (5)
  1. [§III.A.3] The sentence following Eq. (8) is broken: 'indicating that the initial eccentricity provides only the geometric seed, while the for κ3 than for κ2'. Please rewrite.
  2. [§II.D / §III.B] The text uses 'PGCM-c' and 'PGCM-clustered' interchangeably, and 'PGCM-u' for PGCM-uniform. Please define and use a single abbreviation set consistently.
  3. [§I] Typo: 'we present a end-to-end framework' should be 'an end-to-end framework'. Also, 'theagent' appears without a space in §II.D.
  4. [Table I] Table I lists ⟨pT⟩ to three decimals without statistical uncertainties, while the text interprets differences of 1.7–2.2%. Please add uncertainties or explicitly state that the ordering is qualitative; the paired-event analysis in Fig. 5 is more informative.
  5. [Fig. 4] The legend labels 'lowT' and 'highT' are ambiguous. Expand to 'low-T slope' and 'high-T slope' to match the parametrization names used elsewhere.

Circularity Check

2 steps flagged

No hard by-construction circularity: the eta/s result is a genuine controlled-decomposition simulation outcome and the PGCM-u result is a measured forward-model correlation. The moderate concerns are interpretive: darea is a post-hoc diagnostic, the agent's unpublished 'default expectations' seed the qualitative conclusions, and a minor self-citation [83] supports a non-load-bearing premise.

specific steps
  1. other [Sec. III.B (Scenario II), after Eq. (10), darea compactness diagnostic]
    "The agent has defined a new initial state "observable" to explain why PGCM-uniform, despite having an elliptic flow comparable to others, gives such a suppressed final-state ρ(v2 2, [pT ]). This traces to the near-absence of an initial-state correlation between geometric eccentricity and transverse compactness. ... Because the hydrodynamic evolution carries ε2 into v2 and transverse compactness into [pT ], this initial-state decoupling propagates to the final state and accounts for the suppressed ρ(v2 2, [pT ])."

    The explanatory observable darea is defined only after the PGCM-u anomaly in ρ(v2^2,[pT]) is observed, so ρ(ε2^2,darea) = -0.006 is selected to match the outcome. The causal link compactness→[pT] is asserted ('a more compact initial state drives stronger radial flow'), not demonstrated event-by-event, and both correlations are computed on the same events whose centrality window and tuned K-factor are unreported. The explanation is constructed on the data it explains and its quantitative propagation is never shown, so the causal claim is outcome-matched rather than independently derived; the paper itself concedes the orderings are qualitative until centrality and analysis choices are varied. This weakens the claim's independence without equating result and input by construction.

  2. other [Sec. II.A (Agentic workflow architecture), analytical layer of the CLVisc knowledge pack]
    "At the analytical level, it supplies observable extractors and default expectations for how variations of transport parameters should influence selected observables. Comparing extracted results against these expectations enables the agent to test hypotheses and identify potentially informative deviations."

    The knowledge pack steering the agent explicitly contains 'default expectations for how variations of transport parameters should influence selected observables,' and the pack is unpublished. The Scenario I conclusion — 'the high-temperature branch of η/s controls most of the viscous suppression of anisotropic flow' — is of exactly that kind, so the agent's qualitative 'discovery' is seeded by its input. The quantitative support (paired-event Fig. 5) is genuine CLVisc output, so the physics is not fabricated; but because the seeded expectations coincide with the reported conclusions and cannot be inspected, the independence of the interpretive claim is not verifiable. This is a partial input-output alignment in the discovery framing, not an equation-level reduction.

full rationale

I walked both claimed derivation chains against the manuscript's own equations and citations. Scenario I is a clean controlled branch decomposition: the five parametrizations are built so that (const 0.08, low-T slope) share the same above-Tc branch and (V-shape, high-T slope) share the same above-Tc branch; the finding that these pairs remain nearly identical in final observables while const 0.16 groups with the high-T-branch curves is a simulation outcome, not an input. No quantity is fitted to the conclusion. Scenario II's PGCM-u singling-out is likewise a measured forward-model correlation (Fig. 6d), with uncertainties from jackknife; the initial-state correlation ρ(ε2^2,darea) is computed honestly on TRENTo events. The genuine circularity-adjacent weaknesses are interpretive: (i) darea is explicitly introduced post-hoc to explain the anomaly and its propagation chain is asserted rather than derived; (ii) the agent's interpretive layer is seeded by an unpublished knowledge pack containing 'default expectations' that coincide with the paper's qualitative physics conclusions. Both are flagged here, but neither is an equation-level equivalence (no Eq. X = Eq. Y by construction) and no fitted parameter is relabeled as a prediction. On self-citation: Ref. [83] (Q. Wang, L.-G. Pang, X.-N. Wang — three of the four present authors) is cited for the short-range-correlation mechanism behind VMC's small eccentricity, but it is not load-bearing: VMC's small ⟨ε2⟩ and ⟨β2⟩ are measured directly in Table II, external Ref. [82] carries the primary attribution, and the Conclusions explicitly leave the microscopic origin open ('remains open'). The manuscript's own Limitations passage ('should be regarded as qualitative until they are tested against variations of the transport coefficients, centrality definitions, and analysis cuts') and its methodological caveat ('the agent operates within the domain knowledge encoded in the SKILL framework, and its physical interpretations still require expert verification') disclose the main robustness and transparency limits; those are correctness risks rather than circular reductions. The central physics results therefore retain independent computational content, but the post-hoc diagnostic and the seeded expectations warrant a moderate score of 4 rather than 0-2.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 1 invented entities

The central claims rest on the tuned K-factor, hand-chosen scan profiles, TRENTo and ab initio inputs, and the authors-supplied knowledge pack. None of these are derived in the paper; they are imported from prior work or chosen by the authors. The only invented physical entity is the post-hoc darea diagnostic.

free parameters (4)
  • K-factor (O+O entropy scaling) = not stated
    Section II.D: 'scaling the entropy density with a tuned K-factor' to reach dNch/dη≈130; the value is not given, and it normalizes all O+O comparisons.
  • TRENTo parameters w and k = w=0.5 fm, k=1.0 (Scenario I)
    Chosen by hand in Section II.C; standard values but central to initial-condition geometry and all derived flow observables.
  • η/s profile parameters = η_min=0.08 or 0.16; Tmin=Tc=0.15 GeV; slopes unspecified
    Five profiles in Fig. 2 and Section II.C are chosen by hand as scan inputs; they determine the physical conclusions about which temperature branch of η/s matters.
  • Hydrodynamic start and freeze-out settings = τ0=0.6 fm, Tfrz=0.137 GeV
    Conventional fixed settings used in both scenarios; they influence final yields and flow but are not fitted here.
axioms (5)
  • domain assumption Israel-Stewart causal viscous hydrodynamics is an adequate description of QGP expansion.
    Eqs. (1)-(3), Section II.B; the entire simulation program assumes the quark-gluon plasma can be described by viscous hydrodynamics with shear/bulk relaxation.
  • domain assumption TRENTo initial conditions with the specified parameters represent the initial state.
    Section II.C/D; TRENTo is a parametric model, not derived from QCD, and the nucleon width, k parameter, and longitudinal profile are inputs.
  • domain assumption The four ab initio 16O configurations (NLEFT, PGCM clustered/uniform, VMC) are faithful ground-state samples.
    Section II.D; relies on Refs. [77,78]. If the configurations are inaccurate or sampled incorrectly, the O+O conclusions do not transfer to real nuclei.
  • domain assumption With hydro parameters fixed, differences among O+O models are attributable to nuclear-structure input rather than event-selection or normalization differences.
    Section II.D; the controlled-comparison logic requires centrality selection and K-factor scaling not to bias one model over another.
  • ad hoc to paper The agent's knowledge pack encodes correct operational and physical knowledge, so its autonomous choices reflect skill rather than chance or implicit prompting.
    Section II.A supplies 'the physical interpretation of key parameters' and 'default expectations' for observable responses; the autonomy claim depends on this unvalidated, unpublished pack.
invented entities (1)
  • darea (transverse compactness diagnostic) no independent evidence
    purpose: Quantify initial-state size-shape correlation to explain why PGCM-uniform has a suppressed final-state ρ(v2²,[pT]).
    Defined in Eq. (10) after observing the PGCM-u anomaly; no independent or out-of-sample falsifiable handle is provided. It is an exploratory post-hoc observable.

pith-pipeline@v1.3.0-daily-deepseek · 18849 in / 15069 out tokens · 150237 ms · 2026-08-01T00:45:54.093994+00:00 · methodology

0 comments
read the original abstract

We enable large language model (LLM) agents to autonomously perform end-to-end hydrodynamic simulations of the quark-gluon plasma evolution and calculation of final hadron spectra in relativistic heavy-ion collisions. We design a meta skill that allows an agent to explore a project's source code, craft a specialized skill, and iteratively refine it. Applying this meta skill to the (3+1)D viscous hydrodynamic code CLVisc, the agent builds a CLVisc skill encoding its operational knowledge and then independently executes full scientific workflows: designing parameter scans, running simulations, comparing ensemble results, and producing publication-ready figures. Crucially, the agent draws on literature-informed heavy-ion physics to select physically meaningful observables and interpret outcomes without explicit instruction. We demonstrate the pipeline in two scenarios: temperature-dependent shear viscosity over entropy density $\eta/s$, and nuclear-structure effects in O+O collisions at $\sqrt{s_{\mathrm{NN}}} = 5.36$~TeV using four \textit{ab initio} descriptions of $^{16}$O. In both, the agent plans, executes, and analyzes autonomously, devising new initial-state observables to explain final observations and extract qualitative knowledge. The meta skill is agnostic to code versions and Monte Carlo generators, promising future multi-agent systems in high-energy nuclear physics.

Figures

Figures reproduced from arXiv: 2607.27822 by Long-Gang Pang, Qi Wang, Shi Pu, Xin-Nian Wang.

Figure 1
Figure 1. Figure 1: FIG. 1. The scan–compose–validate workflow of the [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. Temperature-dependent shear viscosity [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Rapidity distributions [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4. Flow harmonics [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: summarizes the paired-event relative responses for six quantities: the charged-particle yield Nch, the inte￾grated elliptic and triangular flows v2 and v3, the hydro￾dynamic response efficiencies κ2 = v2/ε2 and κ3 = v3/ε3, and the hydrodynamic duration τdur. The agent’s paired￾event comparison revealed a consistent and physically in￾terpretable hierarchy. For the constant η/s = 0.16 profile, the median shi… view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6. Final-state observables in 0–5% O+O collisions at [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

93 extracted references · 26 linked inside Pith

  1. [1]

    The agentic pipeline began its analysis of Scenario I by organizing the event-averaged final-state observables into a coherent physical narrative

    Event-averaged observables. The agentic pipeline began its analysis of Scenario I by organizing the event-averaged final-state observables into a coherent physical narrative. Informed by its knowl- edge of standard heavy-ion observables, the agent first extracted rapidity distributionsdN/dy for π+, K +, and ¯pacross all five viscosity parametrizations, wi...

  2. [2]

    Event-by-event viscosity response. To isolate the viscosity-induced modification of the hydrodynamic response from initial-state geometry effects, the agent independently devised a paired-event analysis: the same ensemble of 200 TRENTo initial conditions was evolved under all fiveη/s profiles, so that differences in the final-state observables reflect cha...

  3. [3]

    observable

    Dynamical evolution and intermediate hydrodynamic features. To understand why the response efficienciesκn = vn/εn vary from event to event,the agent extended the analysis from final-state observables to intermediate hydrodynamic diagnosticsstoredintheCLVisc bulkinfo.h5output. We focused on two compact quantities: the hydrodynamic duration τdur, defined as...

  4. [4]

    L.-G. Pang, K. Zhou, N. Su, H. Stoecker, H. Petersen, and X.-N. Wang, Classify qcd phase transition with deep learning, Nuclear Physics A982, 867 (2019)

  5. [5]

    C. Gao, S. Höche, J. Isaacson, C. Krause, and H. Schulz, Event generation with normalizing flows, Physical Review D101, 076002 (2020)

  6. [6]

    W.-B.He, Y.-G.Ma, L.-G.Pang, H.-C.Song,andK.Zhou, High-energy nuclear physics meets machine learning, Nu- clear Science and Techniques34, 88 (2023)

  7. [7]

    Boehnlein, M

    A. Boehnlein, M. Diefenthaler, N. Sato, M. Schram, V. Ziegler,et al., Colloquium: Machine learning in nuclear physics, Reviews of Modern Physics94, 031003 (2022)

  8. [8]

    L.-G. Pang, K. Zhou, N. Su, H. Petersen, H. Stoecker, and X.-N. Wang, An equation-of-state-meter of quantum chromodynamics transition from deep learning, Nature Communications9, 210 (2018)

  9. [9]

    Dai, F.-P

    S.-W. Dai, F.-P. Li, L.-G. Pang, G.-Y. Qin, S.-Y. Wei, H.-Z. Zhang, and W. Zhao, Physics-Informed Global Ex- tractionoftheUniversalSmall- xDipoleAmplitude(2026), arXiv:2603.08008 [hep-ph]. 13

  10. [10]

    in Sec. III. III. RESULTS A. Scenario I: Effects of temperature-dependent shear viscosity over entropy density

  11. [11]

    Mikuni, B

    V. Mikuni, B. Nachman, and M. Pettee, Fast point cloud generation with diffusion models in high energy physics, Physical Review D108, 036025 (2023)

  12. [12]

    Raissi, P

    M. Raissi, P. Perdikaris, and G. E. Karniadakis, Physics- informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations, Journal of Computational Physics378, 686 (2019)

  13. [13]

    Dai, F.-P

    S.-W. Dai, F.-P. Li, L.-G. Pang, X.-N. Wang, B.-W. Zhang, and H.-Z. Zhang, Parton Fragmentation Func- tions Extracted with a Physics-Informed Neural Network (2026), arXiv:2601.20177 [hep-ph]

  14. [14]

    S. Ren, C. Xie, P. Jian, Z. Ren, C. Leng, and J. Zhang, Towards scientific intelligence: A survey of LLM-based scientific agents (2025), arXiv:2503.24047 [cs.AI]

  15. [15]

    Edelen and X

    A. Edelen and X. Huang, Machine learning for design and control of particle accelerators: A look backward and forward, Annual Review of Nuclear and Particle Science 74, 557 (2024)

  16. [16]

    Kvapil, G

    J. Kvapil, G. Borca-Tasciuc, H. Bossi, K. Chen, Y. Chen, et al., Intelligent experiments through real-time AI: Fast data processing and autonomous detector con- trol for sPHENIX and future EIC detectors (2025), arXiv:2501.04845 [physics.ins-det]

  17. [17]

    S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, ReAct: Synergizing reasoning and acting in language models, inProc. ICLR 2023(2023)

  18. [18]

    Schick, J

    T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, Toolformer: Language models can teach themselves to use tools, Advances in Neural Information Processing Systems 36(2023)

  19. [19]

    D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes, Au- tonomous chemical research with large language models, Nature624, 570 (2023)

  20. [20]

    A. M. Bran, S. Cox, O. Schilter, C. Baldassari, A. D. White, and P. Schwaller, ChemCrow: Augmenting large-language models with chemistry tools (2023), arXiv:2304.05376 [physics.chem-ph]

  21. [21]

    C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha, The AI scientist: Towards fully automated open- ended scientific discovery (2024), arXiv:2408.06292 [cs.AI]

  22. [22]

    Yamadaet al., The AI scientist-v2: Workshop-level automated scientific discovery via agentic tree search (2025), arXiv:2504.08066 [cs.AI]

    Y. Yamadaet al., The AI scientist-v2: Workshop-level automated scientific discovery via agentic tree search (2025), arXiv:2504.08066 [cs.AI]

  23. [23]

    Zheng, Z

    T. Zheng, Z. Deng, H. T. Tsang, W. Wang, J. Bai, Z. Wang, and Y. Song, From automation to autonomy: A survey on large language models in scientific discovery, inProc. EMNLP 2025, edited by C. Christodoulopou- los, T. Chakraborty, C. Rose, and V. Peng (Association for Computational Linguistics, Suzhou, China, 2025) pp. 17733–17750

  24. [24]

    S. Qiu, S. Guo, Z.-Y. Song, Y. Sun, Z. Cai,et al., PHY- Bench: Holistic evaluation of physical perception and rea- soning in large language models (2025), arXiv:2504.16074 [cs.CL]

  25. [25]

    Novikov, N

    A. Novikov, N. V˜ u, M. Eisenberger, E. Dupont,et al., AlphaEvolve: A coding agent for scientific and algorithmic discovery (2025), arXiv:2506.13131 [cs.AI]

  26. [26]

    Schmidgall, Y

    S. Schmidgall, Y. Su, Z. Wang, X. Sun, J. Wu, X. Yu, J. Liu, M. Moor, Z. Liu, and E. Barsoum, Agent labora- tory: Using LLM agents as research assistants, Findings of the Association for Computational Linguistics: EMNLP 2025 , 5977 (2025)

  27. [27]

    Z. Chen, S. Chen, Y. Ning, Q. Zhang, B. Wang,et al., ScienceAgentBench: Toward rigorous assessment of lan- guage agents for data-driven scientific discovery, inProc. ICLR 2025(2025)

  28. [28]

    M. Tian, L. Gao, S. D. Zhang, X. Chen, C. Fan,et al., Sci- Code: A research coding benchmark curated by scientists (2024), arXiv:2407.13168 [cs.AI]

  29. [29]

    W. Li, J. Ren, L. Cheng, and C. Gong, Autonomous quantum simulation through large language model agents (2026), arXiv:2601.10194 [physics.chem-ph]

  30. [30]

    S. Qiu, J. Deng, Y. Deng, H. Dong, J. Fu, M. Li,et al., PRBench: End-to-end paper reproduction in physics re- search (2026), arXiv:2603.27646 [cs.CL]

  31. [31]

    S. Qiu, Z. Cai, J. Wei, Z. Li, Y. Yin, Q.-H. Cao, C. Liu, M.-x. Luo, X.-B. Yuan, and H. X. Zhu, An end-to- end architecture for collider physics and beyond (2026), arXiv:2603.14553 [hep-ph]

  32. [32]

    Z. N. Ndum, J. Tao, J. Ford, and Y. Liu, Automating monte carlo simulations in nuclear engineering with do- main knowledge-embedded large language model agents, Energy and AI21, 100555 (2025)

  33. [33]

    Ni and M

    B. Ni and M. J. Buehler, Mechagents: Large language model multi-agent collaborations can solve mechanics problems, generate new data, and integrate knowledge, Extreme Mechanics Letters67, 102131 (2024)

  34. [34]

    Heneka, F

    C. Heneka, F. Nieser, A. Ore, T. Plehn, and D. Schiller, Large Language Models – the Future of Fundamental Physics? (2025), arXiv:2506.14757 [astro-ph.CO]

  35. [35]

    E. A. Moreno, S. Bright-Thonney, A. Novak, D. Gar- cia, and P. Harris, Ai agents can already autonomously perform experimental high energy physics (2026), arXiv:2603.20179 [hep-ex]

  36. [36]

    Amram, L

    O. Amram, L. Anzalone, J. Birk, D. A. Faroughy, A. Hallin, G. Kasieczka, M. Krämer, I. Pang, H. Reyes- Gonzalez, and D. Shih, Aspen Open Jets: Unlocking LHC Data for Foundation Models in Particle Physics, Mach.Learn.Sci.Tech.6, 030601 (2024), arXiv:2412.10504 [hep-ph]

  37. [37]

    Zhanget al., Xiwu: A Basis Flexible and Learnable LLM for High Energy Physics (2024), arXiv:2404.08001 [hep-ph]

    Z. Zhanget al., Xiwu: A Basis Flexible and Learnable LLM for High Energy Physics (2024), arXiv:2404.08001 [hep-ph]

  38. [38]

    Fanelli, J

    C. Fanelli, J. Giroux, P. Moran, H. Nayak, K. Suresh, and E. Walter, Physics Event Classification Using Large Language Models (2024), arXiv:2404.05752 [physics.data- an]

  39. [39]

    S. D. Bakshiet al., ArgoLOOM: agentic AI for fundamental physics from quarks to cosmos (2025), arXiv:2510.02426 [hep-ph]

  40. [40]

    McGreivy, B

    J. McGreivy, B. Delaney, A. Beck, and M. Williams, Seeing the Forest Through the Trees: Knowledge Re- trieval for Streamlining Particle Physics Analysis (2025), arXiv:2509.06855 [hep-ex]

  41. [41]

    Rafique, A

    A. Rafique, A. Singh, and R. Srinivas, Large Language Model Integration for Knowledge Retrieval and Interac- tion for the DUNE Experiment, in32nd International Symposium on Lepton Photon Interactions at High Ener- gies: Lepton-Photon 2025(2026) arXiv:2601.05278 [hep- ex]

  42. [42]

    Mermer, J

    I. Mermer, J. Muszyński, J. Możaryn, and K. Rosłon, Proposal of an AI-Based Support Assistant for the ALICE- FIT Detector Setup at CERN (2025), arXiv:2511.17154 [hep-ex]

  43. [43]

    Songet al., Iterated Agent for Symbolic Regression (2025), arXiv:2510.08317 [physics.comp-ph]

    Z.-Y. Songet al., Iterated Agent for Symbolic Regression (2025), arXiv:2510.08317 [physics.comp-ph]

  44. [44]

    Plehn, D

    T. Plehn, D. Schiller, and N. Schmal, MadAgents (2026), arXiv:2601.21015 [hep-ph]

  45. [45]

    Diefenbacher, A

    S. Diefenbacher, A. Hallin, G. Kasieczka, M. Krämer, A. Lauscher, and T. Lukas, Agents of Discovery (2025), arXiv:2509.08535 [hep-ph]

  46. [46]

    T. J. Jat, T. Ghosh, and K. Suresh, Retrieval-Augmented Question Answering over Scientific Literature for the Electron-Ion Collider (2026), arXiv:2604.02259 [hep-ex]

  47. [47]

    Mallampalli and S

    A. Mallampalli and S. Dasu, MITRA: An AI Assistant for Knowledge Retrieval in Physics Collaborations (2026) arXiv:2603.09800 [cs.IR]

  48. [48]

    Tan, T.-J

    J.-X. Tan, T.-J. Miao, M.-H. Zhang, X.-H. Pang, Z.- X. Liu, L.-F. Zhang, S.-H. Chen, and W. Wang, Auto- mated Extraction of Collins-Soper Kernel from Lattice 14 QCD using An Autonomous AI Physicist System (2026), arXiv:2603.22471 [hep-lat]

  49. [49]

    Liang and X.-N

    Z.-T. Liang and X.-N. Wang, Globally polarized quark- gluon plasma in non-central A+A collisions, Phys. Rev. Lett.94, 102301 (2005), [Erratum: Phys.Rev.Lett. 96, 039901 (2006)], arXiv:nucl-th/0410079

  50. [50]

    Knipfer, A

    M. Knipfer, A. Roman, K. T. Matchev, K. Matcheva, and S. Gleyzer, AI Agents for Variational Quantum Circuit Design (2026), arXiv:2602.19387 [quant-ph]

  51. [51]

    Badea, Y

    A. Badea, Y. Chen, and Y.-J. Lee, Agentic AI – Physicist Collaboration in Experimental Particle Physics: A Proof- of-Concept Measurement with LEP Open Data (2026), arXiv:2603.05735 [hep-ex]

  52. [52]

    Agrawal, N

    P. Agrawal, N. Craig, A. Madden, and I. V. Lombera, The FERMIACC: Agents for Particle Theory (2026), arXiv:2603.22538 [hep-ph]

  53. [53]

    Menzo, A

    T. Menzo, A. Roman, G. T. Fleming, S. Gleyzer, K. T. Matchev, and S. Mrenna, Agentic Diagrammatica: To- wards Autonomous Symbolic Computation in High En- ergy Physics (2026), arXiv:2603.26990 [hep-ph]

  54. [54]

    Everettet al.(JETSCAPE Collaboration), Multi- system bayesian constraints on the transport coefficients of qcd matter, Physical Review C103, 054904 (2021)

    D. Everettet al.(JETSCAPE Collaboration), Multi- system bayesian constraints on the transport coefficients of qcd matter, Physical Review C103, 054904 (2021)

  55. [55]

    Liang and X.-N

    Z.-T. Liang and X.-N. Wang, Spin alignment of vector mesons in non-central A+A collisions, Phys. Lett. B629, 20 (2005), arXiv:nucl-th/0411101

  56. [56]

    Adamczyket al.(STAR), GlobalΛhyperon polariza- tion in nuclear collisions: evidence for the most vortical fluid, Nature548, 62 (2017), arXiv:1701.06657 [nucl-ex]

    L. Adamczyket al.(STAR), GlobalΛhyperon polariza- tion in nuclear collisions: evidence for the most vortical fluid, Nature548, 62 (2017), arXiv:1701.06657 [nucl-ex]

  57. [57]

    Becattini, M

    F. Becattini, M. Buzzegoli, T. Niida, S. Pu, A.-H. Tang, and Q. Wang, Spin polarization in relativistic heavy- ion collisions, Int. J. Mod. Phys. E33, 2430006 (2024), arXiv:2402.04540 [nucl-th]

  58. [58]

    J. E. Bernhard, J. S. Moreland, and S. A. Bass, Applying bayesian parameter estimation to relativistic heavy-ion collisions: simultaneous characterization of the initial state and quark-gluon plasma medium, Nature Physics 15, 1113 (2019)

  59. [59]

    Zhang, J

    Y. Zhang, J. Zhang, Y. Suo, Y. Guo, D. Liu, M. Chen, and Y. Chao, Temperature-dependent shear viscosity in a multi-phase transport model for ultrarelativistic heavy-ion collisions at rhic and lhc, Journal of Physics G: Nuclear and Particle Physics46, 055101 (2019)

  60. [60]

    J. E. Parkkila, A. Onnerstad, S. F. Taghavi, C. Mordasini, A. Bilandzic, D. J. Kim, and J. Virta, Bayesian estimation of the specific shear and bulk viscosity of the quark-gluon plasma with additional flow harmonic observables, Physics Letters B835, 137485 (2022)

  61. [61]

    Jiaet al., Imaging shapes of atomic nuclei in high- energy nuclear collisions, Nature635, 67 (2024)

    J. Jiaet al., Imaging shapes of atomic nuclei in high- energy nuclear collisions, Nature635, 67 (2024)

  62. [62]

    Rischke, Influence of shear viscosity of quark-gluon plasma on elliptic flow in ultrarelativistic heavy-ion collisions, Phys

    H.Niemi, G.S.Denicol, P.Huovinen, E.Molnár,andD.H. Rischke, Influence of shear viscosity of quark-gluon plasma on elliptic flow in ultrarelativistic heavy-ion collisions, Phys. Rev. Lett.106, 212302 (2011)

  63. [63]

    Denicol, A

    G. Denicol, A. Monnai, and B. Schenke, Moving forward to constrain the shear viscosity of qcd matter, Phys. Rev. Lett.116, 212301 (2016)

  64. [64]

    J. E. Bernhard, J. S. Moreland, S. A. Bass, J. Liu, and U. Heinz, Quantifying properties of hot and dense qcd matter through systematic model-to-data comparison, Physical Review C94, 024907 (2016)

  65. [65]

    C. Shen, Z. Qiu, H. Song, J. Bernhard, S. Bass, and U. Heinz, The iebe-vishnu code package for relativistic heavy-ion collisions, Computer Physics Communications 199, 61 (2016)

  66. [66]

    Karpenko, P

    I. Karpenko, P. Huovinen, H. Petersen, and M. Bleicher, A 3+1 dimensional viscous hydrodynamic code for rela- tivistic heavy ion collisions, Computer Physics Communi- cations185, 3016 (2014)

  67. [67]

    Schenke, S

    B. Schenke, S. Jeon, and C. Gale, Elliptic and triangular flow in event-by-event (3+1)d viscous hydrodynamics, Physical Review Letters106, 042301 (2011)

  68. [68]

    Novak, K

    J. Novak, K. Novak, S. Pratt, J. Vredevoogd, C. E. Coleman-Smith, and R. L. Wolpert, Determining fun- damental properties of matter created in ultrarelativistic heavy-ion collisions, Physical Review C89, 034917 (2014)

  69. [69]

    Putschke, K

    J. Putschke, K. Kauder, E. Khalaj, A. Angerami, S. Bass, S. Cao, J. Coleman, L. Cunqueiro, T. Dai, L. Du,et al., The JETSCAPE framework (2019), arXiv:1903.07706 [nucl-th]

  70. [70]

    G. Nijs, W. van der Schee, U. Gürsoy, and R. Snellings, Bayesian analysis of heavy ion collisions with the heavy ion computational framework trajectum, Physical Review C103, 054909 (2021)

  71. [71]

    Zhang, S

    Y. Zhang, S. A. Khan, A. Mahmud,et al., Exploring the role of large language models in the scientific method: from hypothesis to discovery, npj Artificial Intelligence1, 14 (2025)

  72. [72]

    L.-G. Pang, H. Petersen, and X.-N. Wang, Pseudorapidity distribution and decorrelation of anisotropic flow within the open-computing-language implementation CLVisc hy- drodynamics, Phys. Rev. C97, 064918 (2018)

  73. [73]

    Wu, G.-Y

    X.-Y. Wu, G.-Y. Qin, L.-G. Pang, and X.-N. Wang, (3+1)- D viscous hydrodynamics at finite net baryon density: Identified particle spectra, anisotropic flows, and flow fluctuations across energies relevant to the beam-energy scan at RHIC, Phys. Rev. C105, 034909 (2022)

  74. [74]

    The hydrodynamic evolution starts at the initial proper timeτ0 = 0 .6fm, and the system evolves until the local temperature drops to the freeze-out temperature Tfrz = 0.137GeV

    with a nucleon width ofw = 0.5fm and a gamma shape parameter k = 1.0. The hydrodynamic evolution starts at the initial proper timeτ0 = 0 .6fm, and the system evolves until the local temperature drops to the freeze-out temperature Tfrz = 0.137GeV. The HotQCD equation of state is adopted throughout the evolution [75]. Once tasked, the agent autonomously mod...

  75. [75]

    D. Everettet al.(JETSCAPE), Phenomenological con- straints on the transport properties of QCD matter with data-drivenmodelaveraging,Phys.Rev.Lett.126,242301 (2021), arXiv:2010.03928 [hep-ph]

  76. [76]

    Everettet al.(JETSCAPE), Multisystem Bayesian constraints on the transport coefficients of QCD matter, Phys

    D. Everettet al.(JETSCAPE), Multisystem Bayesian constraints on the transport coefficients of QCD matter, Phys. Rev. C103, 054904 (2021), arXiv:2011.01430 [hep- ph]

  77. [77]

    Anthropic, Claude Code, https://claude.ai/code (2025), [Online; accessed 2025]

  78. [78]

    Moonshot AI, Kimi-CLI, https://www.kimi.com/code (2025), [Online; accessed 2025]

  79. [79]

    J. S. Moreland, J. E. Bernhard, and S. A. Bass, Alter- native ansatz to wounded nucleon and binary collision scaling in high-energy nuclear collisions, Phys.Rev.C92, 011901 (2015), arXiv:1412.4708 [nucl-th]

  80. [80]

    Bazavovet al.(HotQCD), The equation of state in (2+1)-flavor qcd, Phys

    A. Bazavovet al.(HotQCD), The equation of state in (2+1)-flavor qcd, Phys. Rev. D90, 094503 (2014), arXiv:1407.6387 [hep-lat]

Showing first 80 references.