{"id":"1265edd6-56c1-410a-85bd-5715b59d367c","arxiv_id":"2507.13447","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A Deep Q-Network using matrix-element rewards reconstructs parton assignments in collider events, enabling theory-based tagging and anomaly detection without labels.","lead":"Researchers train a reinforcement-learning agent to assign final-state particles to partons, rewarded at each step by the change in the tree-level matrix element. The method aims to give an interpretable, label-free alternative to black-box classifiers for LHC event reconstruction, W-polarization tagging, and anomaly detection.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No quantitative check that the ME-maximizing assignment equals the true parton assignment; all downstream results inherit this unvalidated premise.","rationale":"The reader's weakest assumption is exactly the load-bearing point: the reward is the tree-level matrix element, and the paper concedes that maximizing it is \"not always the true assignment,\" yet no quantitative assignment-accuracy metric is reported. I agree with the reader's conditional verdict. The concern is not that the method is wrong; it is that the strongest claims—label-free interpretability, optimal likelihood-ratio classification, and robust scaling—all depend on the mapping quality, and that quality is only shown indirectly through mass histograms and matrix-element ratios. The proposed exact-permutation test for t-tbar would settle whether the RL policy matches the ME-argmax and whether the ME-argmax matches Monte Carlo truth. The parton-level limitation is real but is explicitly acknowledged and is not the scientific crux; the missing code and data are artifacts concerns rather than internal correctness issues. Therefore I do not move the verdict: CONDITIONAL remains appropriate pending this quantitative check.","tokens_in":11715,"tokens_out":5547,"duration_ms":75972,"concrete_test":"On the 20k-event t-tbar test set, enumerate all 720 possible particle-to-parton assignments for each event, compute the MadGraph |M|^2 for every assignment using the same settings as training, and record two rates: (i) the fraction of events where the global argmax-ME assignment equals the MC-truth assignment, and (ii) the fraction where the trained RL policy's final assignment equals the global argmax. Report both with statistical uncertainties. If the argmax-to-truth rate is low while the RL-to-argmax rate is high, the reward is demonstrably optimizing an objective that does not recover the true parton assignment. If the RL-to-argmax rate is also low, the policy is not even finding the ME-optimal assignment. Either outcome would quantify how much interpretability and the claimed optimality of Eq. 7 are degraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central premise is stated in Section 3: \"We define the optimal assignment as the one which maximises the matrix element for that process, although this is not always the 'true' assignment.\" Every downstream result—the interpretable event graph, the mass histograms in Figs. 4–6, the W-polarization likelihood ratio in Eq. 7, and the anomaly score in Eq. 8—uses the agent's assignment as though it were physically correct. Yet the paper's validation never directly tests assignment correctness against known parton-level truth. The right-hand plots in Fig. 3 compare the matrix element of the predicted assignment with the matrix element of the true assignment; but because the predicted assignment is chosen to maximize the matrix element, close agreement in matrix-element value does not imply close agreement in assignment. Similarly, good reconstructed mass peaks can arise even when some assignments are wrong, since sub-assignments such as the W-decay quarks may be correct while the b-quark pairing is wrong, as the authors note for t-tbar-W and t-tbar-t-tbar. Thus the claim that the method is \"fully interpretable\" and that Eq. 7 is an optimal classifier in the ideal scenario is not yet supported: the ideal scenario requires the predicted mapping to be the correct mapping, and that has not been measured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a theory-informed reinforcement-learning framework for assigning final-state particles to partons in hadron collider events. A transformer-based Deep Q-Network is trained with a reward equal to the logarithmic change in the tree-level matrix element, and the resulting policy produces a parton-level event graph. The method is demonstrated on parton-level $t\\bar{t}$, $t\\bar{t}W$, and $t\\bar{t}t\\bar{t}$ events, with reconstructed mass peaks shown for intermediate $W$ bosons and top quarks. The same pipeline is then used to construct a $W^+W^-$ polarization tagger via the likelihood ratio in Eq. (7) and an anomaly detector based on the inverse background matrix element in Eq. (8). The central claims are that the method is label-free, fully interpretable, robust across processes, and naturally respects physical symmetries.","tokens_in":11948,"tokens_out":5009,"duration_ms":60522,"significance":"If the central claims are substantiated, the framework would be a useful contribution: it unifies event reconstruction, polarization tagging, and anomaly detection in a single theory-driven pipeline, and it makes the assignment step interpretable by construction. The use of matrix-element rewards rather than labels is a genuine conceptual step, and the demonstration on three processes of increasing combinatorial complexity is a strength. The paper also ships a public code repository, which is valuable for reproducibility. However, the evidence currently falls short of the claims. The validation never directly measures whether the ME-maximizing assignment coincides with the true parton-level assignment, and the claimed classifier performance is not quantified with AUCs or rejection factors. The parton-level approximation is acknowledged, but the 'robust performance' claim is supported only by qualitative histograms. The central idea is plausible and the manuscript is within the scope of the journal, but the missing quantitative validation is load-bearing for the paper's main conclusions.","major_comments":[{"comment":"The core validation compares the matrix element of the predicted assignment with the matrix element of the true assignment (right panels of Fig. 3). This is not an independent test of reconstruction: because the policy is trained to maximize the matrix element, the predicted ME will be close to the true ME whenever the true assignment is near the ME maximum, even if many individual partons are mismatched. The paper itself states in Section 3 that the ME-maximizing assignment is \"not always the true assignment,\" and Section 4 notes that $t\\bar{t}W$ and $t\\bar{t}t\\bar{t}$ events suffer from incorrect W-b-quark matching. The manuscript therefore needs a direct quantitative measure of assignment correctness, for example the fraction of events in which all parton assignments are correct, or per-parton matching accuracies, for each process. Until this is reported, the claims that the policy is \"fully interpretable\" and that Eq. (7) is optimal in the ideal scenario are not supported.","section":"Section 3 and Fig. 3"},{"comment":"The proposed classifiers are shown only as ROC curves, with no reported AUC for the RL-based W-polarization tagger or for the theory-informed anomaly detector. The text quotes AUC 0.904 for the likelihood ratio computed with true assignments and AUC 0.645 for the autoencoder, but does not quote the corresponding numbers for the proposed method. To support \"robust performance\" and \"powerful anomaly-detection performance,\" please report AUC and background-rejection factors, with statistical uncertainties, for the RL-based scores in both applications. This is particularly important because the theoretical optimality of Eq. (7) holds only for the correct particle-parton mapping.","section":"Section 5, Figs. 8 and 9"},{"comment":"The claim that the method \"maintains robust performance across all processes\" is supported only by qualitative mass histograms. There is no per-process quantitative metric, such as reconstruction efficiency with a defined matching criterion, the fraction of correctly assigned events, or the widths of reconstructed mass peaks relative to the truth distribution, and no uncertainty estimate. Given the acknowledged parton-level approximation and the absence of detector effects, \"robust performance\" overstates the evidence presented. Please either add quantitative benchmarks for reconstruction accuracy or soften the claim in the abstract and Section 4.","section":"Section 4 and Abstract"}],"minor_comments":[{"comment":"The text writes \"P rt\" where a summation is intended; please correct the notation and define the sum over time steps explicitly.","section":"Section 2, Eq. (1)"},{"comment":"The sentence \"The ensure that the agent is able to learn from a diverse range of state-action combinations\" contains a typo and should read \"To ensure.\"","section":"Section 2, training procedure"},{"comment":"The right panels are described as histograms, but the axis labels and the meaning of the dashed line (\"initial mapping\") are not defined in the caption; please clarify.","section":"Fig. 3"},{"comment":"The log-ratio reward in Eq. (6) and the inverse matrix-element score in Eq. (8) are undefined if the matrix element vanishes; since tree-level matrix elements can vanish at phase-space boundaries, please state how such cases are handled in the implementation.","section":"Eq. (6) and Eq. (8)"},{"comment":"The notation \"$\\epsilon_b^{-1}(\\epsilon_s=0.3) \\simeq 8$\" is not standard and lacks a space; please define the background-rejection notation and explain how the autoencoder's ROC curve is evaluated.","section":"Section 5, anomaly detection"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of a hep-ph journal and the core idea is interesting, but the current validation is incomplete in a way that affects the central claims. I would ask the authors to add direct assignment-accuracy metrics and quantitative classifier performance numbers before reconsidering it. There are no concerns about citation practice or novelty disclosure beyond what is stated in the report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper has a genuinely new idea — casting parton assignment as a Markov decision process with a log-matrix-element reward — and it demonstrates the concept on three processes of increasing combinatorial complexity. But the validation is thinner than the abstract's claims: the core premise that the ME-maximizing assignment is physically correct is never directly tested, and several quantitative results are missing.\n\nWhat's new: the DQN/transformer combination for parton-to-particle assignment is new, and the downstream uses (polarization tagging, anomaly detection via inverse ME) are natural extensions. The paper is honest about the parton-level approximation and explicitly notes that the ME-maximizing assignment is not always the true one. The mass histograms for ttbar are convincing; for ttbarW and ttbarttbar they are suggestive of good but imperfect reconstruction.\n\nThe soft spots are real. The right-hand panels of Fig. 3 compare the matrix element of the predicted assignment with the matrix element of the true assignment — that's self-referential, because predicted is chosen to maximize the ME. Good mass peaks can coexist with incorrect b-jet pairings, as the authors themselves note. The paper reports no reconstruction efficiency, no assignment accuracy against MC truth, and no AUC for the actual RL-based W tagger or anomaly detector; the quoted 0.904 AUC is for perfect reconstruction, not for the RL mapping. The GitHub link is a promise, not a release. These gaps undercut the 'fully interpretable' and 'robust performance' claims.\n\nThis does not sink the paper — it's a proof-of-concept with acknowledged limitations — but it needs a quantitative re-analysis before the claims are supportable. I'd send it to peer review, and I'd ask for assignment-correctness metrics, ROC/AUC numbers for the RL pipeline, and ideally one detector-level example.\n\nWho should read it: the hep-ph ML community, especially people working on matrix-element methods and anomaly detection. It's a plausible step toward a unified interpretable pipeline, but the current evidence is suggestive, not conclusive. I wouldn't cite it yet, and I'd put it on the reading group as a maybe.","headline":"Clever DQN/ME parton-assignment pipeline that's novel and plausible, but the self-referential validation and missing assignment-accuracy metrics undercut the strong claims.","tokens_in":12435,"tokens_out":3665,"would_cite":false,"duration_ms":39804,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A reinforcement-learning agent that maps final-state particles to partons by maximising the tree-level matrix element can reconstruct LHC events, tag W polarisation, and flag anomalies without labelled data.","keywords":["reinforcement learning","deep Q-network","matrix element method","parton assignment","event reconstruction","LHC physics","anomaly detection","W boson polarization"],"falsifier":"Take the same trained agents and compare their final particle-to-parton assignments, event by event, against Monte Carlo truth labels rather than only comparing reconstructed mass peaks; if the agreement is no better than the best random assignment, or if the policy systematically prefers non-truth assignments in events where the matrix-element-maximising mapping differs from truth, the label-free reconstruction claim is refuted.","tokens_in":11532,"feed_emoji":"⚛️","tokens_out":5535,"duration_ms":59801,"temperature":0.7,"pith_summary":"This paper argues that the hard combinatorial problem of assigning measured particles to the partons of a hypothesized scattering process can be solved without any labelled training data. The agent is a transformer-based Deep Q-Network that repeatedly swaps particle-to-parton assignments, receiving a reward equal to the logarithmic change in the exact tree-level matrix element after each swap. Because the reward comes from first-principles theory, the learned policy is interpretable and every reconstructed particle can be traced to a partonic origin. On simulated $t\\bar{t}$, $t\\bar{t}W$, and $t\\bar{t}t\\bar{t}$ events the reconstructed intermediate masses recover the true distributions, and the same machinery builds a longitudinal $WW$ polarisation tagger and a theory-informed anomaly score. If right, this unifies event reconstruction, tagging, and anomaly detection in one transparent pipeline for the High-Luminosity LHC.","feed_headline":"No-label AI reconstructs LHC events from the quantum matrix element","feed_subtitle":"Rewarded by first-principles amplitudes, the same pipeline reconstructs tops, tags W polarisation, and flags anomalies.","key_machinery":"The load-bearing mechanism is the theory-informed reward: at each step of the Markov decision process the agent chooses a swap (or 'do nothing') action and receives $r_t = \\log(m_{t+1}/m_t)$, the logarithmic change in the tree-level matrix element. A permutation-invariant transformer with a positional encoding for the parton identities approximates the Q-function via the Bellman equation, and the policy greedily selects the action with the highest Q-value over $T=N-3$ steps. This reward is what makes the approach label-free and interpretable: the network learns only the combinatorial mapping, while the matrix element carries all the physics, including spin correlations and interference effects.","core_discovery":"The paper's central claim is that the mapping from final-state particles to matrix-element partons can be treated as an optimal-control problem whose reward is the exact tree-level matrix element itself. The agent's policy is obtained purely from physics: each action changes the assignment, and the reward $r_t = \\log(m_{t+1}/m_t)$ tells the agent whether the new assignment is more consistent with the theory hypothesis. Once trained, the policy has no labels, no handcrafted observables, and no black-box decisions: the network only provides the permutation, while all classification and anomaly scores are computed from the analytic matrix element, so the scores respect the symmetries of the theory. The authors validate this on $t\\bar{t}$, $t\\bar{t}W$, and $t\\bar{t}t\\bar{t}$ reconstruction, on longitudinal $W^+W^-$ tagging, and on $t\\bar{t}$-background anomaly detection, finding robust performance that scales with combinatorial complexity.","pith_inferences":["Beyond the paper, the reward could be upgraded to include higher-order corrections or detector response, which would make the same assignment policy applicable to reconstructed objects rather than parton-level four-momenta.","A testable extension is to compare the agent's final assignment against parton-shower truth labels event-by-event; the paper only shows mass distributions, so an explicit assignment-accuracy number would sharpen the label-free claim.","Because the matrix-element-maximising assignment is not always the true one, the method's practical ceiling depends on how often the two agree in a given process; quantifying that overlap would predict where the approach will fail."],"forward_implications":["Event reconstruction, polarisation tagging, and anomaly detection become one pipeline: the same trained policy supplies the particle-to-parton mapping, and the matrix element supplies the score.","No labels or simulated 'truth' are needed for training, so the method avoids simulation-dependent biases of supervised classifiers.","Because scores are computed from the matrix element, they automatically respect Lorentz invariance and the physical symmetries of the process.","The approach scales to large combinatorial complexity: reconstruction quality remains high from 720 mappings in $t\\bar{t}$ to about $4\\times10^8$ in $t\\bar{t}t\\bar{t}$.","The longitudinal $W^+W^-$ tagger and the $t\\bar{t}$-background anomaly score outperform a plain AutoEncoder on the restricted-mass test, while retaining interpretability of which partons form which resonance."],"supporting_citations":[{"why":"Introduces the dynamical-likelihood/matrix-element approach that this paper recasts as a reinforcement-learning reward.","marker":"[1]"},{"why":"Extends matrix-element reconstruction to general event deconstruction, the paradigm being automated here.","marker":"[5]"},{"why":"Gives the Neyman-Pearson optimality criterion that justifies using matrix-element likelihood ratios for tagging.","marker":"[14]"},{"why":"Defines Deep Q-Learning with experience replay, the training scheme the paper adopts.","marker":"[16]"},{"why":"Shows deep Q-networks can learn complex policies, motivating the transformer-based DQN architecture.","marker":"[17]"},{"why":"Foundational Q-learning algorithm behind the Bellman-equation update.","marker":"[18]"},{"why":"MadGraph generates the simulated events and evaluates the tree-level matrix elements used as rewards.","marker":"[20]"}],"fun_headline_variants":["Reinforcement learning uses matrix element to assign LHC particles","Matrix-element reward makes particle-to-parton mapping label-free","Rewarding AI with the matrix element for LHC event reconstruction","Quantum matrix element guides AI's particle assignment at LHC","No labels needed: AI learns particle assignment from physics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's reward assumes that the assignment with the largest tree-level matrix element for the assumed process is the physical assignment, and the authors note this is not always the true assignment; if this mismatch is frequent, the reconstructed topology and all scores built on it inherit the error.","fun_headline_variants_meta":{"raw":{"variants":["Reinforcement learning uses matrix element to assign LHC particles","Matrix-element reward makes particle-to-parton mapping label-free","Rewarding AI with the matrix element for LHC event reconstruction","Quantum matrix element guides AI's particle assignment at LHC","No labels needed: AI learns particle assignment from physics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00024,"raw_usage":{"total_tokens":1536,"prompt_tokens":982,"completion_tokens":554,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":486}},"tokens_in":598,"tokens_out":554,"duration_ms":6575,"temperature":1.0,"reasoning_tokens":486,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:24:11.357008+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same trained agents and compare their final particle-to-parton assignments, event by event, against Monte Carlo truth labels rather than only comparing reconstructed mass peaks; if the agreement is no better than the best random assignment, or if the policy systematically prefers non-truth assignments in events where the matrix-element-maximising mapping differs from truth, the label-free reconstruction claim is refuted.","supporting_citations":[{"cited_title":"Finding physics signals with event deconstruction","cited_arxiv_id":"1402.1189","evidence_quote":"Extends matrix-element reconstruction to general event deconstruction, the paradigm being automated here."}],"review_version":1}