{"id":"fd7edcaf-7dde-4231-92b1-9fe0db8c3232","arxiv_id":"2509.02502","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"First LHC measurement of Lambda_c baryon elliptic flow shows a 3.6-sigma hint of baryon-meson splitting in the charm sector.","lead":"ALICE reports the first measurement of elliptic flow of Lambda_c baryons at the LHC, together with new D-meson v2 results from Pb-Pb collisions at 5.36 TeV. The data show a hint that charm baryons flow more strongly than charm mesons at high momentum, which would point to quark coalescence in the quark-gluon plasma.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The '3.6-sigma first evidence' claim rests on an undefined significance: no systematic uncertainties, correlations, or trials factor for the post-hoc pT>4 GeV/c window are specified.","rationale":"The reader's conditional verdict is appropriate: the measurement is a solid ALICE preliminary result with a plausible physics signal, but the headline 'first evidence of baryon-meson splitting in the charm sector' carries a statistical weight that is not fully substantiated. My stress-test agrees with the reader on the look-elsewhere concern, while being more skeptical that the feed-down/BBD contamination is the primary risk: because non-prompt v2 is smaller, misassigning non-prompt Lambda_c as prompt would dilute the observed prompt-Lambda_c excess rather than generate it, unless the contamination also shifts the prompt D0 v2 differently. The more pressing gap is the undefined significance: no covariance matrix, no systematic uncertainty breakdown, and no statement about trial factors are provided. The concrete test addresses exactly that gap. I therefore recommend keeping the reader's CONDITIONAL verdict unchanged, with the condition that the significance be recomputed and reported transparently before the 'first evidence' claim is used in the literature.","tokens_in":4099,"tokens_out":4582,"duration_ms":49056,"concrete_test":"Run a null-hypothesis toy: take the published D0 and prompt Lambda_c v2 points with their total statistical-plus-systematic covariance, set the true difference to zero, and compute the distribution of the largest excess over pT thresholds scanned in 1 GeV steps. Also recompute the significance at pT>4 with the full covariance, including the feed-down systematic. If the global p-value is above 0.003 (i.e. below about 3 sigma), the paper should reword 'first evidence' as a hint or indication rather than evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is in Section 4: 'At pT > 4 GeV/c, the prompt Lambda_c-baryon v2 is larger than that of D0 mesons with a significance of 3.6 sigma, which provides the first evidence of baryon-meson splitting in the charm sector.' This is the load-bearing assertion, but the proceedings do not define that significance. It is not stated whether the 3.6 sigma includes (a) systematic uncertainties on v2 and their correlations between D0 and Lambda_c, or (b) any penalty for choosing the pT>4 GeV/c interval after inspecting the data. The text itself notes that below 4 GeV/c the prompt Lambda_c and D0 v2 values are compatible, so the threshold is not a pre-specified interval; if multiple pT thresholds were scanned, the local significance overstates the evidence. A related but directionally weaker concern is feed-down: since the non-prompt v2 is smaller than the prompt v2, leakage of non-prompt Lambda_c into the prompt sample would tend to dilute the excess, not create it; however, a bias in the data-driven feed-down fraction could still affect the correction and alter the difference. The decisive question is whether the 3.6 sigma survives a global significance calculation with the full systematic covariance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a proceedings contribution from the ALICE collaboration reporting elliptic flow (v2) measurements of charm hadrons in Pb-Pb collisions at sqrt(s_NN) = 5.36 TeV from Run 3 data. It presents v2 of prompt D0, D+, Ds+ mesons in several centrality classes and, for the first time at the LHC, prompt and non-prompt Lambda_c+ baryons in the 30-50% centrality class. The analysis uses the scalar-product method with BDT-based candidate selection and a data-driven feed-down correction. The central physics claim, stated in Section 4, is that at pT > 4 GeV/c the prompt Lambda_c+ v2 exceeds the prompt D0 v2 with a significance of 3.6 sigma, which is interpreted as the first evidence of baryon-meson splitting in the charm sector. The measurements are compared with several model predictions.","tokens_in":4325,"tokens_out":3211,"duration_ms":31282,"significance":"If the reported 3.6-sigma effect is robust, the paper delivers a genuinely new result: the first measurement of Lambda_c+ elliptic flow in heavy-ion collisions and a direct experimental handle on charm baryon-meson splitting, which is sensitive to coalescence hadronization and heavy-quark transport in the QGP. The paper's strengths include the use of a well-established scalar-product method, a mature data-driven feed-down procedure from a previous ALICE publication, and comparison with a broad set of model predictions. However, the headline evidence claim is currently undersupported because the paper does not define the quoted significance, does not specify whether systematic uncertainties are included in the displayed error bars, and does not address the post-hoc choice of the pT > 4 GeV/c window. These are load-bearing issues for the central conclusion, so the paper needs revision before the claim can be accepted as stated.","major_comments":[{"comment":"The central claim 'At pT > 4 GeV/c, the prompt Lambda_c+-baryon v2 is larger than that of D0 mesons with a significance of 3.6 sigma' is not supported by a defined significance calculation. Please specify whether the 3.6 sigma is based only on statistical uncertainties or on the full statistical-plus-systematic covariance, and explain how correlations between the D0 and Lambda_c+ measurements are handled. The manuscript currently gives no information on this point, which is essential for judging the claim.","section":"Section 4, Fig. 1 (right)"},{"comment":"The pT > 4 GeV/c interval appears to be chosen after inspecting the data, as the text first states that below 4 GeV/c the prompt Lambda_c+ and D0 v2 are compatible and then quotes the 3.6-sigma difference above that threshold. If multiple pT thresholds or intervals were scanned, the quoted local significance overstates the evidence. Please state whether the threshold was pre-specified, and if not, apply a look-elsewhere correction or report the trials factor.","section":"Section 4, pT range selection"},{"comment":"The prompt and non-prompt separation relies on a multi-class BDT and a data-driven feed-down fraction following Ref. [5], but the manuscript does not quantify the systematic uncertainty associated with this correction. Since the non-prompt v2 is smaller than the prompt v2, simple leakage would dilute the difference, but a bias in the feed-down fraction itself could still affect the extracted prompt Lambda_c+ v2. Please state how the feed-down systematic is evaluated and whether it is included in the uncertainties shown in Fig. 1 (right).","section":"Section 2, feed-down procedure"}],"minor_comments":[{"comment":"Reference [1] (Liu and Liu, Quark-gluon plasma formation time and direct photons) does not appear to support the statement that elliptic flow originates mainly from the initial-state spatial asymmetry; please cite a more appropriate reference for this foundational statement.","section":"Section 1, reference [1]"},{"comment":"The phrase 'combining couples or triplets of tracks' should be 'combining pairs or triplets of tracks'.","section":"Section 2, wording"},{"comment":"The figure captions do not specify whether the error bars represent statistical uncertainties only or total uncertainties. Please add this information to each caption.","section":"Figures 1 and 2"},{"comment":"The text highlights the first measurement of prompt D0 v2 at pT < 1 GeV/c but does not comment on possible low-pT acceptance or efficiency corrections; a brief statement on this would be helpful.","section":"Section 3"},{"comment":"The summary mentions future constraints from RAA and Run 3 data, but it would be useful to also mention whether the current v2 measurements are already compatible with the model predictions in a quantitative sense, rather than only qualitatively.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a conference proceedings contribution, so the level of detail is lower than a full ALICE paper. However, the headline claim of 'first evidence of baryon-meson splitting in the charm sector' is a strong physics claim that goes beyond a preliminary measurement. The lack of a defined significance and the unaddressed look-elsewhere effect are not merely presentational; they directly affect the validity of the evidence claim. I would encourage the authors to either add the missing statistical information or explicitly soften the claim to a '3.6-sigma local significance' and describe it as an indication rather than evidence. The paper is otherwise acceptable as a useful preliminary update."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about arXiv:2509.02502. First, the measurements are genuinely new: first Lambda_c v2 at the LHC, first D0 v2 below 1 GeV/c, and new Run 3 data at 5.36 TeV. That is real progress and the analysis follows established ALICE methods. Second, the headline claim—'first evidence of baryon-meson splitting in the charm sector' at 3.6 sigma—is not backed up by the proceedings as written. The significance is undefined: no statement about whether systematics are included, no note on correlations between D0 and Lambda_c, and no trials factor for choosing the pT > 4 GeV/c window after looking at the data. The text itself says the splitting appears only above 4 GeV/c, so the threshold is post-hoc. That makes the evidence claim weaker than the number alone suggests.\n\nWhat the paper does well is straightforward: cleanly reports new v2 measurements for several charm hadrons in 30–50% centrality, compares them to light-flavor results and models, and notes where models deviate. The model comparisons are qualitative, but that is normal for a proceedings. The feed-down handling follows the previous ALICE non-prompt D0 analysis, with a data-driven fraction; the reference is appropriate. On the direction of possible bias, the stress-test note is right: non-prompt contamination would likely dilute the prompt Lambda_c excess, not create it, so the splitting may survive better than the 3.6 sigma itself implies. Still, a bias in the feed-down correction could move the difference either way.\n\nThe citation pattern looks clean—no self-citation inflation, the key methodological reference is the earlier ALICE paper. The abstract and text are honest about what is measured. The soft spots are the significance definition, the selected pT interval, and the absence of uncertainty breakdowns. For a proceedings, some of that is expected, but the phrase 'first evidence' elevates the claim beyond what the current write-up supports.\n\nWho gets value: heavy-ion phenomenologists working on charm transport and hadronization. This is exactly the kind of measurement that constrains coalescence models. It deserves a serious referee. My recommendation: send it to peer review, but expect the collaboration to either define the significance with full systematics and a trials factor, or soften the 'evidence' language to 'hint' or 'indication.' The measurement itself is solid enough to publish; the interpretation needs calibration.","headline":"First LHC Lambda_c v2 is a real step forward, but the '3.6 sigma first evidence' splitting claim outruns what the proceedings actually define.","tokens_in":4884,"tokens_out":1089,"would_cite":false,"duration_ms":11816,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["25.75.-q","25.75.Ld"],"model":"deepseek-v4-flash","headline":"The paper reports the first LHC measurement of prompt and non-prompt $\\Lambda_c^+$ baryon elliptic flow and a 3.6$\\sigma$ baryon-meson splitting at $p_T>4$ GeV/$c$.","keywords":["elliptic flow","charm quark","Lambda_c baryon","quark-gluon plasma","baryon-meson splitting","Pb-Pb collisions","hadronization","heavy-ion collisions"],"falsifier":"A Monte Carlo closure test would settle this: run the same analysis on simulated Pb--Pb events with a known feed-down fraction and check whether the extracted prompt $\\Lambda_c^+$ $v_2$ at $p_T>4$ GeV/$c$ matches the true value within the quoted uncertainty. An independent decay-length-based tagger should also reproduce the same 3.6$\\sigma$ splitting; if it does not, the BDT-based separation is the likely source.","tokens_in":3889,"feed_emoji":"🌀","tokens_out":14052,"duration_ms":117794,"temperature":0.7,"pith_summary":"These proceedings report the ALICE measurement of the elliptic flow ($v_2$) of charm hadrons in Pb--Pb collisions at $\\sqrt{s_{\\mathrm{NN}}}=5.36$ TeV from Run 3, including the first-ever LHC measurements of prompt and non-prompt $\\Lambda_c^+$ baryon $v_2$. The central result is that for $p_T>4$ GeV/$c$, prompt $\\Lambda_c^+$ baryons have a larger $v_2$ than prompt $D^0$ mesons by 3.6 standard deviations, described as the first evidence of baryon-meson splitting in the charm sector. If the effect is real, charm quarks do not hadronize in isolation; they acquire extra collective motion when they form baryons, most naturally through quark coalescence in the quark-gluon plasma. The paper also reports the first prompt $D^0$ $v_2$ measurement below 1 GeV/$c$, a low-$p_T$ mass ordering relative to light flavor, and a 2.1$\\sigma$ hint that $D_s^+$ $v_2$ is smaller than the non-strange D-meson $v_2$.","feed_headline":"Charm baryons flow more than charm mesons at 3.6 sigma","feed_subtitle":"First Lambda_c elliptic-flow measurement at the LHC shows charm baryons pick up extra flow from the quark-gluon plasma.","key_machinery":"The analysis chains the scalar-product $v_2$ estimator, a multi-class boosted-decision-tree classifier trained following the Ref. [5] strategy, simultaneous fits to the invariant-mass and $v_2$-versus-mass distributions, and a data-driven estimate of the feed-down fraction, followed by a linear fit in feed-down fraction that separates prompt and non-prompt $v_2$ components. The load-bearing part is the prompt/non-prompt separation: the claimed 3.6$\\sigma$ splitting compares prompt $\\Lambda_c^+$ with prompt $D^0$, so any leakage of beauty-decay candidates, which have smaller $v_2$, into the prompt sample would artificially enhance the splitting.","core_discovery":"The paper's central claim is that in 30--50% central Pb--Pb collisions at $\\sqrt{s_{\\mathrm{NN}}}=5.36$ TeV, prompt $\\Lambda_c^+$ baryons have a larger elliptic flow than prompt $D^0$ mesons for $p_T>4$ GeV/$c$, with a significance of 3.6$\\sigma$. This is presented as the first evidence of baryon-meson splitting in the charm sector. The non-prompt (beauty-decay) $\\Lambda_c^+$ and $D^0$ have compatible $v_2$ values that are smaller than their prompt counterparts, which the authors attribute to the larger beauty-quark mass and its longer relaxation time in the medium. The measurements also show a mass ordering at low $p_T$ between charm and light-flavor hadrons, a flat or decreasing trend at intermediate $p_T$, and convergence of different species for $p_T>10$ GeV/$c$, which the authors connect to path-length-dependent parton energy loss.","pith_inferences":["If the 3.6$\\sigma$ splitting survives with more Run 3 data, a natural extension is a finer centrality scan of the $\\Lambda_c^+$/$D^0$ $v_2$ ratio; coalescence models would predict the largest splitting in the centrality range where the partonic phase is longest.","The quoted significance was identified after scanning $p_T$ intervals, so a look-elsewhere correction over the scanned bins and particle species would clarify whether 3.6$\\sigma$ is a single-bin fluctuation.","A coupled analysis of $\\Lambda_c^+/D^0$ ratios and $v_2$ could distinguish recombination from fragmentation more sharply than either observable alone, because the two mechanisms make different joint predictions for yield and flow.","The same BDT plus feed-down machinery could be turned on beauty baryons, for which a baryon-meson splitting would be expected at lower transverse momentum because beauty quarks are heavier and thermalize later."],"forward_implications":["If the effect is real, charm baryons will be confirmed to acquire stronger collective flow than charm mesons at $p_T>4$ GeV/$c$ when the larger 2024--2025 data sample is analyzed.","The ordering baryon $>$ meson, with strange $D_s^+$ slightly below non-strange $D$ mesons, gives hadronization models a multi-species benchmark that transport-plus-coalescence calculations must reproduce simultaneously.","The smaller $v_2$ of non-prompt charm hadrons relative to prompt ones supports the picture that beauty quarks thermalize more slowly than charm quarks in the quark-gluon plasma.","The convergence of charm and light-flavor $v_2$ above 10 GeV/$c$ indicates that at high momentum the flow signal is dominated by path-length-dependent parton energy loss rather than by hadronization.","The centrality dependence of the charm $v_2$ pattern provides a new handle on how collision eccentricity, QGP size, and QGP lifetime combine to set the final flow."],"supporting_citations":[{"why":"Supplies the scalar-product estimator used to extract $v_2$ from azimuthal correlations.","marker":"[2]"},{"why":"Provides the boosted-decision-tree implementation used for the multi-class ML candidate selection.","marker":"[4]"},{"why":"The previous non-prompt $D^0$ $v_2$ measurement whose training strategy and data-driven feed-down method are reused for prompt/non-prompt $\\Lambda_c^+$ separation; this is the key methodological reference.","marker":"[5]"},{"why":"Provides the light-flavor $\\pi^+$ $v_2$ baseline used for the low-$p_T$ mass-ordering comparison.","marker":"[6]"}],"fun_headline_variants":["First Lambda_c flow at LHC shows baryon-meson splitting at 3.6 sigma","Charm baryons pick up extra flow from QGP, ALICE finds 3.6 sigma","Baryon-meson splitting in charm sector: first LHC evidence at 3.6 sigma","Lambda_c v2 exceeds D0 v2 in Pb-Pb, ALICE's first at 3.6 sigma"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 3.6-$\\sigma$ baryon-meson splitting stands on the assumption that the classifier and feed-down correction cleanly separate prompt $\\Lambda_c^+$ baryons from those produced in beauty decays; if even a modest fraction of the smaller-flow non-prompt baryons leaks into the prompt sample, the splitting could be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["First Lambda_c flow at LHC shows baryon-meson splitting at 3.6 sigma","Charm baryons pick up extra flow from QGP, ALICE finds 3.6 sigma","Baryon-meson splitting in charm sector: first LHC evidence at 3.6 sigma","Lambda_c v2 exceeds D0 v2 in Pb-Pb, ALICE's first at 3.6 sigma"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000618,"raw_usage":{"total_tokens":2899,"prompt_tokens":1008,"completion_tokens":1891,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":1785}},"tokens_in":624,"tokens_out":1891,"duration_ms":13738,"temperature":1.0,"reasoning_tokens":1785,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:37:14.995229+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A Monte Carlo closure test would settle this: run the same analysis on simulated Pb--Pb events with a known feed-down fraction and check whether the extracted prompt $\\Lambda_c^+$ $v_2$ at $p_T>4$ GeV/$c$ matches the true value within the quoted uncertainty. An independent decay-length-based tagger should also reproduce the same 3.6$\\sigma$ splitting; if it does not, the BDT-based separation is the likely source.","supporting_citations":[],"review_version":1}