{"id":"5bd03d7c-1c50-41a5-842f-ae49909e547f","arxiv_id":"1909.01650","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In high-multiplicity pp collisions at 13 TeV, muons from charm hadron decays exhibit a nonzero elliptic flow coefficient v2, while muons from bottom hadron decays are consistent with zero v2.","lead":"ATLAS measured how muons from charm and bottom quark decays move in high-multiplicity proton-proton collisions at the LHC. Muons from charm decays show a small but nonzero azimuthal anisotropy, while muons from bottom decays are consistent with zero anisotropy within uncertainties.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The template fit assumes nonflow shape is multiplicity-independent; a data-driven cross-check with an alternative nonflow subtraction is needed to secure the charm-versus-bottom conclusion.","rationale":"The paper is a careful experimental measurement and the template method is standard in the field. The reader's weakest_assumption already identified the multiplicity independence of the nonflow shape, and I agree that this is the most load-bearing assumption. It is load-bearing because the extracted charm v2 is small compared with the nonflow correlations before subtraction, so any evolution of the dijet or resonance shape with multiplicity feeds directly into the cosine ridge term. The paper's Pythia8 closure check and the quoted 15-35% systematic from varying the LM window mitigate the concern but do not eliminate it: a generator-level test is not a data-level test, and the LM-window variation probes reasonable window choices rather than proving shape independence. Still, I see no internal inconsistency, no numerical error, and no evidence of overreach beyond the acknowledged limitation. The authors are transparent about the assumption and have assigned the largest systematic to it, which is the appropriate experimental treatment. Given the state of the art and the paper's self-consistent presentation, I do not find this concern decisive enough to change the reader's ACCEPT verdict. I therefore recommend UNCHANGED, with the concrete scalar-product cross-check as a valuable strengthening step that would settle whether the concern actually lands.","tokens_in":36815,"tokens_out":12036,"duration_ms":139009,"concrete_test":"Re-analyze the 4<pT<6 GeV, 60<=Nrec_ch<120 sample using the scalar-product method: compute v2 of charm-tagged muons from the event plane reconstructed from charged hadrons in the opposite pseudorapidity hemisphere (e.g., |Delta-eta|>2.0 from the muon), keeping the same multiplicity and trigger selections. If the scalar-product v2 is consistent with the template v2 within the combined statistical and systematic uncertainties, the nonflow-shape assumption is supported; if it is significantly lower or consistent with zero, the reported charm v2 is likely contaminated by multiplicity-dependent nonflow.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the template subtraction described in the method section: the low-multiplicity (Nrec_ch<40) correlation shape is scaled by F and subtracted from high-multiplicity (60<=Nrec_ch<120) correlations, with any residual cosine modulation assigned to flow. The assumption that the shape of nonflow correlations (dijets, resonances) is independent of multiplicity is explicitly acknowledged, and varying the LM window produces the largest systematic uncertainty (15-35% in v2,2). In high-multiplicity pp events, the relative contributions and shapes of back-to-back dijets and resonance decays can change with Nrec_ch; if so, the cosine ridge term would absorb the multiplicity-dependent part of the nonflow shape and masquerade as a spurious v2. The reported charm muon v2 is only of order 0.05-0.1, so an unmodeled shape evolution at the level of a few percent could create or remove the signal. The Pythia8 closure check that v2,2 is consistent with zero is reassuring but validates the method only within the generator's nonflow model, not the actual multiplicity dependence in data. Because the bottom-muon result comes from the same subtraction chain, a nonflow bias would also alter the charm-bottom difference that drives the physics conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This Letter reports measurements of the elliptic anisotropy coefficient v2 of muons from heavy-flavor hadron decays in pp collisions at sqrt(s)=13 TeV using 150 pb^-1 of ATLAS data, with separation of charm- and bottom-decay muons via the transverse impact parameter. The analysis uses two-particle muon-hadron correlations with a pseudorapidity gap, subtracts nonflow with a template method that scales the low-multiplicity correlation shape, and divides by the previously measured charged-hadron v2 to obtain the muon v2. The main results are a decreasing inclusive heavy-flavor muon v2 with pT, a charm-muon v2 that is claimed to be significantly nonzero at low pT, and a bottom-muon v2 consistent with zero within uncertainties.","tokens_in":37081,"tokens_out":4405,"duration_ms":49950,"significance":"If the result holds, it provides a first indication that charm quarks participate in the collective elliptic flow of high-multiplicity pp collisions while bottom quarks do not over the measured pT range, giving new constraints on heavy-quark transport in the smallest collision system. The paper is careful and quantitative in several respects: the heavy-flavor separation is cross-checked against FONLL and Pythia8, the muon-hadron correlation analysis includes efficiency corrections and pileup checks, the dominant systematic from the low-multiplicity window choice is quantified, and a Pythia8 closure test shows the extracted v2 is consistent with zero in the absence of flow. The significance of the measurement is, however, tempered by the reliance on the multiplicity-independence of the nonflow shape, which is tested mainly in simulation, and by the absence of a quoted statistical significance for the nonzero charm v2 claim.","major_comments":[{"comment":"The template subtraction assumes that the shape of nonflow correlations is independent of multiplicity, so that the low-multiplicity correlation C_LM can be scaled and subtracted from the high-multiplicity correlation. The only validation reported is the Pythia8 closure check and the variation of the LM window, which gives the largest systematic uncertainty (15-35%). Neither test directly constrains a multiplicity-dependent change of dijet or resonance shapes in data. Since the reported charm-muon v2 is of order 0.05-0.1, a few-percent multiplicity evolution of the nonflow shape could create or cancel the signal. I request a data-driven cross-check (for example, an alternative nonflow-subtraction approach or a comparison using correlations in different event-shape or jet-activity classes) or a quantitative estimate of the maximal nonflow contamination under a controlled multiplicity-dependent model.","section":"Template fit method (C_templ equation)"},{"comment":"The abstract and summary claim a 'significant non-zero' v2 for muons from charm decays, but no significance is quoted anywhere in the text. Figure 4 shows only points with statistical and systematic bands; the number of standard deviations by which the charm v2 deviates from zero, and by which the charm and bottom v2 values differ, should be stated explicitly, including the correlated systematic component. Without this quantitative statement, the central claim is not fully supported.","section":"Summary paragraph and Figure 4"},{"comment":"The conclusion that bottom quarks 'do not participate in the collective behavior' is stronger than what a null measurement at this precision supports. The bottom-muon v2 is consistent with zero within sizeable uncertainties in each Nrec_ch and pT bin, but consistency with zero is not evidence of absence. The wording should be softened to 'no significant v2 is observed for bottom muons,' or an upper limit on the bottom-muon v2 should be provided.","section":"Summary paragraph and Figure 4"}],"minor_comments":[{"comment":"The CERN header contains typographical artifacts ('ORGANISA TION' and 'A TLAS') that should be corrected in the final version.","section":"Title page"},{"comment":"The figures show pT axes extending beyond the measured range, and the caption text would be clearer if the pT ranges for the multiplicity-dependence panels and the multiplicity range for the pT-dependence panels were repeated in the captions rather than only in the main text.","section":"Figure 3 and Figure 4 captions"},{"comment":"The statement that the background fraction is 'fixed in accord with the fit results in Delta(p)/p_ID' is vague; specify how the fixed value and its uncertainty are propagated into the final v2 uncertainties.","section":"Text near d0 fit"}],"recommendation":"major_revision","confidential_remarks":"The nonflow shape-independence assumption is a known challenge for all small-system flow measurements, and the authors' use of the established ATLAS template method with self-citation is standard practice rather than a cause for concern. The requested items—a quantitative significance for the charm v2 and a stronger data-driven test of the nonflow assumption—are feasible within the analysis framework and would materially strengthen the paper's central claim. I therefore recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The new thing here is real: this is the first pp measurement that separates v2 for muons from charm and bottom decays, using the d0 template and momentum-imbalance methods. That alone makes it worth a serious referee. The analysis is careful, the systematic budget is thorough, and the authors openly state the key assumption behind the template fit: nonflow shape is multiplicity-independent, tested only in simulation. I agree with the reader's overall verdict. This is a solid, incremental advance in the small-system collectivity program, not a field reorientation, and it deserves publication.\n\nWhat the paper does well: the d0 separation into charm and bottom fractions is checked against FONLL and Pythia8; the template fit is standard ATLAS methodology; the systematic variations include the low-multiplicity window choice, which is the right thing to vary. The Pythia8 closure check that v2 is zero for heavy-flavor muons is a sensible sanity check. The conclusion about bottom quarks not participating in collectivity is stated with appropriate caution given the uncertainties.\n\nWhere the soft spots are, in proportion: first, the claim of significant nonzero charm v2 is never quantified with a significance or p-value. The figures show error bars, but a reader cannot tell whether \"significant\" means 2-sigma or 5-sigma. That is a concrete, easily fixable omission. Second, the stress-test note is right that the nonflow shape-independence assumption is the load-bearing point. Varying the LM window (0-30 and 20-40) gives a 15-35% systematic, which partially addresses multiplicity dependence, but it does not test the actual shape evolution of nonflow within the HM window. A data-driven cross-check with a different nonflow estimator (e.g., a four-particle cumulant or a different reference multiplicity) would strengthen the conclusion. That is a recommendation for future work, not a fatal flaw. Third, the statement that bottom quarks \"do not participate\" in collective behavior is stronger than the data support: the measured v2 is consistent with zero within large uncertainties, but the uncertainties are large enough that a modest nonzero v2 cannot be excluded. The abstract and summary already soften this with \"consistent with zero,\" but the final sentence could be more careful.\n\nWho this is for: anyone working on heavy-flavor flow, small-system collectivity, or heavy-quark transport models. It is a useful new data point, and I would cite it if I worked in that area. For peer review: yes, send it to a serious referee. The recommended requests would be to quantify the charm v2 significance and to hedge the bottom-quark conclusion to match the sensitivity.","headline":"First pp measurement separating charm and bottom heavy-flavor muon v2: careful analysis, load-bearing nonflow assumption handled honestly but with a couple of soft spots worth a referee's attention.","tokens_in":37557,"tokens_out":1661,"would_cite":true,"duration_ms":20120,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ATLAS observes a significant nonzero elliptic-flow coefficient $v_2$ for muons from charm decays in high-multiplicity $pp$ collisions, while bottom-decay muons are consistent with zero.","keywords":["elliptic flow","azimuthal anisotropy","charm quark","bottom quark","heavy-flavor muons","proton-proton collisions","high-multiplicity events","template fit method"],"falsifier":"Repeat the extraction with the low-multiplicity baseline shifted from $N_{\\mathrm{rec}}^{\\mathrm{ch}}<40$ to $N_{\\mathrm{rec}}^{\\mathrm{ch}}<20$, and compare the resulting charm and bottom $v_2$ values; if the shift moves $v_2$ by more than the quoted uncertainties, the nonflow-shape-independence assumption is falsified.","tokens_in":36618,"feed_emoji":"🌀","tokens_out":7183,"duration_ms":74792,"temperature":0.7,"pith_summary":"This paper reports the first separation of elliptic-flow measurements for muons from charm versus bottom hadron decays in proton-proton collisions at $\\sqrt{s}=13$ TeV. In high-multiplicity events, muons from charm decays show a significant nonzero elliptic anisotropy coefficient $v_2$, while muons from bottom decays are consistent with zero over the measured transverse momentum range. If correct, this means charm quarks are dragged into the collective, geometry-driven expansion pattern in the smallest collision system, whereas the heavier bottom quarks are not in this momentum window. The result gives a new, mass-resolved handle on how quark mass shapes interactions with the medium formed in high-multiplicity $pp$ events.","feed_headline":"Charm muons flow in high-multiplicity pp; bottom muons don't.","feed_subtitle":"First separation of charm and bottom decay muons shows charm participates in collective flow, bottom does not.","key_machinery":"The central object is the template-fit decomposition of per-muon two-particle azimuthal correlation functions into a nonflow component and a flow-modulated ridge: $C_{\\mathrm{templ}}(\\Delta\\phi)=F\\,C_{\\mathrm{LM}}(\\Delta\\phi)+G\\left[1+\\sum_{n=2}^4 2v_{n,n}\\cos(n\\Delta\\phi)\\right]$. This carries the argument by attributing any multiplicity-dependent excess in the correlation function to collective elliptic flow, after subtracting the low-multiplicity baseline that is assumed to contain only nonflow. The charm-bottom separation is then carried by template fits to the muon transverse impact parameter $d_0$, and the muon $v_2$ is obtained from the pair anisotropy through flow factorization $v_n^{\\mu}=v_{n,n}/v_n^h$.","core_discovery":"The central claim is that the elliptic anisotropy $v_2$ of muons from charm-hadron decays in $pp$ collisions at $\\sqrt{s}=13$ TeV is significantly nonzero for high-multiplicity events, while $v_2$ of muons from bottom decays is consistent with zero. Using 150 pb$^{-1}$ of ATLAS data and muons with $4<p_T<7$ GeV and $|\\eta|<2.4$, the analysis separates heavy-flavor muons from light-hadron decay backgrounds through the momentum imbalance $\\Delta p/p_{\\mathrm{ID}}$, and separates charm from bottom muons through template fits to the transverse impact parameter $d_0$. The inclusive heavy-flavor muon $v_2$ shows no strong multiplicity dependence in the 60$\\le N_{\\mathrm{rec}}^{\\mathrm{ch}}<120$ range and decreases as $p_T$ rises from 4 to 7 GeV. The bottom-decay muon $v_2$ is zero within uncertainties, whereas the charm-decay muon $v_2$ is nonzero at lower $p_T$, indicating that bottom quarks do not appear to participate in the collective behavior in these smallest collision systems.","pith_inferences":["Beyond the paper, if the charm-bottom gap persists with more data, the $p_T$ at which bottom flow turns on would map the mass threshold for thermalization in small systems.","Beyond the paper, replacing the low-multiplicity template with a rapidity-separated subevent estimator would directly test whether the nonflow-shape assumption biases the reported charm $v_2$.","Beyond the paper, a fully reconstructed $D$-meson measurement in the same $pp$ dataset could corroborate the charm flow signal without relying on muon decay template ambiguities.","Beyond the paper, an analogous measurement in $p$+Pb collisions at the same muon $p_T$ would show whether the charm-bottom gap widens or narrows as the system size grows."],"forward_implications":["In high-multiplicity $pp$ collisions, charm quarks participate in the same collective elliptic flow pattern that has been observed for light hadrons, reinforcing the hydrodynamic description of the smallest collision systems.","Bottom quarks show no measurable elliptic flow in the $4$\\,--\\,$7$ GeV muon $p_T$ range, implying a mass-dependent threshold for heavy-quark thermalization in $pp$ events.","The inclusive heavy-flavor muon $v_2$ is roughly flat with multiplicity and falls with increasing $p_T$, providing a new differential constraint on heavy-quark transport models.","The measured charm-bottom gap provides the first $pp$-system data point in a regime where transport calculations predict larger $D$-meson than $B$-meson $v_2$ at low $p_T$.","The demonstrated ability to separate charm and bottom contributions via $d_0$ templates can be applied to larger datasets and other collision systems to sharpen the mass-dependence picture."],"supporting_citations":[{"why":"Supplies the two-particle correlation template-fit method and the charged-hadron $v_2$ values used as the denominator in the flow factorization.","marker":"[17]"},{"why":"Established the same template-fitting approach for long-range elliptic anisotropies in $pp$ collisions, which this analysis adapts to muon-hadron pairs.","marker":"[25]"},{"why":"Tests the multiplicity-independence of nonflow shapes, the key assumption behind subtracting the low-multiplicity baseline from high-multiplicity events.","marker":"[27]"},{"why":"Provides the flow-factorization assumption $v_n^{\\mu}=v_{n,n}/v_n^h$ used to convert muon-hadron pair anisotropy into a muon $v_2$.","marker":"[28]"},{"why":"Shows charm and strange hadrons flow in high-multiplicity $p$+Pb collisions, supplying the smaller-system comparison this measurement extends to $pp$.","marker":"[11]"},{"why":"Reports heavy-flavor decay electron anisotropy in $p$-Pb collisions, providing a cross-system reference for the heavy-flavor flow pattern.","marker":"[12]"}],"fun_headline_variants":["Charm muons flow, bottom don't in high-multiplicity pp","ATLAS finds charm-muon elliptic flow in pp; bottom zero","Charm decays show v2 in pp, bottom decays don't","In pp, charm muons have nonzero v2; bottom muons zero","Heavy-flavor flow split: charm yes, bottom no in pp collisions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis assumes the shape of the nonflow background, made of back-to-back dijets and resonance decays, does not change with event multiplicity, so that the low-multiplicity correlation function can be subtracted from the high-multiplicity one; if that shape changes, the extracted $v_2$ would be contaminated.","fun_headline_variants_meta":{"raw":{"variants":["Charm muons flow, bottom don't in high-multiplicity pp","ATLAS finds charm-muon elliptic flow in pp; bottom zero","Charm decays show v2 in pp, bottom decays don't","In pp, charm muons have nonzero v2; bottom muons zero","Heavy-flavor flow split: charm yes, bottom no in pp collisions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000306,"raw_usage":{"total_tokens":1771,"prompt_tokens":982,"completion_tokens":789,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":691}},"tokens_in":598,"tokens_out":789,"duration_ms":6866,"temperature":1.0,"reasoning_tokens":691,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:10:37.846686+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the extraction with the low-multiplicity baseline shifted from $N_{\\mathrm{rec}}^{\\mathrm{ch}}<40$ to $N_{\\mathrm{rec}}^{\\mathrm{ch}}<20$, and compare the resulting charm and bottom $v_2$ values; if the shift moves $v_2$ by more than the quoted uncertainties, the nonflow-shape-independence assumption is falsified.","supporting_citations":[],"review_version":1}