{"id":"8128891b-2185-4ab6-a8f1-be6ca273956b","arxiv_id":"2607.24557","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"With noisy memories, discard-oldest beats early distillation at low coherence, while deferring distillation to the deadline maximizes weighted coherent information at high coherence in one- and two-hop repeaters.","lead":"This paper compares when to distill and swap entanglement in quantum repeaters whose memories decohere over time. It shows that short-lived memories favor discarding old pairs, while longer-lived ones favor delaying distillation until the end of a deadline window.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The headline metric R_S = P_S·max{I_c(E[ρ_out]),0} evaluates coherent information of the *average* state rather than the average coherent information; several of the paper's \"never positive\" claims sit within ~0.002–0.004 fidelity of the I_c=0 threshold, where this convention is decisive.","rationale":"The reader flagged idealized timing (unit-probability swaps, no classical latency) as the weakest assumption. That is the largest external realism gap, and the authors themselves list it as future work, but it does not threaten internal correctness and the paper is explicit about it. I identify a different, more internally load-bearing soft spot: the R_S metric computes coherent information of the mean state rather than the mean coherent information, which is a systematic, strategy-dependent upward bias, and the paper's most quotable negative results (\"never reaches positive weighted coherent information\") hinge on fidelities sitting 0.002–0.004 below a sharp threshold — within the bias and sampling noise. The paper's own ceiling argument (Discard-Swap max fidelity ≈0.813 > F*≈0.811) suggests the \"never positive\" claim may not survive at smaller T/τ than tested. The core qualitative finding — defer the final distillation in the high-coherence regime — has comfortable margins (R≈0.09 vs 0.06) and analytical one-hop support with closed forms and shipped code, so I expect it to hold. Because the reader already issued CONDITIONAL and my concern adds a different condition (metric-robustness check) without overturning the verdict, the verdict should remain CONDITIONAL — i.e., UNCHANGED. Credit: closed-form appendix, public Mathematica/Python code, and honest treatment of the distribution-level subtleties in Sec. III.E/IV.B are real strengths.","tokens_in":17630,"tokens_out":3683,"duration_ms":123386,"concrete_test":"Recompute Fig. 8 from the saved per-realization outputs (code at [29]) using the per-sample metric R̃ = (1/N)Σ_j w_j·max{I_c(F_j),0} instead of P_S·max{I_c(F̄),0}, and additionally extend the Discard-Swap and D-ASAP-S-ASAP curves to T ≲ 0.1/λ with τ=100 s. If either strategy yields R̃ > 0 at short deadlines, or if the ordering of near-threshold strategies flips between the two conventions, the \"never positive weighted coherent information\" claims are artifacts of the averaging convention and the tested T-grid; if the zeros persist at 0.809-vs-0.811 margins across both metrics, the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central two-hop conclusions (Sec. IV.A, Fig. 8) rest on R_S = P_S · max{I_c(E[ρ_out|success]),0} (Eq. 8–9). For isotropic states, I_c as a function of fidelity, g(F)=1−h2(F)−(1−F)log2 3, is concave, so by Jensen I_c(E[ρ]) ≥ E[I_c(ρ)] — the metric systematically overstates the operational (hashing) rate, by an amount that grows with the width of the fidelity distribution. The width differs by strategy (Sec. III.E reports σ=0.056 for Discard-Oldest vs 0.048 for Distill-ALAP one-hop; Fig. 9 shows Discard-Swap unimodal-broad vs S-ASAP-D-ASAP narrow+bimodal), so the bias is not uniform across the strategies being ranked. This matters because the paper's sharpest claims are knife-edge relative to the I_c=0 threshold F*≈0.811: (i) D-ASAP-S-ASAP \"never reaches positive\" R because its F̄ \"settles just below F*, at ≈0.809\" (Sec. IV.A) — a 0.002 margin, likely within Monte Carlo error at 10^5 samples and certainly within the Jensen gap; (ii) the paper itself notes Discard-Swap's noiseless ceiling is F_swap(F0,F0)≈0.813 > F*, which implies that at sufficiently small T/τ the Discard-Swap average state *does* have positive I_c, so \"never positive in any regime tested\" is plausibly an artifact of the tested T-grid rather than a structural fact. The qualitative \"defer final distillation\" conclusion (R≈0.092 vs 0.061 for the best vs worst distillation strategies) has enough margin to survive, but the zero-vs-positive claims that distinguish the strategies most sharply are metric- and threshold-sensitive. The authors are transparent about the definition and even flag the tail-weight/support distinction in Sec. III.E, which mitigates but does not remove the issue.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper studies how finite memory coherence time shapes the timing and ordering of entanglement distillation and swapping in multiplexed repeater building blocks. For a one-hop link it derives closed-form expressions (Appendix A) for the success probability, expected output fidelity, and weighted coherent information R_S = P_S·max{I_c(E[ρ_out|success]),0} of three strategies (Distill-ASAP, Distill-ALAP, Discard-Oldest) under exponential heralded arrivals and depolarizing memory noise, finding that Discard-Oldest wins at short coherence times while Distill-ALAP wins on fidelity and, over most of the deadline range, on R_S at long coherence times. For a two-hop chain it compares seven strategies by Monte Carlo (10^5 samples), finding that strategies deferring the final distillation to the deadline T achieve the highest R_S (D-ALAP-S-ALAP at short T, S-ASAP-D-ALAP at longer T), while Discard-Swap and D-ASAP-S-ASAP never reach positive R_S in any regime tested.","tokens_in":18102,"tokens_out":2465,"duration_ms":86681,"significance":"If the results hold, the paper provides a clean, well-isolated answer to a practically relevant scheduling question — when to distill and when to swap under memory decoherence — at the level of the elementary links from which larger hierarchies are built. Notable strengths: the one-hop analysis is fully analytical with closed forms in Appendix A; the policies studied are explicitly causal; the fidelity-distribution discussion (Secs. III.E, IV.B) is careful, including the observation that distillation tails exceed F_0 while Discard-Oldest is hard-bounded at F_0; and the evaluation code (Mathematica notebooks plus Python Monte Carlo) is publicly available [29], making the numerical claims reproducible. The robust qualitative conclusion — defer the final distillation to the consumption time — has comfortable margins (R≈0.092 vs 0.061 across distillation strategies) and is a useful link-level design principle. The weaker points are concentrated in the sharpest zero-vs-positive two-hop claims, as detailed below.","major_comments":[{"comment":"Sec. II.C, Eqs. (8)–(9): the headline metric evaluates the coherent information of the *average* output state, I_c(E[ρ]), but justifies it operationally via the hashing bound [28]. The hashing rate for an ensemble of output states is E[I_c(ρ)], not I_c(E[ρ]). For isotropic states I_c(F) = 1−h2(F)−(1−F)log2 3 is concave in F, so by Jensen I_c(E[ρ]) ≥ E[I_c(ρ)] — the metric systematically overstates the operational rate, by an amount that grows with the width of the fidelity distribution. The width differs across the strategies being ranked (Sec. III.E: σ=0.056 Discard-Oldest vs 0.048 Distill-ALAP; Fig. 9: broad unimodal vs narrow bimodal), so the bias is not uniform. The metric is defensible as a figure of merit in its own right, but the hashing justification is incorrect as stated and should be corrected; ideally the paper should also report P_S·E[max{I_c,0}] alongside, at least for the","section":"Sec. II.C, Eqs. (8)-(9)"},{"comment":"Sec. IV.A, Fig. 8: the claim that D-ASAP-S-ASAP 'never reaches positive weighted coherent information' rests on its conditional fidelity settling at F̄≈0.809, i.e. 0.002 below the threshold F*≈0.811. With 10^5 Monte Carlo samples, the standard error on a weighted mean fidelity is plausibly of this order, and no error bars or confidence intervals are reported anywhere in Figs. 6–8. A knife-edge claim of this kind requires a stated uncertainty budget (standard errors on F̄ and R_S, ideally per data point). As written, the zero-vs-positive distinction for this strategy is not established by the evidence shown.","section":"Sec. IV.A, Fig. 8"},{"comment":"Sec. IV.A and abstract: 'Discard-Oldest-then-Swap never reaches positive weighted coherent information in any regime tested' is plausibly an artifact of the tested (λ, τ, T) grid rather than structural. The paper itself notes the noiseless ceiling F_swap(F0,F0)≈0.813 > F*≈0.811, so for T/τ small enough that both newest pairs arrive nearly fresh (e.g. λ large relative to 1/T and T≪τ), the Discard-Swap average state must have positive I_c. The tested grids (λ=10/s with T up to tens of seconds; distributions at λ=1/s) do not probe this corner. The authors should either extend the grid to locate the crossover, or qualify the claim as grid-dependent in the abstract, Sec. IV.A, and Sec. V.","section":"Sec. IV.A"},{"comment":"Secs. II.A and V: swaps are assumed deterministic and classical-communication latency is set to zero, so all operation times are governed by arrivals and the deadline T. This is acknowledged as a simplification, but it is load-bearing for the ASAP-vs-ALAP rankings: with probabilistic swaps (e.g. 1/2 for linear optics) or classical round-trips comparable to τ, the benefit of deferring operations to T must be weighed against heralding/confirmation delays, and the identity of the leading strategy can change. The paper would be substantially strengthened by a sensitivity estimate — even a back-of-envelope rescaling or one Monte Carlo rerun at swap success probability 1/2 — showing whether the 'defer final distillation' conclusion survives.","section":"Secs. II.A, IV, V"}],"minor_comments":[{"comment":"Figs. 6–8 are referenced in the text but, in the version reviewed, the panels lack visible legends/labels in the reproduced figures; ensure each two-hop figure clearly identifies all seven strategies (the color/style mapping is only inferable from the caption).","section":"Figs. 6-8"},{"comment":"Equation referencing style is inconsistent: 'Eq. 9' vs '(9)' vs 'Eq. (8)'; please standardize.","section":"Throughout"},{"comment":"Sec. III.D, Eq. (23): the limiting success probability P≈P_dist(F0,F0) is stated for λτ≫1, but the comparison is made at λ=10/s, τ=100 s (λτ=1000) — fine — while Fig. 3 shows the same plateau structure at other τ; it would help to state explicitly which regimes Eq. (23) is meant to describe.","section":"Sec. III.D"},{"comment":"Sec. II.A: the memory-noise expression ρ(∆t)=[N⊗N](ρ0) is given inline in a long sentence; consider displaying it as a numbered equation since it underlies Eq. (2) and all subsequent dynamics.","section":"Sec. II.A"},{"comment":"Sec. IV: state the Monte Carlo sample count for Figs. 6–8 explicitly (10^5 is stated only for the distribution plots in Sec. IV.B), and clarify whether arrival times are drawn per segment independently with exactly two parallel sources, as implied by Sec. IV's opening paragraph.","section":"Sec. IV"},{"comment":"Abstract/Sec. V: 'over most of the deadline range' for one-hop R_S should point to Fig. 4 or specify the range; as stated it is vague.","section":"Abstract, Sec. V"},{"comment":"Reference [29]: 'accessed: 2026' gives no date; please include a full access date and a commit hash or version tag for reproducibility.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is competent, well-scoped, and ships reproducible code; the analytical one-hop portion is solid. My hesitation is concentrated in the two-hop headline claims: the R_S metric's hashing justification is formally incorrect (Jensen gap), and the two most quotable results — D-ASAP-S-ASAP at 0.002 below threshold and Discard-Swap 'never positive' — are knife-edge relative to the stated simulation budget and the tested parameter grid. These are fixable within the paper's scope (report both metric conventions, add uncertainty estimates, extend the grid), but they touch the central claims, hence major rather than minor revision. The robust 'defer final distillation' message will survive any reasonable resolution."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful bit here is a head-to-head of causal ASAP/ALAP distill-and-swap schedules on multiplexed 1–2 hop links under depolarizing memory noise, with closed forms for the one-hop trio and Monte Carlo for seven two-hop policies. Prior cutoff and distillation-timing work is adjacent; what is new is the systematic ordering comparison plus the discard baseline under a shared deadline T.\n\nThey do the one-hop analysis properly. Exponential arrivals, standard Werner/CNOT maps, expected fidelity, success probability, and weighted coherent information are derived cleanly; Appendix A has the closed forms, and the fidelity distributions in III.E line up with the analytics. Code is pointed to. Two-hop is simulation-only, which is honest given the state space. The qualitative takeaway that survives is that deferring the final distillation to T tends to win on R_S in the high-coherence regime, with D-ALAP-S-ALAP and S-ASAP-D-ALAP trading the lead by deadline, while pure discard-then-swap is weak on quality.\n\nSoft spots, in proportion. Swaps are unit-probability and classical latency is ignored; the authors flag both. That can reorder ASAP vs ALAP when round-trips are comparable to τ or when linear-optics swaps sit at 1/2. More subtle: R_S uses I_c of the average state, not the average of I_c. For isotropic states g(F) is concave, so Jensen overstates the hashing rate, and the bias is strategy-dependent because fidelity widths differ. Several sharp claims sit within ~0.002 of F*≈0.811 (D-ASAP-S-ASAP “never positive” at F̄≈0.809; Discard-Swap’s noiseless ceiling ≈0.813). Those zero-vs-positive distinctions are metric- and grid-sensitive. The broader “defer final distillation” ranking has more margin and is still the part I would keep.\n\nThis is for people building near-term repeater control policies, not for asymptotic rate theory. Math and citations look solid; self-cites are background, not load-bearing for the rankings. I would send it to referees. Engage if you care about schedule design under memory noise; treat the knife-edge “never positive” lines as provisional until someone checks E[I_c] and probabilistic swaps.","headline":"Solid systems paper with clean one-hop analytics and useful two-hop schedule rankings; the “never positive R” claims sit too close to the I_c threshold to take at face value.","tokens_in":19094,"tokens_out":604,"would_cite":true,"duration_ms":15757,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"When quantum memories are noisy, delaying distillation until the end usually beats purifying early or just discarding old pairs.","keywords":["quantum repeaters","entanglement distillation","entanglement swapping","noisy quantum memories","operation scheduling","coherent information","multiplexed links"],"falsifier":"Re-run the same one-hop analytics and two-hop Monte Carlo with swap success probability 1/2 (linear optics) or with classical round-trip times comparable to the memory coherence time; if late distillation no longer leads on weighted coherent information, the central timing claim fails.","tokens_in":18759,"feed_emoji":"⚛️","tokens_out":943,"duration_ms":16839,"temperature":0.7,"pith_summary":"Near-term quantum repeaters must create long-distance entanglement with memories that decohere. That forces a timing choice: purify pairs as soon as they arrive, wait until a deadline, or simply throw away the older pair. This paper works out those tradeoffs on the smallest building blocks—a single link and a two-hop chain. On one hop, short memory times favor discarding the older pair; longer coherence favors waiting to distill at the deadline, which yields higher output quality and more usable entanglement even though success is less frequent. On two hops the same late-distillation idea wins on the paper’s main figure of merit, while a discard-then-swap baseline never produces positive weighted coherent information in the regimes tested. The point is practical: decoherence changes which schedule is best, and the link-level rules found here are meant to guide larger repeater designs.","feed_headline":"Delay distillation when memories last; discard when they don't","feed_subtitle":"Noisy-memory repeaters: late purification wins most often; discard-then-swap never yields usable entanglement","key_machinery":"Weighted coherent information R_S = P_S · max{I_c(E[ρ_out|success]), 0}, which multiplies success probability by the coherent information of the average successful output state. This single score balances how often a schedule works against how much distillable entanglement the output carries, and is the quantity used to rank ASAP, ALAP, and discard policies.","core_discovery":"Memory decoherence reshapes the optimal timing of distillation and swapping. In the low-coherence regime, discarding the older entangled state gives higher expected output fidelity than distilling early or late. In the high-coherence regime, delaying distillation to the end of the time window gives the highest expected fidelity and, over most deadlines, the highest weighted coherent information, at the cost of lower success probability. On two-hop chains the best weighted coherent information likewise comes from strategies that defer the final distillation, with Distill-ALAP-then-Swap-ALAP and Swap-ASAP-then-Distill-ALAP leading at different operating points, while Discard-Oldest-then-Swap n","pith_inferences":["Adding realistic classical latency will likely shrink the region where pure ALAP wins, because waiting to the deadline then costs an extra round-trip of decoherence.","If more than two pairs per link are stored, the same late-versus-early tension will reappear in how pairs are matched for multi-copy distillation.","Combining these timing rules with existing memory-cutoff policies is the natural next control layer for multiplexed repeaters."],"forward_implications":["Link-level schedules should default to late distillation when memories are relatively long-lived, and to discarding the older pair when coherence is short.","On two-hop segments, Discard-Oldest-then-Swap is a poor default: it never yields positive weighted coherent information in the tested regimes.","Which of Distill-ALAP-then-Swap-ALAP versus Swap-ASAP-then-Distill-ALAP wins depends on the deadline and generation rate, so operating point must be checked rather than assumed.","The same ASAP/ALAP ordering principles are intended to compose hierarchically into longer repeater chains."],"fun_headline_variants":["Low coherence: discard beats distill; high coherence: delay wins","Defer distillation for best weighted coherent info on noisy memories","Discard-oldest never yields positive coherent info on two-hop chains","ALAP distillation leads when memories hold; ASAP discard when they fade","Memory decoherence flips optimal distill-vs-discard timing in repeaters"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Swaps are assumed to always succeed and classical communication delays are ignored, so operation times depend only on when pairs arrive and on a fixed deadline.","fun_headline_variants_meta":{"raw":{"variants":["Low coherence: discard beats distill; high coherence: delay wins","Defer distillation for best weighted coherent info on noisy memories","Discard-oldest never yields positive coherent info on two-hop chains","ALAP distillation leads when memories hold; ASAP discard when they fade","Memory decoherence flips optimal distill-vs-discard timing in repeaters"]},"model":"grok-4.5","effort":"low","cost_usd":0.001995,"raw_usage":{"total_tokens":1033,"prompt_tokens":942,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":19948000,"prompt_tokens_details":{"text_tokens":942,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":942,"tokens_out":71,"duration_ms":2794,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T11:42:12.487416+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same one-hop analytics and two-hop Monte Carlo with swap success probability 1/2 (linear optics) or with classical round-trip times comparable to the memory coherence time; if late distillation no longer leads on weighted coherent information, the central timing claim fails.","supporting_citations":[],"review_version":1}