{"id":"23c6c8b5-de8d-4caf-83ee-ea8ce8fc6258","arxiv_id":"2506.20514","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A mode-selective Raman quantum memory in warm cesium vapor estimates the separation between two spectral lines with up to a 34-fold precision enhancement over direct intensity detection, resolving separations down to 1/20 of the linewidth.","lead":"This paper demonstrates a quantum memory made of warm cesium vapor that can pick out specific light patterns, or modes, with very low crosstalk, and uses it to measure the gap between two overlapping spectral lines far more precisely than standard spectroscopy. The technique reaches a precision about 34 times better than direct intensity measurement at separations as small as one twentieth of a linewidth, while also storing and retrieving the light on demand.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 34-fold enhancement may be inflated because the crosstalk parameters are calibrated on the same click dataset later bootstrapped for the MSE; the paper does not state that calibration counts were excluded.","rationale":"The paper is a careful experimental demonstration with a plausible physical mechanism, honest uncertainty accounting, and a clear acknowledgment that finite-statistics bias can make MSE fall below the CRLB. The central claim of a 34-fold precision enhancement at ϵ = 0.05 is internally consistent with the quoted crosstalk and with the theoretical F/FDI curve at that separation, so the headline number is not obviously an artifact of the biased-MSE-versus-unbiased-CRLB comparison. The load-bearing vulnerability is instead the data-handling pipeline: the crosstalk matrix used by the MLE is calibrated on 1.6×10^5 counts per separation, and the MSE is later bootstrapped from a full click dataset of only about 2×10^5 counts per separation, with no explicit statement that calibration and evaluation sets are disjoint. Because the MLE's truncation boundary depends directly on the fitted leakage parameter 1−α, even modest sampling fluctuations shared between calibration and evaluation can bias the reported MSE downward. This is a concrete, testable concern rather than a speculation about intent; if the calibration counts were already excluded, the concern dissolves. The proposed data-splitting check would settle it directly. The reader's weakest assumption identified the same issue as the primary risk, so the conditional verdict remains appropriate without modification.","tokens_in":22123,"tokens_out":11504,"duration_ms":132382,"concrete_test":"Recompute the MSE at ϵ = 0.05 and 0.1 with strict data splitting: partition each separation's recorded clicks into a calibration subset and a disjoint evaluation subset before any bootstrapping, for example fitting α and β on the first 8×10^4 counts per separation and evaluating on the remaining 1.2×10^5 counts, or using a separate calibration dataset if one exists. If the out-of-sample MSE is higher than the reported value by more than the stated uncertainty, the headline enhancement is partly a training-on-test artifact. As a minimal check, the authors should state the overlap between the 1.6×10^5 calibration counts and the resampled evaluation sets; if the overlap is nonzero for any bootstrap draw, repeat the analysis with the overlap removed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The evaluation pipeline for the headline enhancement is circular unless the calibration counts are excluded. In Methods (System calibration), the crosstalk parameters α and β of matrix M are fitted by least squares to 1.6×10^5 detected counts per separation. In Methods (Data collection and analysis), the MSE is then computed by resampling N = 2×10^3, 10^4, or 10^5 counts from the full click dataset for 50 bootstrap replicates. The paper never states that the 1.6×10^5 calibration counts were removed from that dataset. Since the total recorded counts per separation are only about 2×10^5, an evaluation sample of 10^5 counts must overlap the calibration data almost completely if the counts are not excluded. The MLE in Eq. S6 is highly sensitive to the fitted leakage 1−α through the truncation boundary N1 ≥ (1−α)N, so fitting α on the very clicks later used for evaluation can let sampling fluctuations in those clicks lower the apparent MSE. This would bias the reported (34±4)-fold enhancement at ϵ = 0.05 and the PER = 4.4 dB sensitivity claim. A separate concern about comparing a biased MSE to the unbiased DI CRLB does not appear to drive the headline number: at N = 10^5 and ϵ = 0.05 the measured MSE×N is close to the mode-filtering CRLB, so the enhancement is not an artifact of estimator bias at that operating point.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an experimental demonstration of frequency-domain superresolution using a warm-cesium Raman quantum memory as a coherent temporal-mode filter. The signal is prepared as an incoherent mixture of two equal-intensity Gaussian spectral lines of width σ = 5.30 MHz, the memory is used to project onto HG0 and HG1 temporal modes with measured crosstalk as low as 0.34%, and maximum-likelihood estimation from the two-mode count statistics is used to estimate the normalized separation ϵ. For 10^5 detected photons the authors report MSE×N below the direct-intensity CRLB, PER = 4.4 ± 0.5 dB at ϵ = 0.05 (corresponding to 265 kHz), and a (34 ± 4)-fold precision enhancement over direct detection at ϵ = 0.05. The paper also benchmarks an asymptotic superresolution parameter s ≈ 37 against other platforms and demonstrates robustness over storage times of 150–250 ns and with HG0/HG1 retrieval modes.","tokens_in":22365,"tokens_out":9325,"duration_ms":93736,"significance":"If the reported numbers survive a clean calibration/evaluation split, this is a valuable experimental advance: it demonstrates superresolved frequency-separation estimation in the MHz–GHz band with a memory platform that also offers on-demand readout, buffering, and user-defined mode conversion. The work includes substantial data (21 separations, four phases, two storage modes), bootstrapped error bars, an explicit MLE treatment including the bias induced by the non-negativity constraint on ϵ, and useful supplementary simulations of storage-efficiency crosstalk and control-field leakage. The comparison with prior QPG, GEM-based, and spectral-inversion results is informative, and the authors are candid about the low end-to-end efficiency and the crosstalk-efficiency trade-off. The main reservation is procedural: the headline enhancement may be evaluated on the same counts used to calibrate the crosstalk matrix, and the paper must demonstrate that this is not the case.","major_comments":[{"comment":"The evaluation pipeline for the headline MSE and the (34 ± 4)-fold enhancement is not shown to be free of training/evaluation overlap. The paper states that approximately 2×10^5 clicks are recorded per separation, that α and β are calibrated with 1.6×10^5 counts per separation, and that the MSE is then computed by bootstrapping N = 2×10^3, 10^4, or 10^5 counts 'from the full click dataset.' No statement is made that the calibration counts were excluded from the evaluation samples. For N = 10^5, any bootstrap sample drawn from the full dataset overlaps the calibration set almost completely. Because the MLE boundary condition N1 ≥ (1−α)N in Eq. S6 depends sensitively on α, fitting α on the same clicks later used for evaluation can let sampling fluctuations in those clicks bias the reported MSE downward, which would inflate the reported enhancement. Please clarify the dataset partition and, if the calibration counts were not excluded, repeat the MSE evaluation on a held-out subset and report whether the (34 ± 4)-fold enhancement and the PER = 4.4 dB value at ϵ = 0.05 survive.","section":"Methods, 'Data collection and analysis' and 'System calibration'"},{"comment":"Equation (8) is not the inverse Fourier transform of Eq. (7) as claimed. With the spectral lines centered at ω0 ± ϵσ/2 in Eq. (7), the temporal envelope should contain cos(ϵσt/2 − φ/2), not cos(ϵσt − φ/2) as printed. If the signal pulses were actually carved using Eq. (8) as written, the generated line separation would be twice the nominal value, making the experimental ϵ labels inconsistent with the model used in the MLE. The close agreement between the MLE estimates and ground truth in Fig. 2 suggests that the experiment may have used the correct form and Eq. (8) is a typographical error, but the authors should reconcile the equation with the actual pulse-carving waveform.","section":"Methods, 'Signal preparation', Eq. (8)"}],"minor_comments":[{"comment":"The axis label 'Nomalized intensity' is misspelled; it should read 'Normalized intensity.'","section":"Supplementary Figs. S3 and S4"},{"comment":"The variance term in MSE(ϵ, N) = Var(ϵ) + b(ϵ, N)^2 should be written as Var(ϵ̂); as printed it reads as the variance of the true parameter rather than the variance of the estimator.","section":"Eq. (4)"},{"comment":"Please define the vertical extent of the shaded regions (for example, ±1 standard deviation of the estimator about the true ϵ) in the caption; currently the reader must infer this from the main text.","section":"Fig. 2 caption"},{"comment":"The uncertainty on the calibrated crosstalk value 0.34% and on the fitted parameters α and β is not reported; since Eq. S6 is highly sensitive to 1−α, please provide these uncertainties or a sensitivity analysis showing that the reported MSE is robust to their variation.","section":"Methods, 'System calibration'"}],"recommendation":"major_revision","confidential_remarks":"The key issue is the potential training/evaluation overlap in the MSE analysis: the manuscript never states that the 1.6×10^5 calibration counts per separation were excluded from the later bootstrap evaluation, and at N = 10^5 the overlap would be near-total. Because the headline (34 ± 4)-fold enhancement and the PER sensitivity claim depend on this MSE, the revision needs an unambiguous dataset partition and, if necessary, a held-out reanalysis. The rest of the experimental work is substantial and the paper is otherwise suitable in scope for the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a serious read: it is the first demonstration of super-resolved frequency estimation using a warm-vapor Raman memory as a coherent mode filter, with on-demand storage and mode conversion. The bandwidth (MHz–GHz) fills a genuine gap between ultrafast quantum pulse gates and narrowband gradient-echo memories. The experimental work is careful: pulse carving with frequency-response correction, noise subtraction, multiple photon budgets, and bootstrap error bars. The measured crosstalk of 0.34% is impressive, and the MLE analysis follows the standard crosstalk-limited Fisher information framework.\n\nThe main soft spot is the one the stress-test flags: the crosstalk matrix M is calibrated on 1.6×10^5 counts per separation from the same click dataset later resampled for the bootstrapped MSE. The paper never states that the calibration counts were excluded. With total clicks around 2×10^5 per separation, a bootstrap sample of 10^5 substantially overlaps the calibration set. Since the MLE truncation boundary depends sensitively on the fitted leakage 1−α, fitting α on the very clicks used for evaluation can lower the apparent MSE and inflate the reported (34±4)-fold enhancement. This is a real methodological concern, not a nitpick. That said, the qualitative result is likely robust: at N=10^5 and ϵ=0.05 the measured MSE×N is close to the mode-filtering CRLB, so beating direct intensity detection is not a pure artifact. But the authors should be asked to clarify whether calibration data were excluded, and if not, to redo the analysis with a held-out calibration set.\n\nA second, smaller issue is that the enhancement factor divides a biased MSE by an unbiased DI CRLB. The authors acknowledge this and caution about overinterpretation at very small separations. It would be cleaner to report the estimator bias separately and compare against a bias-corrected DI benchmark.\n\nThe comparison with prior work in Fig. 4B is fair, and the self-citation to their EEVI protocol is appropriate. The theoretical superresolution parameter s≈37 is computed from the measured crosstalk—that is not circular, it is just a performance characterization.\n\nVerdict: this deserves peer review. The platform is genuinely new, the execution is careful, and the bandwidth niche is real. The calibration-evaluation overlap is a fixable but important gap that referees must see addressed. I would bring it to the group; it should spark a useful discussion about crosstalk calibration and estimator bias in mode-selective metrology.\n\nRecommendation: engage with it seriously, and require the calibration-overlap clarification before trusting the headline number.","headline":"New and well-executed platform demonstration for frequency superresolution, but the headline enhancement needs a check for calibration/evaluation overlap.","tokens_in":22994,"tokens_out":3762,"would_cite":true,"duration_ms":36053,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A warm-vapor Raman quantum memory filters temporal modes to estimate the separation between two spectral lines with a 34-fold precision gain over direct intensity measurement.","keywords":["frequency superresolution","Raman quantum memory","mode-selective filtering","Hermite-Gaussian temporal modes","maximum likelihood estimation","Rayleigh's curse","cesium vapor","quantum metrology"],"falsifier":"Recompute the bootstrapped MSE at $\\epsilon=0.05$ after fitting $\\alpha$ and $\\beta$ on a separately acquired calibration run, never resampling those calibration clicks for the evaluation; if the $(34\\pm4)$-fold enhancement over direct intensity does not survive, the claimed superresolution advantage is not established.","tokens_in":21872,"feed_emoji":"⚛️","tokens_out":11335,"duration_ms":112955,"temperature":0.7,"pith_summary":"The paper tries to establish that a mode-selective Raman quantum memory in warm cesium vapor can perform frequency super-resolution: it stores the optimal temporal mode of a weak optical signal with high fidelity, retrieves it on demand, and uses the retrieved photon statistics to estimate the separation between two overlapping spectral lines. The central result is a measured precision enhancement of $(34 \\pm 4)$-fold over direct intensity measurement at separation $\\epsilon = 0.05$, where $\\epsilon$ is the separation divided by the linewidth, corresponding to a resolvable separation of 265 kHz for 5.30 MHz lines. If true, this would break the practical barrier set by Rayleigh's curse for MHz-to-GHz optical signals without requiring cryogenics or magnetic field gradients, while adding on-demand storage and mode conversion that earlier mode-filtering platforms lack. The paper also reports mode crosstalk as low as 0.34% and a Fisher-information superresolution parameter approaching about 37.","feed_headline":"Raman memory beats direct intensity by 34-fold","feed_subtitle":"A warm-cesium quantum memory filters optimal temporal modes, resolving 265 kHz separations in 5.30 MHz lines.","key_machinery":"The load-bearing object is the mode-selective Raman memory treated as a single-mode quantum memory: in the low-coupling regime the storage Green's function has one dominant singular value, so the write-in control pulse selects a single temporal mode that is coherently mapped to a collective spin wave and retrieved on demand by a second control pulse. Around this, the argument uses a crosstalk matrix $M$ with parameters $\\alpha$ and $\\beta$ that maps ideal HG projection probabilities to measured ones, a Fisher-information formula showing how leakage $1-\\alpha$ destroys precision at small separations, and maximum likelihood estimation with mean squared error and parameter-to-error ratio as figures of merit.","core_discovery":"On the paper's own terms, the claim is that an atomic Raman memory operating in the low-coupling regime is a coherent temporal-mode filter, and that this filter is enough to beat direct intensity spectroscopy at sub-linewidth frequency separation estimation. The signal is prepared as an incoherent mixture of two equal-intensity Gaussian lines of known width $\\sigma = 5.30$ MHz; the memory projects it onto Hermite-Gaussian temporal modes HG0 and HG1 through the shaping of the control pulse, with measured crosstalk from HG0 to HG1 of $0.34\\%$. Maximum likelihood estimation on the retrieved counts $N_0$ and $N_1$ yields estimates of the normalized separation $\\epsilon$, and over 50 bootstrapped resamplings at photon budgets from $2\\times10^3$ to $10^5$ the mean squared error falls below the direct-intensity Cramér-Rao bound, giving a $(34\\pm4)$-fold enhancement at $\\epsilon=0.05$ and a $(28\\pm6)$-fold enhancement at $\\epsilon=0.1$. The paper further shows the estimator bias is dominated by the non-negativity constraint on $\\epsilon$ and diminishes as photon number grows.","pith_inferences":["The calibration-aware caveat: the paper fits the crosstalk parameters $\\alpha$ and $\\beta$ on $1.6\\times10^5$ detected counts per separation from the same click dataset that is later resampled for the 50 bootstrapped MSE evaluations; if those calibration counts are not excluded, the MLE is effectively evaluated on its training data. An independent calibration run would test how much of the $(34\\pm","Model robustness: the signal model assumes exactly two incoherent, equal-intensity Gaussian lines of known width, so for real spectral targets with unknown widths or intensities the reported enhancement is an upper bound rather than a guaranteed operating point.","Efficiency trade-off: the paper's end-to-end efficiency is only about 0.3% because of the filtering needed to suppress the 9.2 GHz-distant control field, so the advantage per input photon is weaker than the per-detected-photon enhancement suggests.","Architectural extension: because the filtering mode is selected by the control pulse, a cascade or loop of the same memory could sort onto more than two Hermite-Gaussian modes and extend the technique from two-line separation to fuller time-frequency characterization; the paper gestures at this extension but does not demonstrate it."],"forward_implications":["Resolving power: separations down to $\\epsilon=0.05$ (about 265 kHz on a 5.30 MHz line) can be distinguished at $10^5$ detected photons, with a parameter-to-error ratio of $4.4\\pm0.5$ dB.","Precision gain: the mode-filtered estimate outperforms direct intensity measurement at small separations, with the Fisher-information ratio approaching about 37 as $\\epsilon\\to0$, a value the paper benchmarks above previously reported time-frequency superresolution platforms.","Bandwidth niche: the memory operates in the MHz-to-GHz range, filling a gap between ultrafast quantum pulse gates and gradient-echo-memory interferometry, while adding on-demand storage, retrieval, and mode conversion.","Noise diagnostics: crosstalk and control-field leakage are the dominant error sources, so reducing crosstalk through higher-efficiency storage protocols directly raises attainable precision.","Generality: programmable temporal mode filtering opens the route to multi-parameter estimation of more complex spectral features and to distributed quantum sensor networks."],"supporting_citations":[{"why":"Supplies the quantum estimation theory that coherent mode-projection measurements escape Rayleigh's curse and defines the quantum Fisher information that the HG basis is designed to saturate.","marker":"[14]"},{"why":"Provides the Green's-function and singular-value description of multimode Raman storage used to justify the single-mode, control-shaped filtering at low coupling.","marker":"[55]"},{"why":"Establishes how measurement crosstalk limits superresolution and defines the parameter-to-error-ratio sensitivity criterion used in the benchmarks.","marker":"[22]"},{"why":"Introduces the superresolution parameter and the quantum-memory temporal-imaging comparison point at tens-of-kHz bandwidth.","marker":"[43]"},{"why":"Supplies the quantum pulse gate implementation of mode-selective time-frequency measurement against which the MHz-to-GHz memory is benchmarked.","marker":"[40]"},{"why":"Points to the light-matter-interference efficiency-enhancement protocol that the discussion proposes for reducing crosstalk while raising storage efficiency.","marker":"[60]"},{"why":"Supports the detuning choice that suppresses four-wave-mixing noise, which otherwise would add crosstalk-like background counts.","marker":"[71]"},{"why":"Provides the higher-order asymptotic bias analysis used to explain the finite-statistics MLE bias at small separations.","marker":"[53]"},{"why":"Gives the biased Cramér-Rao inequality that turns the estimator bias into the reported mean-squared-error bound.","marker":"[54]"}],"fun_headline_variants":["Quantum memory boosts frequency precision 34-fold","Mode-selective memory resolves sub-linewidth features","Warm cesium memory beats intensity by 34x","Super-resolved frequency estimation via quantum memory","Atomic memory filter yields 34x precision enhancement"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported enhancement rests on the assumptions that the signal is exactly an incoherent mixture of two equal-intensity Gaussian lines of known width and that the memory's crosstalk matrix is constant and fixed by a calibration drawn from the same dataset that later seeds the bootstrapped error estimates.","fun_headline_variants_meta":{"raw":{"variants":["Quantum memory boosts frequency precision 34-fold","Mode-selective memory resolves sub-linewidth features","Warm cesium memory beats intensity by 34x","Super-resolved frequency estimation via quantum memory","Atomic memory filter yields 34x precision enhancement"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001177,"raw_usage":{"total_tokens":4865,"prompt_tokens":944,"completion_tokens":3921,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":560,"completion_tokens_details":{"reasoning_tokens":3850}},"tokens_in":560,"tokens_out":3921,"duration_ms":29945,"temperature":1.0,"reasoning_tokens":3850,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:47:21.784082+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the bootstrapped MSE at $\\epsilon=0.05$ after fitting $\\alpha$ and $\\beta$ on a separately acquired calibration run, never resampling those calibration clicks for the evaluation; if the $(34\\pm4)$-fold enhancement over direct intensity does not survive, the claimed superresolution advantage is not established.","supporting_citations":[{"cited_title":"Tsang, R","cited_arxiv_id":null,"evidence_quote":"Supplies the quantum estimation theory that coherent mode-projection measurements escape Rayleigh's curse and defines the quantum Fisher information that the HG basis is designed to saturate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Green's-function and singular-value description of multimode Raman storage used to justify the single-mode, control-shaped filtering at low coupling."},{"cited_title":"Gessner, C","cited_arxiv_id":null,"evidence_quote":"Establishes how measurement crosstalk limits superresolution and defines the parameter-to-error-ratio sensitivity criterion used in the benchmarks."},{"cited_title":"Mazelanik, A","cited_arxiv_id":null,"evidence_quote":"Introduces the superresolution parameter and the quantum-memory temporal-imaging comparison point at tens-of-kHz bandwidth."},{"cited_title":"Donohue, V","cited_arxiv_id":null,"evidence_quote":"Supplies the quantum pulse gate implementation of mode-selective time-frequency measurement against which the MHz-to-GHz memory is benchmarked."},{"cited_title":"Enhancing Quantum Memories with Light-Matter Interference","cited_arxiv_id":"2411.17365","evidence_quote":"Points to the light-matter-interference efficiency-enhancement protocol that the discussion proposes for reducing crosstalk while raising storage efficiency."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the detuning choice that suppresses four-wave-mixing noise, which otherwise would add crosstalk-like background counts."},{"cited_title":"Hervas, A","cited_arxiv_id":null,"evidence_quote":"Provides the higher-order asymptotic bias analysis used to explain the finite-statistics MLE bias at small separations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the biased Cramér-Rao inequality that turns the estimator bias into the reported mean-squared-error bound."}],"review_version":1}