{"id":"f0a81541-83ea-4671-b683-34319f7456a4","arxiv_id":"2411.14076","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper introduces a boson sampling validation protocol that distinguishes quantum sampling from classical look-alikes by fitting polynomial curves to the connectivity growth of a network built from the observed samples.","lead":"This paper proposes a new way to check that a boson sampling machine really samples from the correct quantum distribution: build a network from the collected outcomes and watch how its connectivity grows as more samples arrive. If it works, quantum advantage experiments get a simple, fast extra validation test that avoids the heavy permanents computation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The scale tests in Fig. 6 do not separate ideal boson sampling from the rejection and Monte-Carlo classical samplers of [40], so the abstract's claim of distinguishing 'classically simulated cases' is broader than the demonstrated evidence.","rationale":"The reader's verdict is CONDITIONAL, and my stress-test confirms that conditions are needed; I therefore recommend no change to the verdict label. However, the load-bearing concern I identify is not the same as the reader's weakest assumption. The reader focused on the empirical fitting form of Eq. (4), stability with respect to R, and the no-repetition assumption. Those are legitimate internal-validity concerns about whether cluster separation might be an artifact of analysis choices. My concern is about external validity of the central claim: even taking the fitted coefficients at face value, the paper's own Figure 6 shows that ideal boson sampling and the classical rejection/Monte-Carlo samplers from [40] are not separated. Since [40] is a paper demonstrating classical algorithms that can spoof boson sampling, a validation protocol that cannot reject those algorithms does not justify the abstract's unqualified claim. This is not an ad hominem or an accusation of inconsistency; the authors are honest about the limitation in the Discussion. It is, however, the most direct threat to the headline claim, because it concerns what the protocol can actually distinguish at the scale where validation matters. The concrete test I propose would settle this by quantifying the overlap between the ideal and approximate-sampler clouds. If the overlap is substantial, the claim should be narrowed or the protocol augmented with additional features; if separation is actually present but not obvious from the figure, the claim would be strengthened. Either way, the CONDITIONAL verdict remains appropriate until this is checked.","tokens_in":10173,"tokens_out":5610,"duration_ms":59726,"concrete_test":"Reproduce Figure 6 using the published GitHub code and the Bristol data from [40], then compute a quantitative separation measure (e.g., pairwise Bhattacharyya distance or cross-validated classifier error) between the fitted (alpha_mu, beta_sigma) clouds for ideal brute-force boson sampling versus each of the rejection and Monte-Carlo classical samplers at n=20, m=400, R=36. If those clouds overlap or yield classification error well above zero, the central claim must be narrowed to 'distinguishable from distinguishable-particle sampling, and from uniform/mean-field at small n' rather than 'distinguished from classically simulated cases.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of the abstract is that boson sampling filling behavior can be 'computationally efficiently distinguished from classically simulated cases.' The strongest scale evidence, Section 4 and Figure 6, only shows separation between ideal boson sampling and distinguishable-particle sampling at n=7,12,20. The same figure also includes the rejection sampler and two Monte-Carlo samplers from the Bristol dataset [40]; the text explicitly states that these are 'still much closer to the real boson sampling than to a set-up with distinguishable particles.' Those rejection/Monte-Carlo samplers are exactly classically simulated algorithms designed to spoof boson sampling, and the whole point of [40] is that such algorithms can mimic ideal boson sampling. A validation protocol that cannot reject these classically simulated spoofers does not support the broad claim of distinguishing from classically simulated cases. The paper is transparent about this in the Discussion, where the claim is narrowed to distinguishing distinguishable-particle sampling, but the abstract and the stated validation protocol are not so limited. A secondary gap is that the mean-field sampler, the hardest standard mock-up, is only tested at n=5 and is absent from the larger n=7,12,20 tests. The reader's concerns about the fitted form of Eq. (4), manual optimization of R, and the no-repetition assumption are real, but the approximate-classical-sampler gap is more load-bearing because it is visible in the paper's own main result and directly undermines the headline claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new boson sampling validation protocol based on the sample-space filling behavior of wave-function networks. For a set of N samples, the authors construct a graph with edges activated when the L1 distance between two samples is below a radius R, then examine the mean μ and standard deviation σ of the degree distribution as functions of N. They assert that in the accessible regime the filling curves obey μ = α_μ N and σ = α_σ N + β_σ N^2 (Eq. 4), and that the fitted coefficients (α_μ, α_σ, β_σ) are intrinsic fingerprints of the sampling procedure. The protocol is demonstrated at n=5, m=25 for boson sampling, distinguishable-particle sampling, and mean-field sampling, with error ellipsoids reported as non-overlapping. Using the Bristol dataset of Ref. [40], the authors test the protocol for n=7, 12, and 20 photons in 49-, 144-, and 400-mode interferometers, reporting separation between ideal boson sampling and distinguishable-particle sampling. The abstract claims that boson sampling filling behavior 'can be computationally efficiently distinguished from classically simulated cases.'","tokens_in":10491,"tokens_out":2500,"duration_ms":25762,"significance":"If the central claim holds, the approach would provide a simple, permanent-free, sample-efficient validation tool that can complement existing protocols for photonic boson sampling experiments. The paper's strengths include the use of publicly available experimental data from [40], open-source code, and a concrete demonstration that a network-based statistic separates distinguishable-particle sampling from boson sampling at moderately large scales (up to n=20, m=400). The method's computational cost is stated to be dominated by pairwise distances, with potential improvements via nearest-neighbor search. The significance is conditional, however, because the demonstrated separation does not cover the full range of classical mock-up distributions claimed in the abstract, and because the fitted polynomial form and the manual choice of R are not yet backed by a derivation or a sensitivity analysis.","major_comments":[{"comment":"The abstract's claim that the filling behavior can be 'distinguished from classically simulated cases' is broader than the demonstrated evidence. In Figure 6, the only clear separation shown for n=7, 12, 20 is between ideal boson sampling and distinguishable-particle sampling; the rejection sampler and the two Monte-Carlo samplers are explicitly described in Section 4 as 'still much closer to the real boson sampling than to a set-up with distinguishable particles.' These samplers are exactly the classically simulated spoofers that a validation protocol should reject. The Discussion narrows the claim to distinguishable-particle sampling, but the abstract and the protocol description do not. Please either narrow the central claim to match the evidence or extend the experimental tests to show that the protocol can reject approximate classical samplers.","section":"Abstract and Section 4, Figure 6"},{"comment":"The low-order polynomial form (4) is asserted without derivation or justification. The claim that α_μ, α_σ, and β_σ are intrinsic to the sampling procedure and independent of N is load-bearing, but the paper's own results show that the form is not stable: in Section 4, for n≥7, α_σ vanishes within errors, so the fitted model reduces from three parameters to two. The coefficients are also fit to the same kind of data that they classify, and the errors grow substantially when averaging over different unitaries (Figure 5b). A theoretical derivation of (4), or at minimum a systematic stability analysis over fit windows, R, and unitary ensembles, is needed to support the claim that these coefficients are intrinsic fingerprints rather than artifacts of the fitting procedure.","section":"Section 3.3, Eq. (4) and Table 1"},{"comment":"The assumption that the sample set contains no repetitions (Xi≠Xj for all i≠j) is inconsistent with the n=5, m=25 demonstration. The sample space size is C(29,5)=118,755, and up to 2,000 samples are drawn. Under sampling with replacement, the probability of no collision is extremely small (approximately exp(-2000^2/(2·118755)) ≈ 10^-8). If the data were post-processed to remove repetitions, or if the samples were generated without replacement, this must be stated, and the impact on the filling curves and on the fitted parameters must be analyzed. If repetitions are instead neglected in the analysis, the approximation error needs to be quantified.","section":"Section 3.1, no-repetition assumption"},{"comment":"The activation radius R is described as an adjustable parameter that 'should be optimised to get the best result,' and the caption of Figure 6 states that R was 'optimised manually.' The reported separation is therefore shown only at a manually tuned operating point, with no sensitivity analysis. Since R directly controls the degree distribution and hence the fitted coefficients, the protocol's robustness to the choice of R (and to the fit window) must be demonstrated. Without such an analysis, the non-overlap of error ellipsoids in Figures 5 and 6 could reflect the tuning procedure rather than an intrinsic property of the sampling distributions.","section":"Section 3.1 and Figure 6 caption"},{"comment":"The mean-field sampler, which the paper itself identifies as the most stringent classical mock-up distribution (Section 2.2), is tested only at n=5, m=25 and is absent from the larger n=7, 12, 20 tests. The paper notes that [40] contains no mean-field data, but this leaves open the question of whether the protocol scales to the hardest classical case. Please either include mean-field data at larger sizes, or clearly state the absence as a limitation of the current evidence and temper the protocol's claimed applicability.","section":"Section 4"}],"minor_comments":[{"comment":"Figure 3 shows dependence of cloud position on N, but the caption does not identify which curves correspond to which sampling procedure. Please add a legend or describe the line styles and colors in the caption.","section":"Section 3.2, Figure 3"},{"comment":"The error bars in Table 1 are reported for each fitted coefficient, but the number of sampling iterations and the number of unitary realizations used for the 'single U' and 'diff U' cases are not stated. Please provide these details, as they are needed to interpret the error ellipsoids in Figure 5.","section":"Section 3.3, Table 1"},{"comment":"The text says that the coefficients 'do not depend on the number of samples received by the validation program,' but the coefficients are themselves obtained by fitting over a range of N. Please clarify that this means the coefficients are asymptotically independent of the final sample count in the fitted regime, and specify the range of N used for the fits in each case.","section":"Section 3.3, Eq. (4)"},{"comment":"The transition from three fitted parameters (α_μ, α_σ, β_σ) at n=5 to two parameters (α_μ, β_σ) at n≥7 is stated but not explained. A brief comment on why α_σ vanishes for larger systems would help the reader understand whether this is a physical effect or an artifact of the fit.","section":"Section 4"},{"comment":"The phrase 'permanent-free' is used in the introduction and discussion, but the main text says 'does not involve the calculation of permanents' (Section 3.3). For consistency, use the same term throughout, and define it at first use.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a potentially useful and simple validation tool, and the authors are transparent about many limitations in the Discussion. However, the abstract overstates the scope of the demonstrated results, and the load-bearing assumptions about the polynomial form (4), the no-repetition condition, and the manual tuning of R require either theoretical backing or systematic sensitivity analysis. The absence of mean-field data at larger sizes further limits the evidence. These issues are fixable within the manuscript's scope, and I would encourage a revision that narrows the claims and adds the missing analyses. The public code and use of the Bristol dataset are positive features that should be preserved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper has a genuinely simple idea — characterize samplers by the fitted slope and curvature of the WFN degree statistics as a function of sample count — and it demonstrates that the resulting fingerprint separates ideal boson sampling from distinguishable-particle sampling at up to n=20, m=400, using public data from Neville et al. That is a real, reproducible result, and the authors deserve credit for shipping code and being transparent about which data they used.\n\nThe trouble is the abstract. It says the filling behavior can be 'efficiently distinguished from classically simulated cases,' but the scale tests only separate boson sampling from distinguishable particles. The same Bristol data contain a rejection sampler and two Monte-Carlo samplers — exactly the kind of classical spoofers a validation protocol should reject — and Figure 6 shows those sit essentially on top of the boson sampling cluster. The authors acknowledge this in the Discussion, so the gap is not hidden, but the headline claim and the validation protocol as stated are broader than the demonstrated evidence. The mean-field sampler, the hardest mock-up, also disappears after n=5, so we don't know how the fingerprint behaves against it at scale.\n\nThere are two further soft spots that are real but more minor. Equation (4) is an empirical ansatz: we get no derivation or error analysis for the polynomial form, and the paper's own n>=7 results show the quadratic term vanishes, so the fingerprint itself changes form with system size. The activation radius R is tuned manually, and there is no sensitivity analysis, so the separation we see is at a selected operating point. Finally, the no-repetition assumption in Section 3.1 is problematic at n=5, m=25: with ~2,000 samples and only ~1.2e5 outcomes, collisions are near-certain, so either the data were filtered (which would bias the filling curve) or the assumption is silently violated.\n\nNone of this undermines the core observation. The filling-curve fingerprint is plausible and the clean separation against distinguishable particles is worth knowing about. But as it stands, the paper promises something the data don't show. A revised version that narrows the abstract, adds a sensitivity analysis for R, tests the mean-field sampler at scale, and addresses the repeated-samples policy would be a solid contribution. I'd send it to review, with clear instructions to focus on those points.","headline":"A simple, reproducible filling-curve fingerprint that separates boson sampling from distinguishable particles at scale, but the abstract overclaims against classical spoofers that the paper's own figure shows are not separated.","tokens_in":11008,"tokens_out":2645,"would_cite":false,"duration_ms":27582,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A boson-sampling validator based on how samples fill the outcome space can distinguish genuine boson sampling from classically simulatable alternatives using only measured samples and pairwise distances, with no permanent computations.","keywords":["boson sampling","validation protocol","wave function network","sample space filling","distinguishable particles","mean-field sampling","quantum advantage","degree distribution"],"falsifier":"Take the n=7, m=49 setup and sweep the activation radius R from just above 0 to large values; if the boson-sampling and distinguishable-particle clusters in the $(\\alpha_\\mu,\\beta_\\sigma)$ plane overlap for any R other than the manually chosen 8, or if $\\alpha_\\sigma$ becomes nonzero when the sample count is extended beyond 18,000, the intrinsic-fingerprint claim fails.","tokens_in":9926,"feed_emoji":"🎯","tokens_out":8278,"duration_ms":75024,"temperature":0.7,"pith_summary":"Boson-sampling experiments must convince skeptics that their outputs come from the true multi-photon distribution rather than from a classically simulatable substitute, and this paper proposes a cheap method for doing so. It builds a graph, called a wave function network, from the measured samples and tracks how the mean and spread of the sample degree distribution grow as more samples are collected. The paper claims these growth curves have a fixed low-order polynomial shape and that their fitted coefficients are fingerprints of the sampling mechanism. For a 5-photon, 25-mode circuit the fitted error ellipsoids for boson sampling, distinguishable-particle, and mean-field samplers do not overlap; for 7, 12, and 20 photons in 49, 144, and 400 modes, boson sampling separates from distinguishable particles on public data. If the claim is right, experimenters can validate large boson-sampling runs without computing permanents and without huge sample counts.","feed_headline":"Filling curves tell true boson sampling from classical fakes","feed_subtitle":"Network degree growth separates quantum output from distinguishable and mean-field samplers using measured samples only.","key_machinery":"The wave function network is a graph whose vertices are the collected samples $X_i$, with an edge between two samples when their L1 distance $d(X_i,X_j)=\\sum_p |n^{(i)}_p-n^{(j)}_p|$ is below an activation radius $R$. The method measures the average degree $\\mu$ and the standard deviation $\\sigma$ of the degree distribution of this graph as the number of samples $N$ grows, fits the two growth curves with the polynomials of Eq. (4), and uses the fitted coefficients as the discriminating features. The practical load is just the pairwise distance computation between samples, with no permanents and no trained classifier; this is why the method can be applied at 20 photons in 400 modes.","core_discovery":"The central claim is that the filling of sample space by boson-sampling outputs carries an intrinsic signature of the wave function, visible in the growth of a network built from samples. With the L1 metric and a fixed activation distance $R$, the mean degree of the network grows as $\\langle\\mu\\rangle(N)=\\alpha_\\mu N$, and the standard deviation of the degree grows as $\\langle\\sigma\\rangle(N)=\\alpha_\\sigma N+\\beta_\\sigma N^2$ over the sample numbers available in practice. The coefficients $\\alpha_\\mu,\\alpha_\\sigma,\\beta_\\sigma$ are proposed as snapshot-count-independent fingerprints of the sampler inside the black box. The paper shows that for $n=5,m=25$ these coefficients form non-overlapping error ellipsoids separating boson sampling from distinguishable-particle and mean-field sampling, both for a fixed unitary and when averaged over unitaries; the uniform case is noted as trivially excluded by single-particle observables. Using the data in [40], it further shows that with two surviving coefficients $\\alpha_\\mu$ and $\\beta_\\sigma$, boson sampling separates from distinguishable particles for $n=7,12,20$ photons in $m=49,144,400$ modes, within a collision-free subspace. The authors conclude that the approach is an efficient, permanent-free validation protocol for current and near-term experiments.","pith_inferences":["A natural next test, not reported in the paper, is to measure the same coefficients on Gaussian boson sampling data, where the same network construction applies and a fingerprint might generalize.","The method's reliance on a manually chosen $R$ could be replaced by an automatic rule that picks $R$ to maximize the separation between the fitted ellipsoids; that would make the protocol parameter-free and testable on unknown black boxes.","Because the coefficients are fitted from growth curves, one could look for scale-invariant combinations such as $\\beta_\\sigma/\\alpha_\\mu^2$ that might transfer across system sizes, giving experimenters a calibrated decision threshold rather than relative cluster separation."],"forward_implications":["The protocol can reject distinguishable-particle and mean-field hypotheses without evaluating permanents, making it a practical pre-screening test for experimental data.","Because it uses fitted coefficients rather than raw cloud positions, validation does not require matching the exact number of collected samples between different black boxes.","The method remains usable when data are restricted to a collision-free subspace, which is a harder case for validation and the natural regime for large photonic experiments.","Results on 20 photons in 400 modes suggest the test can be applied at sizes where direct distribution verification is impossible."],"supporting_citations":[{"why":"Introduces wave function networks and snapshot-based analysis; the method's construction rests on it.","marker":"[36]"},{"why":"Earlier application of wave function networks to boson sampling showing that degree clouds can separate cases; the paper extends this to fitted filling curves.","marker":"[35]"},{"why":"Public sample data at n=7,12,20 and m=49,144,400 used to test the protocol on larger circuits.","marker":"[40]"},{"why":"Defines the mean-field sampler, one of the three classically simulatable distributions the protocol must reject.","marker":"[38]"},{"why":"Defines uniform and distinguishable-particle mock-up distributions and the validation task.","marker":"[23]"},{"why":"Specifies that a useful validation protocol should require polynomial resources; the paper's efficiency claim is measured against this requirement.","marker":"[32]"}],"fun_headline_variants":["Filling curves unmask false boson samplers","Sample-space filling validates true quantum advantage","Network degree growth exposes classical samplers in boson tests","Permanent-free validation via filling behavior","Filling fingerprint distinguishes boson sampling from fakes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's separation depends on the asserted polynomial forms in Eq. (4), with coefficients that are stable across sample count N, activation radius R, and interferometer realization; if any of these stabilities breaks, the distinguishable error ellipsoids may reflect analysis choices rather than an intrinsic property of the sampler.","fun_headline_variants_meta":{"raw":{"variants":["Filling curves unmask false boson samplers","Sample-space filling validates true quantum advantage","Network degree growth exposes classical samplers in boson tests","Permanent-free validation via filling behavior","Filling fingerprint distinguishes boson sampling from fakes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1465,"prompt_tokens":1003,"completion_tokens":462,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":389}},"tokens_in":619,"tokens_out":462,"duration_ms":5038,"temperature":1.0,"reasoning_tokens":389,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:36:06.182095+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the n=7, m=49 setup and sweep the activation radius R from just above 0 to large values; if the boson-sampling and distinguishable-particle clusters in the $(\\alpha_\\mu,\\beta_\\sigma)$ plane overlap for any R other than the manually chosen 8, or if $\\alpha_\\sigma$ becomes nonzero when the sample count is extended beyond 18,000, the intrinsic-fingerprint claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces wave function networks and snapshot-based analysis; the method's construction rests on it."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier application of wave function networks to boson sampling showing that degree clouds can separate cases; the paper extends this to fitted filling curves."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Public sample data at n=7,12,20 and m=49,144,400 used to test the protocol on larger circuits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the mean-field sampler, one of the three classically simulatable distributions the protocol must reject."},{"cited_title":"Comput.14 1383–1423","cited_arxiv_id":null,"evidence_quote":"Defines uniform and distinguishable-particle mock-up distributions and the validation task."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Specifies that a useful validation protocol should require polynomial resources; the paper's efficiency claim is measured against this requirement."}],"review_version":1}