{"id":"0de9bf77-3cfb-4aca-bd7a-e9a1ea7db603","arxiv_id":"2608.10304","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A linear optical frontend improves classification at a small sensor bottleneck through coherent, nonlocal field mixing, producing quadratic features that can outperform a trained linear preprocessor.","lead":"A simulation study on MNIST shows that a linear optical frontend helps downstream classification mainly by reshaping the joint statistics of the sensor readout, measured by the Bhattacharyya distance, not by making individual pixels more separable. The paper identifies which physical resources, coherence and general nonlocal field mixing, are needed to access quadratic interference features that can beat the best trained linear preprocessor.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 93% super-linear result is demonstrated only for an unconstrained complex matrix (Sec. III.F), not for a physically constrained nonlocal device; the abstract's claim that 'such systems can yield significant performance gains' therefore exceeds the evidence.","rationale":"The reader identified the physical realizability of the general operator M as the weakest assumption, and my stress-test concurs: this is the load-bearing point for the central claim. The paper's own text is honest about the gap, labeling M as an empirical upper bound and deferring physical realization to future work. However, the abstract and parts of the discussion state the super-linear advantage as a property of 'nonlocal, coherent optical systems' without the necessary caveat. This overstatement is consequential because the entire claimed advantage over the linear ceiling is demonstrated only in the idealized operator class; a physically constrained nonlocal device might not achieve it. The proposed test—training a Maxwell-consistent or multi-plane parameterization under the same protocol—would directly settle whether the advantage is physical or an artifact. Since the reader already conditioned acceptance on this issue, no verdict change is needed; the conditional accept stands, with the condition being exactly that the super-linear advantage must be shown in a physically realizable model before it is stated as a system capability. My assessment agrees with the reader's weakest assumption rather than introducing a new objection.","tokens_in":15949,"tokens_out":4328,"duration_ms":52655,"concrete_test":"Replace the unconstrained M in Sec. III.F with a physically constrained nonlocal linear forward model, for example a trainable 3D susceptibility distribution (or a multi-plane diffractive network with a finite number of phase planes and the same 4 mm total thickness), trained under the same coherent MNIST protocol with passivity enforced (singular values ≤ 1). If the accuracy at the 2×2 sensor falls to the 87% linear ceiling or below, the super-linear claim is not supported for physical systems. As a complementary check, attempt to factor the optimized M into the cascade form of Eq. (2) (free-space propagation, a finite-thickness local phase mask, and propagation) with physically plausible parameters; if no such factorization exists, the advantage is an artifact of the unconstrained operator class.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that coherent, nonlocal optical systems can surpass the best trained linear preprocessor rests on simulations of an unconstrained complex matrix M in Sec. III.F. This operator is trained with no physical constraints beyond an optional rescaling to norm ≤ 1: it need not satisfy passivity in the stricter scattering-matrix sense, reciprocity, causality, locality of the material response, finite device thickness, or any restriction on the number of accessible spatial channels. A volumetric metamaterial, the suggested realization, is a local dielectric medium obeying Maxwell's equations; its scattering operator is not an arbitrary matrix. The paper explicitly acknowledges this gap (Sec. I.D.2: 'Interesting open questions remain regarding which specific physical properties such a device would need'), and the abstract's 'such systems can yield significant performance gains' is stronger than the demonstrated 'an idealized operator reaches 93%'. The 5–6% advantage over the trained linear ceiling (87% at 2×2) may therefore be an artifact of an operator class that includes non-physical matrices, rather than a property of any realizable optical frontend. This is not an internal inconsistency; the paper is careful to label M as an empirical upper bound. But the headline claim, as phrased, asserts a capability of physical systems, and the load-bearing support for that assertion is missing. The concern would be settled by checking whether the super-linear advantage survives when M is replaced by a physically constrained nonlocal linear model.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript studies hybrid optical-electronic classifiers in which a linear optical frontend (a phase mask, a metasurface with an added nonlocal k-space filter, or an idealized general operator) transforms a coherent or incoherent field before intensity detection and an MLP backend. The central thesis is that a well-designed frontend improves classification accuracy by reshaping the joint readout statistics, quantified by the Bhattacharyya distance, which rank-orders configurations consistently with measured accuracy (Spearman rho = 0.97 at 2x2). Because the sensor measures intensity, the readout is quadratic in the input field; the quadratic cross-terms require coherence and nonlocal field mixing. The authors report that an unconstrained complex matrix operator under coherent illumination reaches 93% at a 2x2 sensor, surpassing a trained linear preprocessor ceiling of 87%, while the physically modeled metasurface frontends reach 80% and incoherent systems 71%. They interpret this as evidence that coherent nonlocal systems can yield super-linear performance, with the gains concentrated in the small-sensor bottleneck regime.","tokens_in":16221,"tokens_out":8783,"duration_ms":99800,"significance":"If the upper-bound result holds, the paper offers a useful design principle for hybrid optical-electronic inference: optimize linear optical frontends to maximize joint Bhattacharyya separability, a correlation-aware, training-free metric. The demonstration that per-pixel separability can be unchanged while joint separability and accuracy improve (Fig. 2) is valuable and well supported. The paper also carefully distinguishes inference from imaging, and the claim that intensity detection is a computational resource for classification is physically interesting. Strengths include the nonparametric validation of the Gaussian separability estimate at 2x2 (Spearman 0.89-0.99 between Gaussian and nonparametric pairwise ranks, and 0.97 against accuracy) and the clearly specified simulation and training protocol. The principal weakness is that the headline 'super-linear' claim rests on an idealized operator class whose physical realizability is left open, which limits the strength of the conclusions as currently stated.","major_comments":[{"comment":"The claim that coherent nonlocal optical systems can surpass the best trained linear preprocessor is load-bearing but is supported only by the unconstrained complex matrix M in C^784x784 of Sec. III.F. This operator is trained without reciprocity, causality, or restrictions on the number of accessible spatial channels or the locality of the material response; the body explicitly leaves open which physical properties a realizing device would need (Sec. I.D.2). The abstract's statement that 'such systems can yield significant performance gains' is therefore stronger than the demonstrated 'an idealized operator reaches 93%.' Please either constrain M to a physically plausible class (e.g., complex symmetric/reciprocal scattering matrices, matrices generated by a finite-thickness local dielectric slab, or operators with explicit channel limitations) and show that the 5-6% advantage over the 87% linear ceiling survives, or reframe the abstract, title, and discussion so that the super-linear advantage is attributed to an idealized operator class as an empirical upper bound rather than to physical systems.","section":"Abstract; Sec. I.D.2; Sec. III.F"},{"comment":"The batch-normalization argument correctly removes the overall scale of M, so rescaling to norm <= 1 addresses passivity in the operator-norm sense. It does not, however, address structural constraints: a reciprocal passive medium would require a symmetric scattering matrix in an appropriate basis, and a physical Green's function must satisfy causality and limited transverse channel count. A concrete test would be to train M restricted to complex symmetric matrices, or to matrices produced by a slab of local dielectric material with finite thickness, and to report whether the 93% accuracy at 2x2 persists. Without such a test, the super-linear performance could be an artifact of nonphysical nonreciprocal or acausal transformations.","section":"Sec. III.F; Sec. I.D.2"},{"comment":"The manuscript's own Fig. 4 shows that the physically modeled coherent frontends, including the metasurface with and without a nonlocal k-space filter, do not surpass the trained linear ceiling; only the unconstrained operator does. This discrepancy between the modeled physical systems and the headline claim is central. The Discussion should state explicitly that no currently modeled physical frontend achieves the super-linear advantage, and should identify arbitrary nonlocal field routing as a hypothesis for future work rather than a demonstrated capability.","section":"Fig. 4; Sec. I.D.2"}],"minor_comments":[{"comment":"The term 'linear optical frontend' could mislead because the intensity readout is quadratic; please state explicitly at first use that the frontend is linear in the field, not in the measured intensity.","section":"Sec. I.A"},{"comment":"The statement that 'real and complex operators give statistically indistinguishable accuracy' should report the actual accuracies and a measure of spread across seeds, rather than only a qualitative claim.","section":"Sec. III.F"},{"comment":"For the LDA projections in Fig. 3(b), please state whether the projection was fit on the readout distribution and whether the same projection is used for all panels; the caption currently leaves this unclear.","section":"Fig. 3"},{"comment":"For a computational study of this type, a public repository containing the trained-model readouts and analysis scripts would substantially strengthen reproducibility; 'available from the corresponding authors on reasonable request' is generally insufficient for a modern physics-optics journal.","section":"Sec. III.J"},{"comment":"The statement that under incoherent illumination 'intensity kernels are non-negative' is correct, but please clarify that this refers to the intensity point-spread function (the autocorrelation of the field kernel), not to the field kernel itself.","section":"Sec. I.D.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the body is more careful than the abstract, but the central physical claim needs either additional constrained simulations or a substantial reframing. I recommend major revision rather than rejection because the gap is fixable: the authors can temper the claims or add a physically constrained operator class. The separability analysis and the distinction between inference and imaging are solid contributions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new piece here is the extension of the separability–accuracy correspondence from the two-class, shared-covariance incoherent case (Ref. [21]) to ten classes, full per-class covariances, coherent illumination, and nonlocal operators. The Bhattacharyya distance is a sensible choice, and the validation is careful: Spearman rho of 0.97 against accuracy at 2×2, plus a nonparametric kNN check that recovers the same ranking. That part of the paper is methodologically solid, even though no code or data are shipped.\n\nThe conceptual contribution is Eq. (3): a linear optical frontend followed by intensity detection produces features quadratic in the input, and the interference cross-terms require both nonlocality and coherence. This cleanly explains why a local phase mask under incoherent light only remaps intensities, and why a k-space amplitude filter adds nothing. The contrast with imaging—where non-mixing, focusing optics are optimal—is well drawn and gives the work broader significance.\n\nSoft spots, in proportion: the abstract claims that nonlocal, coherent optical systems \"can yield significant performance gains, surpassing the best trained linear preprocessor.\" That claim leans on the 93% accuracy for the unconstrained complex matrix M in Sec. III.F. The paper is honest about this being an idealized empirical upper bound, and Sec. I.D.2 explicitly says the physical properties such a device would need are open questions. But the abstract's phrasing overstates what has been demonstrated. Nothing here shows that a physical passive, reciprocal, or even a realistically constrained nonlocal device reaches the super-linear regime. The stress-test concern is real, though not fatal—the paper would be strengthened by replacing M with a physically constrained model, or at least softening the abstract.\n\nAlso minor: the experiments use MNIST with a scalar-diffraction model, and the Fashion-MNIST supplement is only a partial answer. The claims about design principles are plausible but rest on a narrow empirical base.\n\nWho is this for? Anyone working on hybrid optical–electronic inference or meta-optic encoders. It clarifies what the frontend should compute and gives a cheap, training-free objective for optimizing it. I'd send it to peer review: the core message is useful, the metric validation is sound, and the super-linear claim, if confirmed later, matters. Recommend accepting with revision, mainly tightening the abstract and adding either a physical realization or a more constrained numerical experiment to back the headline claim. This deserves a serious referee.","headline":"A clean simulation study that identifies when and why a linear optical frontend helps classification, but the headline super-linear advantage rests on an idealized operator whose physical realizability is left open.","tokens_in":16809,"tokens_out":2661,"would_cite":true,"duration_ms":27335,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A linear optical frontend can beat the best trained linear preprocessor only when it is nonlocal and coherent.","keywords":["hybrid optical-electronic inference","linear optical frontend","Bhattacharyya distance","optical nonlocality","optical coherence","intensity detection","metasurface","quadratic features"],"falsifier":"Simulate or fabricate a passive volumetric linear optical frontend whose field operator approximates the unconstrained matrix operator on the MNIST task with a 2x2 sensor. If coherent readout accuracy does not exceed the 87% trained-linear ceiling, or if a local metasurface alone already matches the general operator under coherent light, the central super-linear advantage would be refuted.","tokens_in":15690,"feed_emoji":"🔬","tokens_out":5765,"duration_ms":52847,"temperature":0.7,"pith_summary":"Hybrid inference systems that send light through an optical frontend before a digital classifier need to know what the optics should compute. The paper shows that the right goal is reshaping the class-conditional statistics of the sensor readout, and that a training-free statistical distance, the Bhattacharyya distance, measures this reshaping and predicts downstream accuracy. Because a camera records intensity, any linear field transformation produces features that are quadratic in the input field; the extra discriminative information lives in interference cross-terms that only nonlocal, coherent optics can access. The paper's central result is that such systems can surpass the best trained linear preprocessor of the same dimension, with an idealized unconstrained linear operator reaching 93% accuracy at a 2x2 sensor against an 87% linear ceiling.","feed_headline":"Coherent nonlocal optics beat the best linear preprocessor","feed_subtitle":"Field-mixing interference lifts 2x2-sensor accuracy to 93%, past the 87% linear ceiling.","key_machinery":"The central object is the sensor readout equation $I_j = \\sum_k |M_{jk}|^2 |u_k|^2 + \\sum_{k\\neq l} M_{jk} M_{jl}^* u_k u_l^*$, which shows that intensity detection makes a linear field map $M$ quadratic in the input. The diagonal sum is a linear intensity map, while the interference cross-terms carry class-discriminative information only when $M$ couples distinct input points onto one pixel (nonlocality) and the input has a deterministic relative phase (coherence). The paper pairs this with the Bhattacharyya distance $D_B = -\\ln \\int \\sqrt{p(y)q(y)}\\,dy$, a training-free separability metric whose Bayes-error bound makes it a task-relevant predictor of accuracy. The machinery works by showing that correlation-aware joint separability, not per-pixel separability, tracks downstream accuracy, with a Spearman rank correlation of $\\rho = 0.97$ across the eight 2x2 configurations.","core_discovery":"Under coherent illumination, a sufficiently general nonlocal linear operator can exceed the accuracy of any linear preprocessor because the intensity readout is quadratic in the input field. The quadratic cross-terms in Eq. (3) let the frontend encode class information in inter-pixel correlations rather than per-pixel intensities: a local metasurface raises downstream accuracy from 0.52 to 0.80 while leaving per-pixel Bhattacharyya separability essentially unchanged, with joint separability rising from 0.76 to 1.64. An unconstrained complex matrix operator under coherent light reaches 93% at a 2x2 sensor, beating the trained linear ceiling of 87%, while under incoherent light the same operator collapses to the metasurface level because the cross-terms average away. The paper therefore identifies an empirical upper bound for linear optical preprocessing and locates the physical resources needed to approach it: coherence combined with general nonlocal field mixing.","pith_inferences":[],"forward_implications":["Optical frontends only help when the sensor is a strong bottleneck; performance gains saturate for sensor sizes $N \\gtrsim 6$, since optics cannot add information, only re-encode it.","Under incoherent illumination the readout is a non-negative linear map of input intensity, so no amount of operator generality can beat the linear ceiling; the coherence gap and the field-mixing gap vanish together.","Trainable k-space amplitude filters, on either side of the mask, do not improve inference accuracy for $N \\geq 2$ because amplitude attenuation removes task-relevant photons or cross-term weight rather than adding class separation.","A per-pixel separability metric is blind to the optical advantage; only joint, correlation-aware measures such as the Bhattacharyya distance predict accuracy.","The best linear preprocessor of matched dimensionality (87% at 2x2) is a digital reference that requires signed weights, while physical intensity readouts are non-negative; coherent nonlocal field mixing reaches 93%.","","pith_inferences:","Extensions the paper leaves implicit: a training-free separability objective could replace end-to-end co-optimization, optimizing the frontend alone to maximize joint Bhattacharyya distance before training the backend; the paper suggests this route but does not test it."],"supporting_citations":[{"why":"Defines the Bhattacharyya distance used throughout as the training-free separability metric.","marker":"[19]"},{"why":"Provides the Bayes-error upper bound that makes the Bhattacharyya distance a task-relevant predictor of accuracy.","marker":"[26]"},{"why":"Establishes the proven separability-accuracy correspondence for incoherent local frontends that this paper extends to ten classes and to coherent nonlocal cases.","marker":"[21]"},{"why":"Exemplifies natural-light metasurface vision processors operating as non-negative intensity maps, the incoherent regime the paper contrasts with coherent field mixing.","marker":"[11]"},{"why":"Shows that intensity detection makes focusing or permutation optics optimal for imaging, the contrast that sharpens the mixing result for inference.","marker":"[18]"},{"why":"Shows that incoherent intensity transfer functions peak at zero spatial frequency, explaining why incoherent k-space filters cannot perform edge extraction.","marker":"[32]"},{"why":"Identifies interferometric mixing as information-optimal for coherent intensity measurements, an analogous effect to the super-linear inference regime.","marker":"[36]"},{"why":"Describes nonlocal metasurfaces and nonlocality in photonic materials, the physical platform for the general nonlocal operators considered here.","marker":"[14,15]"}],"fun_headline_variants":["Coherent nonlocal optics beat linear preprocessors","Nonlocal coherence lifts optical inference past linear limit","Quadratic field mixing: why nonlocal coherence wins","Optical frontend: coherence plus nonlocality breaks linear ceiling"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The super-linear claim assumes that an unconstrained complex matrix operator can be realized, at least approximately, by a physical passive linear optical device; the paper identifies this as an open question rather than a demonstrated device.","fun_headline_variants_meta":{"raw":{"variants":["Coherent nonlocal optics beat linear preprocessors","Nonlocal coherence lifts optical inference past linear limit","Quadratic field mixing: why nonlocal coherence wins","Optical frontend: coherence plus nonlocality breaks linear ceiling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000297,"raw_usage":{"total_tokens":1703,"prompt_tokens":910,"completion_tokens":793,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":728}},"tokens_in":526,"tokens_out":793,"duration_ms":8369,"temperature":1.0,"reasoning_tokens":728,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:29.800756+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate or fabricate a passive volumetric linear optical frontend whose field operator approximates the unconstrained matrix operator on the MNIST task with a 2x2 sensor. If coherent readout accuracy does not exceed the 87% trained-linear ceiling, or if a local metasurface alone already matches the general operator under coherent light, the central super-linear advantage would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Exemplifies natural-light metasurface vision processors operating as non-negative intensity maps, the incoherent regime the paper contrasts with coherent field mixing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that incoherent intensity transfer functions peak at zero spatial frequency, explaining why incoherent k-space filters cannot perform edge extraction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies interferometric mixing as information-optimal for coherent intensity measurements, an analogous effect to the super-linear inference regime."}],"review_version":1}