{"id":"1c3d6591-b9fa-4b74-9e7d-9fe3421b1767","arxiv_id":"2607.22867","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A split-latent variational autoencoder shows that optical scattering distributes image information for occlusion robustness and encodes focal depth in single speckle patterns.","lead":"Scattering light through a medium normally blurs images, but this paper shows it can also help: it spreads a picture's information across the camera so that masking the center still lets an AI reconstruct the original. The same speckle patterns also carry a readable signature of how deep the object sits in the scattering medium, which a specially split neural network can decode.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The focal-depth claim is confounded by an unspecified implementation: Section III.B never states how the three axial positions were realized, so the classifier may be labeling intensity, magnification, or alignment differences rather than depth.","rationale":"The reader's weakest-assumption analysis identifies exactly the same load-bearing gap, and I agree with it. The second central claim, that scattering can enable focal-depth discrimination, rests entirely on the three datasets differing only in axial position. Because the paper does not specify the mechanism that changes focal depth, the split-latent classifier's near-98% accuracy and the latent-space geometry in Tables II and III could be driven by any number of correlated image changes. This is an experimental specification issue, not a mathematical inconsistency, so it is addressable with a controlled repeat of the acquisition. I am recommending UNCHANGED rather than moving the verdict because the reader already conditioned acceptance on this assumption; the proposed control experiment would settle it. I would note in passing that the occlusion-robustness result and the use of five seeds are reasonable supporting evidence, and the missing code, data, and hyperparameters remain a reproducibility concern but are secondary to the focal-depth confound.","tokens_in":6825,"tokens_out":5548,"duration_ms":51304,"concrete_test":"Re-run the Section III.B procedure under three controlled variants in the same setup: (i) translate only the scattering medium along the optical axis, (ii) translate only the camera, and (iii) refocus by changing the SLM phase curvature, keeping all other settings fixed. Train the split-latent VAE separately on each variant. The focal-depth claim survives only if variant (i) alone reproduces the reported high-scattering accuracy near 98% and the Table II/III ordering, while variants (ii) and (iii) yield near-chance classification or clearly non-monotonic behavior. If (ii) or (iii) also produces high accuracy, the current dataset cannot separate true depth from focus or alignment confounds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing part of the central claim is the Section III.B demonstration that scattering enables focal-depth discrimination. For that claim to hold, the three datasets in Fig. 5(a) must differ only in the axial position of the projected MNIST pattern inside the scattering medium. The text never states whether the depth variation was produced by moving the medium, moving the camera, changing the SLM phase curvature, or refocusing the 4f relay. Each of those implementations changes magnification, total intensity, illumination spot position, or defocus, any of which is a perfectly legible label for the split-latent classifier. The rising silhouette scores and inter-cluster distances in Tables II and III would then reflect that confound rather than the 3D scattering path, so the claim that 'heavier scattering produces better-separated focal depth clusters' is built on an unverified data-generation assumption. The overlapping error bars in Table I also weaken the monotonic-alpha claim, but the focal-depth confound is the greater risk to the central contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether optical scattering can play a beneficial role in image reconstruction and depth discrimination. The authors generate MNIST-based speckle images under four conditions (no scattering and low/moderate/high scattering through ZnO-doped PDMS) and train VAEs to reconstruct the original digits. In the first set of experiments, center-occlusion masks of increasing radius are applied to the inputs, and the VAE-based reconstruction is scored by a pre-trained SOPCNN classifier; the authors report that scattering slows the accuracy degradation and fit the accuracy-vs-radius curves with a sigmoid (no scattering) and power laws (scattering), with an exponent alpha that is said to increase with scattering strength. In the second set of experiments, a split-latent VAE divides the latent space into a digit subspace and a focal subspace, with a classifier head on the focal subspace, and is trained on three focal-depth versions of the scattered MNIST patterns. The authors report that focal classification accuracy, silhouette scores, and inter-cluster distances in the focal latent subspace increase with scattering strength, reaching about 98% accuracy under high scattering, while digit reconstruction accuracy drops to about 85%. The central claims are that scattering distributes object information across the detector plane, enhancing robustness against pixel loss, and that scattering can enable focal-depth discrimination.","tokens_in":7061,"tokens_out":4718,"duration_ms":41412,"significance":"If substantiated, the robustness and depth-discrimination results would be a useful step toward exploiting scattering as an information-coding mechanism rather than only as a nuisance. The split-latent VAE approach is a reasonable and interpretable method for attempting to separate content from depth-related information, and the use of five random seeds with reported uncertainties in several tables is a strength. However, the paper's second central claim rests on an unspecified experimental implementation of 'focal depth' and on metrics obtained from a model that was explicitly trained with a focal classifier on the same labels. The first central claim is partially supported by the data, but the claimed monotonic trend of the fitted exponent is not statistically supported. The work is a proof-of-concept with potential impact in scattering-based sensing, but its conclusions require substantially more experimental and statistical detail before they can be accepted.","major_comments":[{"comment":"The manuscript never states how the three 'focal depths' were physically realized. The text says 'we projected the same MNIST pattern onto each of three focal depths of the scattering medium,' but the implementation is not described: whether the medium was moved axially, the SLM pattern curvature was changed, the camera or relay lens was translated, or the 4f system was refocused. Each of these implementations changes magnification, total intensity, illumination spot position, or defocus, any of which provides a legible label for the split-latent classifier. Because the central claim that scattering enables focal-depth discrimination depends on the three datasets differing only in axial position inside the scattering volume, this omission is load-bearing. The authors must specify the focal-depth variation method and provide control analyses (for example, showing that intensity, magnification, and alignment are matched across the three datasets, or that the classifier cannot separate the three conditions without scattering).","section":"Section III.B (Fig. 5(a))"},{"comment":"The claim that the fitted exponent alpha increases monotonically with scattering density (0.692→0.757) is not supported by the data. The low and moderate scattering values, 0.692±0.030 and 0.694±0.022, are statistically indistinguishable; only the high-scattering value (0.757±0.047) is clearly larger. The sentence 'alpha increases monotonically with scattering density (0.692→0.757), approaching the uniform limit alpha=1 at the heaviest scattering condition' should be replaced with a two-level comparison (low/moderate vs high) and the error bars should be propagated into any statement about a trend. In addition, the sigmoid and power-law functional forms are selected post hoc without a derivation or a goodness-of-fit comparison against alternative models; this does not invalidate the qualitative robustness result, but alpha should not be presented as a validated physical parameter without model-selection evidence.","section":"Table I and Section III.A"},{"comment":"The focal-depth discriminability metrics are computed from a latent subspace that was explicitly trained with a dedicated classifier head and a cross-entropy focal loss on the same three depth labels. Silhouette scores, inter-cluster distances, and classification accuracy on the held-out test set therefore partly reflect the model's success in fitting those labels, rather than an intrinsic property of the speckle patterns. The held-out test set and cross-seed consistency rule out training overfitting, but they do not rule out that the model is using any available cues (including the confounds discussed above) to achieve separation. To support the claim that scattering 'enables' depth discrimination, the authors should either (i) train a VAE without the focal classifier and focal loss and show that the raw encodings or raw speckle patterns still cluster by depth, or (ii) apply an unsupervised clustering method (e.g., k-means) to encodings from a classifier-free VAE and report the resulting adjusted Rand index or similar metric. Without such a control, the conclusions in Tables II and III are circular with respect to the training objective.","section":"Section III.B, Tables II and III, Fig. 6(b)"}],"minor_comments":[{"comment":"The phrase 'not only the coal depth information' appears to contain a typo; it should read 'not only the focal depth information.'","section":"Section II.B.1.b"},{"comment":"Fig. 6 reports SOPCNN accuracy and focal classification accuracy without error bars or standard deviations, even though the text states that five independent training seeds were used. Please add error bars or report mean±std in the text, as done for Tables I–III.","section":"Fig. 6"},{"comment":"The sentence 'each experimental condition was run across five independently drawn random seeds' should specify what constitutes a condition (per dataset and per focal-depth set) and whether the SOPCNN pre-trained classifier was retrained for each seed or fixed.","section":"Section II.B.2"},{"comment":"The caption states that the fits were performed on average accuracies over five seeds, but the plotted data do not show individual-seed spread or confidence intervals. Including the data points or error bars would help the reader judge the goodness of the sigmoid and power-law fits.","section":"Fig. 4"},{"comment":"The description of the 'No Scattering' dataset as a direct SLM-to-camera transmission is ambiguous given that the 4f relay and zero-order blocker are part of the optical path. Please clarify whether the no-scattering case used the same 4f relay and zero-order blocking without the scattering medium, since a zero-order blocker changes the effective image of the phase-only SLM.","section":"Section II.A"},{"comment":"Reference [12] is an arXiv preprint; if a peer-reviewed version exists, it should be cited instead. The text also cites [13] for the split-latent architecture but does not specify which aspects of the architecture are borrowed versus newly introduced; a sentence in Section II.B.1.b clarifying this would improve reproducibility.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's second central claim—that scattering enables focal-depth discrimination—rests on an unspecified optical implementation and on metrics derived from a model trained with the same labels it is used to evaluate. The first issue is the more serious one, because no amount of statistical care can fix a confounded data-generation protocol. I recommend that the editor require the authors to describe the focal-depth implementation in detail and to add at least one control experiment or unsupervised analysis before the manuscript is considered for publication. The robustness-under-occlusion result is more solid, though the monotonic-alpha claim needs to be softened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The occlusion-robustness result is solid and the split-latent VAE is a genuinely new architecture, but the focal-depth claim is undermined by an underspecified experimental setup. Fig. 5(a) never says how the three focal depths were produced, so the classifier may be reading intensity, magnification, or alignment differences instead of depth.\n\nThe good parts first. The circular-masking experiment is a clean proof-of-concept that scattering spreads digit information across the detector. Fitting no-scattering data with a sigmoid and scattered data with a power law is a reasonable way to compare information distributions, and the visual examples in Fig. 3 are convincing. The split-latent VAE is the real novelty: forcing clean reconstruction from the digit subspace alone is a clever architectural constraint, and jointly training a focal classifier on the other subspace is a sensible way to test depth decodability. Five seeds, held-out test metrics, and silhouette/inter-cluster distance checks are above the usual bar.\n\nThe soft spots are in proportion. The biggest is the focal-depth confound. The full text gives one sentence about 'projecting the same MNIST pattern onto each of three focal depths of the scattering medium,' with no mention of whether the medium, camera, SLM phase curvature, or 4f relay was changed. Each of those changes other optical parameters. This is load-bearing: if the three datasets differ in total intensity or magnification, the ~98% classification and high-scattering cluster separation are explained without any physical depth signal. This is fixable with a proper experimental methods paragraph and ideally control measurements, but right now the second central claim does not hold.\n\nSecond, the monotonic alpha claim is overstated. Table I shows 0.692 ± 0.030, 0.694 ± 0.022, and 0.757 ± 0.047. Low and moderate are statistically indistinguishable; the trend only clearly appears at high scattering. Saying alpha increases monotonically with scattering density is not supported by the data as presented.\n\nThird, the functional forms were chosen post hoc. The sigmoid for no-scattering and power law for scattering fit the data, but there is no independent reason given for those choices. Minor.\n\nFourth, the focal classification is trained with a classifier head and focal loss on the same labels, so the separation is partly fitted. The held-out test set and cross-seed consistency reduce this concern, but the architecture only shows that the latent space can encode depth when explicitly pressured to do so, not that depth is intrinsically decodable.\n\nFinally, no code, data, or hyperparameters are provided, which makes reproduction harder than it should be for a methods-focused paper.\n\nWho this is for: people working on speckle imaging, scattering-based sensing, or interpretable latent spaces. It is a proof-of-concept, not a definitive result. I would send it to peer review because the architecture is novel and the occlusion result deserves an audience, but the referee should require the experimental details and a revision of the alpha claim. If this comes back with the focal-depth method specified and the data available, it could be a solid contribution.","headline":"Scattering-as-feature proof-of-concept with a genuinely novel split-latent VAE, but the focal-depth experiment is underspecified enough that the paper's second central claim is not yet established.","tokens_in":7551,"tokens_out":3222,"would_cite":false,"duration_ms":27348,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that optical scattering, usually treated as an obstacle, can make captured images more robust to occlusion and can encode the focal depth of an object hidden inside the scattering volume.","keywords":["optical scattering","speckle imaging","variational autoencoder","split-latent VAE","focal depth discrimination","occlusion robustness","MNIST","information distribution"],"falsifier":"Repeat the focal-depth experiment while changing only the axial distance of the projected pattern relative to the scattering medium—for instance, translating the medium along the optical axis with the camera, lenses, illumination, and SLM untouched. If the focal classifier drops to chance once magnification and intensity are held constant, the claimed depth encoding is not a property of scattering; if the pattern must be refocused and any refocusing changes image scale or brightness, those confounds need to be controlled and shown not to drive the result.","tokens_in":6628,"feed_emoji":"💡","tokens_out":7715,"duration_ms":66558,"temperature":0.7,"pith_summary":"The paper is trying to establish that scattering is not only a lossy corruption but a physical process that reshapes how information is stored in light. By projecting MNIST digits through scattering media of three strengths and decoding the resulting speckle patterns with a variational autoencoder, the authors show that object information spreads across the whole detector plane: scattered images survive circular occlusion far better than unscattered ones, with heavier scattering moving the accuracy-versus-masked-area curve toward the uniform-information limit. They also claim that the same speckle patterns carry a decodable signature of the axial focal depth where the pattern was projected inside the medium, and that a split-latent VAE can separate this depth information from the digit content. These results matter because they suggest a way to build imaging systems that use obstacles and volumetric scattering as part of the sensing mechanism, rather than only as something to undo.","feed_headline":"Scattered light keeps images readable and reveals depth","feed_subtitle":"Speckle patterns keep digits readable under occlusion and encode focal depth up to 98 percent.","key_machinery":"The load-bearing object is the speckle field generated by multiple scattering—light redistributed into a grainy intensity pattern that carries object information across the full detector rather than in compact central features. The load-bearing architecture is the split-latent VAE, a variational autoencoder whose latent vector is split into two disjoint subspaces: a digit subspace $\\mathbf{z}_{\\mathrm{digit}}$ used exclusively to decode clean MNIST images, and a focal subspace $\\mathbf{z}_{\\mathrm{focal}}$ used together with the digit subspace to decode scattered images and, through a jointly trained classifier head, to predict focal depth. This split enforces a functional separation: the digit subspace must carry the object content, while the focal subspace absorbs depth-dependent and scattering-dependent variation. The occlusion analysis fits the scattered-data accuracy curves to a power law whose exponent quantifies how uniformly information is distributed, and the depth analysis measures cluster separation in the focal subspace via silhouette scores and centroid distances.","core_discovery":"On its own terms, the paper claims two concrete discoveries. First, information in a speckle pattern is spatially distributed: under a centered circular mask whose radius reaches 192 pixels of a 384×384 image, reconstruction of unscattered MNIST collapses to roughly 10% SOPCNN accuracy, while scattered datasets remain substantially decodable, and the fitted power-law exponent $\\alpha$ rises from 0.692 (low) to 0.757 (high) scattering—toward 1, the value for perfectly uniform information distribution. Second, scattering imprints axial position into the speckle field: using a split-latent VAE whose digit latent subspace alone reconstructs clean images and whose focal subspace is trained by a classifier head to predict one of three focal depths, the authors obtain focal classification accuracy that rises with scattering strength, reaching about 98% for high scattering, with silhouette scores and inter-cluster distances increasing accordingly. The trade-off is explicit: stronger scattering improves depth discrimination while degrading digit reconstruction accuracy from roughly 96–97% to about 85%.","pith_inferences":["The authors test only three discrete focal depths; a natural extension is continuous axial localization, which would reveal whether the focal latent manifold is ordered by physical depth or merely separates the three trained positions.","The occlusion study uses centered circular masks; testing random pixel dropout or arbitrary geometric occlusions would show whether the robustness is a general consequence of information spreading or specific to the mask shape.","The observed asymmetry between adjacent focal planes under high scattering may encode physical properties of the medium, such as its thickness or particle distribution; measuring those properties could turn the asymmetry into a calibration signal.","If the depth signal is truly carried by speckle statistics rather than by alignment artifacts, the same split-latent approach could be applied to three-dimensional objects instead of MNIST planes projected at three depths, testing whether real-volume depth is similarly separable."],"forward_implications":["Imaging systems with scattering in the optical path can tolerate large central occlusions or dead sensor regions, because the information needed to reconstruct the object is spread across the speckle field rather than localized.","A single speckle exposure can be read for two kinds of information at once: the object identity and its axial position inside the scattering volume, without a separate depth-sensing path.","Scattering strength becomes a tunable design parameter: heavier scattering buys better focal-depth separation (up to about 98% classification accuracy) at the price of reconstruction quality (SOPCNN accuracy down to about 85%), so a system can be optimized for one goal or the other.","Because the digit and focal subspaces are trained to be functionally separate, the split-latent VAE provides an interpretable route to disentangling object content from viewing conditions in scattering-based imaging."],"supporting_citations":[{"why":"Supplies the variational autoencoder formalism, ELBO objective, and reparameterization trick used throughout.","marker":"[7]"},{"why":"Supplies the stochastic backpropagation training method for the VAE.","marker":"[8]"},{"why":"Prior demonstration that depth information survives scattering and can be recovered from a single speckle pattern.","marker":"[11]"},{"why":"Prior deep-learning result that 3D phase images can be recognized through scattering, supporting the depth-signal premise.","marker":"[12]"},{"why":"The domain-invariant VAE whose latent-partitioning idea the split-latent architecture builds on.","marker":"[13]"},{"why":"Provides the SOPCNN classifier used to measure whether reconstructed images preserve digit identity.","marker":"[14]"}],"fun_headline_variants":["Scattering distributes image info and embeds focal depth","Speckle patterns preserve images and carry depth information","Scattering turns speckles into robust, depth-aware images","Optical scattering aids imaging by distributing and encoding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The focal-depth result assumes that the three datasets differ only in the axial position of the projected pattern inside the scattering medium, with no unintended changes in intensity, magnification, or camera alignment supplying the classification signal.","fun_headline_variants_meta":{"raw":{"variants":["Scattering distributes image info and embeds focal depth","Speckle patterns preserve images and carry depth information","Scattering turns speckles into robust, depth-aware images","Optical scattering aids imaging by distributing and encoding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000299,"raw_usage":{"total_tokens":1702,"prompt_tokens":893,"completion_tokens":809,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":745}},"tokens_in":509,"tokens_out":809,"duration_ms":8120,"temperature":1.0,"reasoning_tokens":745,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:28:37.625984+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the focal-depth experiment while changing only the axial distance of the projected pattern relative to the scattering medium—for instance, translating the medium along the optical axis with the camera, lenses, illumination, and SLM untouched. If the focal classifier drops to chance once magnification and intensity are held constant, the claimed depth encoding is not a property of scattering; if the pattern must be refocused and any refocusing changes image scale or brightness, those confounds need to be controlled and shown not to drive the result.","supporting_citations":[{"cited_title":"J., Mohamed, S., & Wierstra, D","cited_arxiv_id":null,"evidence_quote":"Supplies the stochastic backpropagation training method for the VAE."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior demonstration that depth information survives scattering and can be recovered from a single speckle pattern."},{"cited_title":"Recognizing three-dimensional phase images with deep learning","cited_arxiv_id":"2107.10584","evidence_quote":"Prior deep-learning result that 3D phase images can be recognized through scattering, supporting the depth-signal premise."},{"cited_title":"M., Louizos, C., & Welling, M","cited_arxiv_id":null,"evidence_quote":"The domain-invariant VAE whose latent-partitioning idea the split-latent architecture builds on."},{"cited_title":"Stochastic Optimization of Plain Convolutional Neural Networks with Simple methods","cited_arxiv_id":"2001.08856","evidence_quote":"Provides the SOPCNN classifier used to measure whether reconstructed images preserve digit identity."}],"review_version":2}