{"id":"0cc58360-536a-43d1-bff0-885c9d589527","arxiv_id":"2601.18219","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Lensfree laser holography plus a five-network deep-learning ensemble scored HER2 status at 84.9% four-class / 94.8% binary accuracy on blinded patient cores, with Monte-Carlo-dropout uncertainty used to reject low-confidence predictions.","lead":"A UCLA team scanned breast-cancer tissue with a lens-free laser system and trained an AI to read HER2 cancer-marker scores from the resulting holograms, reaching 84.9% four-class accuracy on 412 blinded tissue cores. The ~$980 device (laser excluded) aims to bring automated HER2 scoring to clinics that cannot afford full digital pathology scanners.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 84.9% accuracy is measured after an unquantified QC filter that discards poor-staining/artifact cores; without knowing the exclusion rate, the headline overstates deployable performance.","rationale":"The reader's weakest_assumption identifies the same primary issue: unreported exclusion of cores due to staining/artifact quality, making the accuracy conditional on an unquantified filter. I focused on this rather than the pathologist-ground-truth concern because the QC filter directly changes the test-set denominator and is uniquely within the authors' control to report; label noise affects essentially all HER2 IHC studies, whereas the QC filter is a pipeline-specific selection that can be quantified and corrected. The concern is not that the authors acted improperly — the Methods explicitly disclose the QC step — but that the central numerical claim is not tied to a deployable population without knowing how much was discarded. A concrete re-evaluation on an unselected cohort, or at minimum a reported exclusion count, would settle whether the 84.9% generalizes. The existence proof is not broken: even on a curated set, lensfree holography plus deep learning can classify HER2 with reasonable accuracy, and the binary results and blue-channel results provide internal consistency. But the abstract's 'practical pathway' claim is overstated if the QC filter is large. Since the paper already received a CONDITIONAL verdict and my concern reinforces rather than redirects that verdict, I recommend no change to the reader's verdict.","tokens_in":14881,"tokens_out":4021,"duration_ms":46039,"concrete_test":"Report the total number of cores extracted from the 15 TMA slides and the number discarded by the QC filter, broken down by exclusion category. Then, using the same trained M=5 ensemble and the validation-selected FOM thresholds, run inference on the full set of cores from those slides (including discarded cores where labels can be recovered by the same pathologist-verified spreadsheet) and compute 4-class and binary accuracy on the complete unselected cohort. If the unselected accuracy is within ~2 points of 84.9%/94.8% and the exclusion rate is small (<5%), the headline stands. If accuracy drops materially or the exclusion rate exceeds ~10–20%, the abstract should be reframed as reporting QC-filtered performance rather than deployable accuracy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — 84.9% 4-class / 94.8% binary accuracy on blinded test cores — is conditional on an unquantified data-exclusion step. In Materials and Methods, 'Data preparation and labeling,' the authors state: 'We discarded tissue cores with poor staining quality (e.g., uneven or weak chromogen deposition, excessive background, or tissue folds) and cores affected by imaging artifacts.' The number or fraction of discarded cores is never reported. With 15 TMA slides each containing ~100–150 cores, the raw pool is roughly 1,500–2,250 cores, yet only 1,273 training/validation + 412 test = 1,685 cores remain. The missing cores must have been removed by this QC filter, but the count is absent. If the filter removed a large fraction of real-world slides, the 412 'blinded' cores are a selected subset, and the 84.9% figure does not estimate performance on unselected clinical material. The paper's resource-limited deployment framing relies on this filter being a small minority. All downstream numbers, including the uncertainty-corrected 84.9%, are measured only on cores that survived manual QC. This is the most load-bearing gap: it directly changes the denominator of the headline accuracy. The lack of confidence intervals is secondary; the Z3 operating-point selection is also a risk, but the QC filter is a more direct threat to the claim as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports an automated HER2 immunohistochemistry scoring pipeline that combines a compact lensfree holographic microscope with an EfficientNet-B0 classifier and Monte Carlo dropout uncertainty quantification. The system images ~1,250 mm² of tissue in ~15 minutes at an estimated hardware cost below $1,000 (excluding the tunable laser). On a blinded, patient-separated test set of 412 tissue microarray cores, the authors report 4-class HER2 testing accuracy of 84.9% and binary accuracy of 94.8% after FOM-based rejection of low-confidence predictions, with a 30.4% correction rate and 7.2% loss rate. The paper also compares performance against a brightfield whole-slide scanner baseline and analyzes monochrome (single-wavelength) holography and the effect of digital propagation distance.","tokens_in":15026,"tokens_out":2477,"duration_ms":30431,"significance":"If the reported accuracies hold on unselected clinical material, this would be a meaningful demonstration that a low-cost, lensfree imaging platform can approach the scoring performance of conventional brightfield scanners, with potential value for resource-limited settings. The use of a blinded test set with patient-level separation, the inclusion of an uncertainty-based rejection mechanism, and the direct comparison with a brightfield scanner are notable strengths. However, the headline accuracy is conditional on an unquantified quality-control filter and on threshold choices that require clearer reporting; these gaps currently limit the strength of the claim.","major_comments":[{"comment":"The paper states that 'We discarded tissue cores with poor staining quality ... and cores affected by imaging artifacts,' but it never reports how many cores were discarded. With 15 TMA slides of ~100–150 cores each, the raw pool is 1,500–2,250 cores, while only 1,273 + 412 = 1,685 cores remain. If the discarded fraction is substantial, the 84.9% test accuracy is measured on a selected subset and overstates performance on unselected slides. Please report the exact exclusion count/rate and, ideally, a sensitivity analysis or a description of the excluded cases. This is the most load-bearing missing number for the central claim.","section":"Materials and Methods, 'Data preparation and labeling'"},{"comment":"The FOM thresholds [12.1, 16.2, 11.6, 21.2] were 'determined on the validation set to achieve >30% correction rate,' and the test-set correction rate is reported as 30.4%. Thus the headline correction rate is the test realization of the tuning target. Although the thresholds were apparently applied blindly to the test set, this makes the 30.4% figure an expected consequence of the selection rule rather than an independent discovery. Please report validation-set correction/loss rates, describe the threshold-selection procedure more explicitly (e.g., the grid searched, the selection criterion), and provide confidence intervals for test-set correction and loss rates.","section":"Materials and Methods, 'Uncertainty quantification and filtering using MC dropout' and Results, 'Uncertainty quantificat"},{"comment":"Figure 6 shows that Z3 = 2.4 mm was chosen as the operating point and that it gives the highest test accuracy among the tested distances. The manuscript does not state whether Z3 was selected using a validation set or whether it was selected by inspecting test-set performance. If Z3 was tuned on the test set, the reported 84.9%/94.8% accuracies are optimistically biased, and the comparison with Z3 = 0 mm or Z3 = 2.84 mm is not a valid out-of-sample comparison. Please clarify the selection procedure for Z3 and, if it was test-set-selected, report a validation-based or nested evaluation.","section":"Results, 'Impact of the digital propagation distance on the automated HER2 scoring performance'"},{"comment":"The ground-truth labels are described as 'pathologist-verified ... by at least 3 certified pathologists,' but no consensus rule is given (e.g., majority vote, adjudication, or exact agreement requirement). Given the known inter-observer variability in HER2 IHC scoring, especially at the 1+/2+ boundary (ref. 40), the label noise could materially affect the reported class accuracies, particularly the 2+ class (68.0% before uncertainty rejection). Please specify the labeling protocol and, if possible, report inter-observer agreement or a label-noise sensitivity analysis.","section":"Materials and Methods, 'Data preparation and labeling' and Discussion"}],"minor_comments":[{"comment":"The brightfield comparison reports 82.8%/84.3% 4-class accuracies, but no confidence intervals or significance tests are provided. Given n=412, the difference between 84.9% (lensfree) and 82.8% (brightfield) is within sampling noise. Please add confidence intervals or a paired test to support the 'comparable' claim.","section":"Results, 'Automated HER2 scoring using lensfree holography' and Fig. S3/S4"},{"comment":"The text reports rejection rates for correctly classified cores by class, but not the class-wise acceptance counts for misclassified cores. Adding the accepted/rejected confusion-matrix counts in a supplementary table would help readers assess whether the rejection rule is clinically sensible.","section":"Results, 'Uncertainty quantification for HER2 scoring'"},{"comment":"The Basler CMOS sensor is listed with a unit price of $584 but a total of $689. Please reconcile this discrepancy or clarify whether the total includes additional components (e.g., cables or lens mounts).","section":"Table 1"},{"comment":"A few typographical issues: the abstract states '~1,250 mm^2' while the Methods say '~12.5 cm²' (equivalent, but should be consistent); Fig. 6 caption appears truncated in the manuscript text; and Eq. (1) uses subscripts/superscripts that are hard to parse. Please correct these.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The central engineering demonstration is plausible and the hardware cost/throughput argument is attractive. The main issues are reporting gaps rather than internal contradictions, but the unquantified QC exclusion and the validation-set-based threshold tuning directly affect how the reader should interpret the headline numbers. I think a major revision is appropriate: the authors should report exclusion counts, validation-set thresholds/performance, and the Z3 selection procedure. If those additions confirm the current numbers, the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The paper's core claim — that HER2 IHC scoring can be done from lensfree holograms — is real and survives scrutiny. The intensity-only angular-spectrum back-propagation, without phase retrieval, gives a six-channel input that clearly carries enough information; the binary accuracy is 88.8% with a single model, 92.5% with the ensemble, and 92.5% with blue-only holograms, which is a strong sign the result is not a fluke. The comparison to a brightfield scanner using the same pipeline (80.8% vs 82.8% 4-class) is reasonable, and the paper is upfront that the tunable laser is excluded from the cost estimate.\n\nThe main soft spot is the quality-control filter. The methods say they discarded cores with poor staining or imaging artifacts, but never report how many. With ~1,500–2,250 raw cores and 1,685 retained, the missing count matters. If the filter removes a large fraction, the 84.9% is measured on an easier subset. This is the one gap that could change deployment claims, and it needs to be quantified in revision.\n\nTwo secondary issues. First, Fig. 6 shows that Z3 = 2.4 mm is the accuracy peak, but the paper does not state whether that operating point was chosen on validation or test. If it was test-selected, the downstream numbers carry optimistic bias; the authors need to clarify. Second, no confidence intervals or significance tests are given, so the near-parity with brightfield is within noise either way. The 30.4% correction rate is also a validation-tuned target, but since they report alternative thresholds with 20.2% correction, that is more a framing choice than a flaw.\n\nOverall, this is a proof-of-concept with genuine novelty and a clear path to practical relevance. It is not a mature clinical system, but it does not claim to be. I would send it to peer review: the existence proof is solid and the open questions are answerable with additional analysis.","headline":"First lensfree-HER2 scoring with genuine novelty, but the headline accuracy is conditional on an unquantified quality filter and the propagation-distance selection needs clarification.","tokens_in":15721,"tokens_out":2918,"would_cite":true,"duration_ms":31534,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HER2 scoring of breast cancer tissue can be performed from lens-free holograms by a $980 device with deep learning, reaching 84.9% four-class accuracy and 94.8% binary accuracy on a blinded 412-core test set.","keywords":["HER2 scoring","lensfree holography","digital pathology","deep learning","uncertainty quantification","Monte Carlo dropout","immunohistochemistry","breast cancer"],"falsifier":"Apply the trained ensemble to an uncurated, consecutive series of HER2 IHC slides (no exclusions for staining quality or artifacts) and compare the resulting four-class accuracy and rejection rate to the reported 84.9% and 30.4%; if the accuracy drops materially or the rejection rate balloons, the curated-test estimate does not generalize. Alternatively, quantify how many cores were discarded in data preparation — if it exceeds roughly 15% of all cores, the headline accuracy is optimistic.","tokens_in":14573,"feed_emoji":"🔬","tokens_out":6008,"duration_ms":54408,"temperature":0.7,"pith_summary":"This paper establishes that HER2 scoring — the immunohistochemistry test that decides whether breast cancer patients qualify for trastuzumab — can be performed from lens-free holograms captured by a compact $980 device, with deep learning replacing the microscope. On a blinded test set of 412 unique tissue cores, the system achieves 84.9% four-class accuracy (0, 1+, 2+, 3+) and 94.8% binary accuracy, matching or slightly exceeding a brightfield whole-slide scanner run through the same neural pipeline. The authors show that the hologram's digitally back-propagated complex field, fed as a six-channel RGB tensor to an EfficientNet-B0 ensemble, carries enough diagnostic information to classify HER2 expression without a physical objective lens. They further show that Monte Carlo dropout uncertainty estimates allow the system to abstain on low-confidence predictions, correcting 30.4% of errors while discarding only 7.2% of correct ones.","feed_headline":"Lens-free holography hits 84.9% in HER2 breast-cancer grading","feed_subtitle":"Using laser diffraction and neural nets, a $980 no-lens device approaches scanner accuracy on blinded cores.","key_machinery":"The load-bearing element is lensfree inline holography paired with digital back-propagation: a compact laser-illuminated CMOS sensor records interference patterns of the stained tissue without any objective lens, and the angular-spectrum method propagates the intensity back toward the sample plane to synthesize a complex field. This six-channel RGB complex field — not a reconstructed image — is the direct input to an EfficientNet-B0 classifier, bypassing phase retrieval. The paper's other machinery is a confidence-aware five-network ensemble with a positive-class override, and a Monte Carlo dropout uncertainty score (FOM = baseline confidence divided by the standard deviation of 256 stochast","core_discovery":"The central claim is that the diagnostic information needed for HER2 IHC classification survives in lens-free inline holograms, provided the network is trained on a digitally back-propagated complex-field representation. The system records diffraction patterns of DAB-stained tissue under red, green, and blue laser illumination on a monochrome CMOS sensor, then uses angular-spectrum propagation (at a fixed digital distance Z3=2.4 mm) to turn each hologram into a six-channel tensor (real and imaginary parts per color). An ensemble of five EfficientNet-B0 classifiers, with a HER2-positive-sensitive fusion rule, yields 80.8% four-class accuracy; applying class-specific Monte Carlo dropout uncert","pith_inferences":["Editorial inference: the 84.9% figure is measured on a curated dataset from which an unquantified number of poor-staining or artifact-laden cores were removed; a deployment on consecutive uncurated slides would likely lower accuracy, and the system's real tolerance to staining variability is untested.","Editorial inference: the near-parity of blue-only illumination with full RGB suggests that DAB phase contrast at short wavelengths is the dominant signal; this could be validated by acquiring the same cores under blue light and checking whether the accuracy gap between blue and RGB shrinks as staining intensity varies.","Editorial inference: the same uncertainty-abstention scheme could be transferred to other equivocal IHC biomarkers (e.g., PD-L1, Ki-67) or to other low-cost imaging platforms, since the FOM is model-agnostic."],"forward_implications":["If correct, the result means a ~$980, lens-free device can deliver HER2 scoring accuracy comparable to a commercial brightfield scanner, removing the optics cost and bulk from the bottleneck of digital pathology.","Clinical workflows could adopt the FOM threshold as an explicit 'abstain' rule: slides below threshold are re-read by a pathologist, which corrects 30.4% of misclassifications while discarding only 7.2% of correct ones.","The blue-channel-only result (75.0% four-class, 92.5% binary) indicates a single-wavelength version could cut hardware cost further while retaining most binary clinical value.","The strong dependence on digital propagation distance (80.8% at Z3=2.4 mm vs 56.3% at Z3=0) shows the complex-field representation is what carries the classification signal, not the raw hologram.","Because the system needs no mechanical focusing or high-precision alignment, it is compatible with decentralized, low-infrastructure settings — the stated motivation of the work."],"fun_headline_variants":["No-lens holography grades HER2 at 84.9% with uncertainty checks","Lensless HER2 scoring: 84.9% accurate, uncertainty-aware","Portable holography nails HER2 scoring with built-in uncertainty","Cost-effective lensfree HER2 grading: 84.9% accuracy","Uncertainty-aware HER2 scoring: 84.9% on lensfree holograms"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Everything rests on the assumption that the unreported number of tissue cores removed for poor staining or imaging artifacts is small enough not to distort accuracy, and that the pathologist-consensus labels are correct ground truth even though HER2 1+/2+ boundaries are notoriously subjective.","fun_headline_variants_meta":{"raw":{"variants":["No-lens holography grades HER2 at 84.9% with uncertainty checks","Lensless HER2 scoring: 84.9% accurate, uncertainty-aware","Portable holography nails HER2 scoring with built-in uncertainty","Cost-effective lensfree HER2 grading: 84.9% accuracy","Uncertainty-aware HER2 scoring: 84.9% on lensfree holograms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001588,"raw_usage":{"total_tokens":6194,"prompt_tokens":792,"completion_tokens":5402,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":5307}},"tokens_in":536,"tokens_out":5402,"duration_ms":36330,"temperature":1.0,"reasoning_tokens":5307,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T08:02:47.160286+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the trained ensemble to an uncurated, consecutive series of HER2 IHC slides (no exclusions for staining quality or artifacts) and compare the resulting four-class accuracy and rejection rate to the reported 84.9% and 30.4%; if the accuracy drops materially or the rejection rate balloons, the curated-test estimate does not generalize. Alternatively, quantify how many cores were discarded in data preparation — if it exceeds roughly 15% of all cores, the headline accuracy is optimistic.","supporting_citations":[],"review_version":1}