{"id":"e8d29b64-87d2-4728-a633-e039a0671777","arxiv_id":"2509.03193","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A miniaturized fiber-scanning endoscope paired with a spatio-temporal DenseNet estimates localized tissue elasticity from sparse wave-field images, with lower phantom errors than conventional phase-tracking elastography.","lead":"Researchers built a 2.2 mm fiber-scanning endoscope and a deep learning pipeline that estimates tissue stiffness from rapidly changing wave fields in real time. The method reduced elasticity errors on gelatin phantoms compared with conventional phase tracking and works without knowing the wave direction, a step toward stiffness-guided minimally invasive surgery.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth transfer from separate indentation phantoms is the unvalidated lynchpin; reported errors are relative to nominal concentration labels, not co-located modulus measurements.","rationale":"The reader's weakest_assumption correctly identifies the ground-truth transfer problem: indentation labels come from separate phantoms, not the imaged samples. This is the most load-bearing concern because every reported error (4.48±3.63 kPa in the abstract, Table II averages, and the per-concentration MAEs) is computed relative to those labels. If the label transfer is invalid, the headline accuracy numbers are not meaningful, and the claim of a 'robust, local, contactless, and non-directional estimate of tissue elasticity' is not substantiated for absolute modulus values. The paper provides no co-located validation, and the discussion explicitly acknowledges that indentation moduli and shear-wave-derived moduli 'differ and are complex to model'—yet delegates this conversion to the neural network without demonstrating that the conversion is consistent across phantom batches. Other issues (e.g., the impossible reported precision 0.94±2.51 for a proportion, and the conventional baseline being fit to the test data) are real but secondary; the baseline fitting actually makes the DL comparison conservative, and the precision typo does not affect the central accuracy claim. The proposed concrete test directly checks whether the transfer assumption holds by comparing network predictions to per-phantom or co-located indentation values. If the test shows a large discrepancy, the paper would need to be revised to report accuracy relative to a validated ground truth, potentially downgrading the central claim. If the test shows agreement, the concern is resolved. Given that the concern is addressable and the paper is otherwise technically sound, the reader's CONDITIONAL verdict is appropriate; no change to the verdict is needed from this pass.","tokens_in":13604,"tokens_out":6007,"duration_ms":68677,"concrete_test":"Recompute the reported MAE using per-phantom indentation measurements taken on the exact same phantoms that were imaged (or at least on multiple samples from each imaging phantom batch), instead of the single concentration-level label. Specifically, after imaging a homogeneous phantom, perform indentation on that same phantom (or a sample from the same poured batch) and compare the network prediction to that co-located modulus. If the resulting MAE is substantially larger than the reported 4.48 kPa (e.g., by more than the inter-phantom standard deviation of indentation moduli within a concentration), the ground-truth transfer assumption fails and the central error claim is not validated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (Abstract, Table II) is that the pseudo-3D DL method achieves a mean absolute error of 4.48±3.63 kPa relative to Young's moduli obtained from indentation tests. However, the indentation tests were performed on separately manufactured cylindrical phantoms (Section II-F), not on the same phantoms that were imaged. The imaged homogeneous and inclusion phantoms are never co-located with a mechanical measurement. Because the network's training labels are a single indentation modulus per gelatin concentration (Table I, e.g., 8.28 kPa for G3%, 97.22 kPa for G15%), the model is effectively trained to predict a nominal concentration-level modulus, not the actual modulus of the specific phantom being scanned. Any batch-to-batch variation in gelatin preparation, temperature history, or aging between the indentation phantoms and the imaging phantoms directly biases the learned mapping and all reported error values. The inclusion-phantom DICE score (0.91±0.03) does not rescue this, since it is based on thresholded relative contrast, not absolute accuracy. Without a co-located validation, the claim of a 4.48 kPa error is unsupported and could be off by the magnitude of the batch-to-batch variation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a fiber scanning endoscope (FSE) with a 2.2 mm outer diameter that acquires high-speed OCT images of propagating shear wave fields in either 2D+t (line) or pseudo 3D+t (conical) scan patterns. The authors propose an end-to-end spatio-temporal DenseNet that takes raw phase data (or ST maps) and directly regresses Young's modulus, avoiding the need for explicit wave-direction estimation, triggering, or spatial calibration of the excitation position. They compare this against a conventional FFT-based phase-velocity estimator with and without angle correction, using leave-one-phantom-out cross-validation on homogeneous gelatin phantoms. The headline result is that the pseudo-3D DL method achieves a mean absolute error of 4.48±3.63 kPa, compared with 19.75±21.82 kPa for the conventional FFT method and 11.33±12.78 kPa for the angle-corrected version. A held-out test on inclusion phantoms reports a DICE score of 0.91±0.03, and ex-vivo porcine heart images demonstrate the feasibility of mapping adipose and muscle tissue. The central claim is that a robust, local, contactless, non-directional elasticity estimate is feasible with a miniature FSE.","tokens_in":13788,"tokens_out":6789,"duration_ms":75320,"significance":"If the reported accuracy holds, this is a significant contribution to interventional elastography. The work is, to the best of the reviewer's knowledge, the first application of a fiber scanning endoscope to shear wave OCE, and the deep learning pipeline eliminates several practical bottlenecks of conventional methods: it does not require knowledge of the excitation position, does not need triggered acquisition, and avoids explicit wave-direction reconstruction. The experimental design is a clear strength: the leave-one-phantom-out cross-validation, the held-out inclusion-phantom test, the systematic variation of imaging position with a robot, and the reported real-time inference rate (about 14 Hz) all support the practical feasibility claim. The DL approach also consistently outperforms the conventional baseline, even when the baseline is given the excitation position and a test-set-fitted calibration, which makes the relative comparison conservative and credible. The paper is clearly of interest to the IEEE TMI readership, but the absolute accuracy claims require further validation.","major_comments":[{"comment":"The ground-truth Young's moduli used for both training and evaluation come from indentation tests on separately manufactured cylindrical phantoms, not from the phantoms that were actually imaged. The text states this explicitly: 'indentation tests performed on additionally manufactured cylindrical gelatin phantoms containing the same gelatin to water ratio.' Thus the reported MAEs (e.g., 4.48±3.63 kPa) are errors relative to nominal concentration-level labels. Batch-to-batch variation in gelatin preparation, temperature history, or aging could bias every reported error, and the magnitude of this bias is unquantified. Since the absolute error is the paper's headline quantitative claim, this is a load-bearing issue. The authors should either provide co-located mechanical measurements on the imaged phantoms (or samples from the same batch) or explicitly reframe the claims as relative to a c","section":"Section II-F / III / Table I"},{"comment":"The conventional baseline's linear calibration E = k·v + q with k=24.2, q=-16.4 is stated to be obtained by minimizing the absolute error between FFT estimates and indentation values, apparently on the same evaluation data used to report the final errors. This is a test-set calibration, giving the baseline an advantage that the DL method does not receive. While the DL method still outperforms the baseline, the comparison is not on a methodologically equal footing. The calibration should be performed within each cross-validation fold, or at minimum the authors should explicitly state that this calibration is a best-case upper bound for the conventional approach and quantify the sensitivity of the coefficients.","section":"Section II-D"}],"minor_comments":[{"comment":"The entry '12.13±79.1' appears to be a typographical error; the standard deviation 79.1 is implausibly large compared with the surrounding values. Likely it should be '12.13±7.91'.","section":"Table II, G15% row"},{"comment":"The reported precision value '0.94 ± 2.51' for the FFT+AC method is impossible: precision cannot exceed 1, and a standard deviation of 2.51 is invalid for a quantity constrained to [0,1]. This is likely a typo (e.g., 0.94±0.25) and must be corrected.","section":"Figure 7 / Results"},{"comment":"The sentence 'As a baseline, we used densely connected neural networks [31]' is vague because the paper immediately describes a custom DenseNet. Please clarify whether the literature baseline refers to the architecture family or a specific prior implementation, and describe any modifications made here.","section":"Section II-E"},{"comment":"The description 'We trained in total four networks for each fold while randomly choosing a phantom from each gelatin concentration for validation' is unclear. It should explain why four networks are trained (e.g., four random validation splits) and how their predictions are combined or selected.","section":"Section II-G"},{"comment":"The sentence 'apart from the 3% gelatin concentration the MAE is always best when the FSE is operated in pseudo 3D+t scan mode' is contradicted by Table II, where the pseudo 3D+t method has the lowest MAE at all concentrations, including G3%. This should be corrected.","section":"Section III"},{"comment":"The sentence 'The size of our scan field already represents a typical laparoscopic image size of 53 × 40 mm at a working distance of 120 mm [44]' is confusing, as the scan field diameter is stated as approximately 1.5 mm. Please rephrase to clarify the intended comparison (e.g., the size of the elasticity map, the field of view after mosaicking, or the lateral resolution).","section":"Section IV"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a genuinely novel application and a solid experimental framework, and the relative superiority of the DL approach over the conventional baseline is credible even with the baseline given the excitation position and a test-set calibration. The central concern is the absence of co-located mechanical validation: all absolute errors are tied to nominal concentration labels from separate indentation phantoms. I believe this is fixable by adding a co-located validation set, or by carefully reframing the claims and quantifying the expected batch variability. A second methodological point, the test-set-fitted baseline calibration, should also be addressed. I would encourage the editor to seek a revision rather than reject, as the engineering contribution and the general methodology appear sound. The manuscript also contains several typographical errors and one internally contradictory statement about the G3% results, which should be cleaned up in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth taking seriously. The hardware is genuinely new: the first fiber-scanning endoscope for quantitative shear wave elastography, with a 2.2 mm outer diameter and a conical pseudo-3D scan pattern that keeps temporal sampling at 5 kHz. The end-to-end spatio-temporal DenseNet is a sensible way to handle that sparse data, and the authors did the validation right in the ways that matter most. Leave-one-phantom-out cross-validation, held-out inclusion phantoms with a DICE score, and systematic robot-controlled variation of the imaging position all support the central claim: a miniaturized probe can produce local, non-directional elasticity maps in real time without knowing the excitation location.\n\nThe soft spots are real but not fatal. The biggest one is the ground-truth transfer. The reported MAE values, including the headline 4.48 ± 3.63 kPa, are computed against indentation tests on separate cylindrical phantoms with the same gelatin-to-water concentration, not on the actual imaged samples. The model is effectively trained to predict a concentration-level nominal modulus. If a batch of gelatin differs in stiffness by even a few percent, every absolute error shifts by that amount. The relative comparison to the conventional baseline is still meaningful because both use the same labels, and the inclusion DICE test is contrast-based and robust to an overall offset. But the absolute accuracy claim is not yet supported. A co-located mechanical measurement on the imaged phantom would close this gap.\n\nA few smaller issues. The conventional baseline's linear calibration, k=24.2 and q=-16.4, is fit to minimize error on the evaluation data, so the comparison looks fair but the baseline's performance is optimistically estimated. The reported precision of 0.94 ± 2.51 for a proportion is impossible; likely a typo. And no code or data is released, which doesn't help reproducibility.\n\nWho gets value? Anyone working on endoscopic OCT, elastography, or surgical navigation. The paper deserves a serious referee. The ground-truth transfer issue is addressable and, if fixed, the central result should hold. My recommendation: engage with it, ask for the co-located validation, and treat the 4.48 kPa as provisional until then.","headline":"A genuinely new endoscopic elastography probe with a mostly sound validation; the absolute accuracy numbers rest on an unverified ground-truth transfer, so treat them as provisional rather than proven.","tokens_in":14397,"tokens_out":2617,"would_cite":true,"duration_ms":31896,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 2.2 mm endoscope combined with end-to-end deep learning can map soft-tissue elasticity in real time, without knowing wave direction.","keywords":["optical coherence elastography","fiber scanning endoscope","spatio-temporal deep learning","DenseNet","shear wave imaging","Young's modulus mapping","minimally invasive surgery","pseudo 3D+t scanning"],"falsifier":"Indent the exact phantom regions imaged by the endoscope immediately after scanning and compare the co-located Young's modulus with the network's predictions; if the mean absolute error against those directly measured values substantially exceeds the reported 4.48±3.63 kPa, the phantom-to-phantom transfer assumption fails.","tokens_in":13408,"feed_emoji":"🩺","tokens_out":6935,"duration_ms":67536,"temperature":0.7,"pith_summary":"This paper tries to show that a miniaturized 2.2 mm fiber-scanning endoscope, when paired with a deep learning pipeline, can estimate local soft-tissue elasticity in real time from complex, multi-directional wave fields—something current ultrasound and MRI elastography cannot do during minimally invasive surgery. The authors acquire OCT phase data at 5.05 kHz using a conical scan pattern that captures pseudo-3D information, and feed it into a spatio-temporal DenseNet that directly outputs Young's modulus. On gelatin phantoms spanning 3–15% concentrations, the deep learning approach achieves a mean absolute error of 4.48±3.63 kPa without any estimate of wave direction, compared with 19.75±21.82 kPa for a conventional FFT-based phase-velocity method. The paper also reports successful detection of a stiff inclusion (DICE 0.91±0.03) and demonstrates qualitative elasticity maps on ex-vivo porcine heart tissue, arguing that this combination makes contactless, direction-independent elastography feasible for surgical navigation.","feed_headline":"Miniature endoscope maps tissue stiffness in real time","feed_subtitle":"A 2.2 mm probe with spatio-temporal deep learning beats conventional elastography without knowing wave direction.","key_machinery":"The load-bearing object is the conical (pseudo 3D+t) scan pattern generated by deflecting the imaging fiber with a piezoelectric tube at 5.05 kHz, yielding a circular trajectory that samples a propagating wave field in three spatial-plus-time dimensions. The second component is a spatio-temporal DenseNet—a real-valued densely connected convolutional network with about 110,000 parameters—that takes phase-difference volumes (depth × lateral × time) and regresses Young's modulus. The network learns the mapping implicitly, absorbing the probe's scan geometry and the unknown wave directions during training on phantoms whose stiffness is labeled by indentation tests, so no explicit velocity estima","core_discovery":"The central claim is that a fiber-scanning endoscope with a conical scan pattern, together with an end-to-end spatio-temporal convolutional network, can convert raw optical-coherence phase data of a diffuse multi-frequency wave field directly into quantitative Young's modulus estimates, without knowing the excitation position or wave propagation direction. The paper reports that for pseudo 3D+t scanning the mean absolute error is 4.48±3.63 kPa, roughly a four-fold improvement over the conventional 2D FFT approach (19.75±21.82 kPa), and that the estimates are spatially uniform across the field of view. It further claims this is the first application of a fiber scanning endoscope to shear wave","pith_inferences":["The reported accuracy depends on the assumption that the imaged phantoms have exactly the same stiffness as the separately indented replicate phantoms; a co-located indentation on the imaged samples would be a stronger test of the mapping.","Since the network implicitly absorbs the probe's scan geometry and wave-field statistics, changing the scan pattern, probe hand-built imperfections, or excitation frequencies would likely require retraining or fine-tuning before the error figures carry over.","The multi-frequency excitation and diffuse wave fields suggest the network may be able to estimate viscoelastic or frequency-dependent properties, not just a single Young's modulus, if the training labels are extended.","The approach could transfer to other excitation sources, such as instrument-integrated piezoelectric actuators, as long as the resulting wave fields are captured by the same conical scanning geometry."],"forward_implications":["A 2.2 mm endoscope can generate localized elasticity maps at roughly 14 Hz, compatible with continuous scanning during minimally invasive surgery.","The method removes the need to know or calibrate the excitation location and wave propagation direction, simplifying probe design and deployment.","On phantoms with a stiff inclusion, deep learning estimates segment the inclusion with DICE 0.91±0.03, compared with 0.64±0.10 for the conventional method.","Because the network is trained end-to-end, the same architecture could be retrained on clinical labels to perform tissue classification directly, avoiding indentation tests.","Ex-vivo porcine heart results suggest the approach can distinguish adipose and muscle tissue, pointing toward surgical navigation applications."],"supporting_citations":[{"why":"Establishes the 2D+t deep-learning baseline for OCT elastography that this work extends to pseudo-3D scanning.","marker":"[7]"},{"why":"Shows that 4D spatio-temporal deep learning can estimate elasticity in volumetric OCT data, motivating the 3D+t design.","marker":"[8]"},{"why":"Provides the spatio-temporal deep-learning approach and the indentation-test protocol used to label phantom stiffness.","marker":"[14]"},{"why":"Demonstrates forward-viewing resonant fiber-optic scanning endoscope designs that enable high-speed FSE OCT.","marker":"[21]"},{"why":"Supplies the FSE design and the spatial-calibration procedure whose imperfections the deep network learns to absorb.","marker":"[23]"},{"why":"Defines the conventional FFT-based phase-velocity estimation from space-time maps that serves as the comparison baseline.","marker":"[26]"},{"why":"Provides the diffuse shear-wave spectroscopy method relevant to estimating velocity from multi-directional complex wave fields.","marker":"[27]"},{"why":"Introduces the densely connected convolutional network architecture on which the spatio-temporal network is based.","marker":"[31]"}],"fun_headline_variants":["Fiber scope + deep learning yields real-time stiffness maps","AI endoscope measures tissue stiffness without wave direction","Mini probe with deep nets cuts elastography error fourfold","2.2 mm scope uses AI for real-time tissue stiffness"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The Young's modulus labels that train and evaluate the network are measured on separate cross-sectional gelatin cylinders, and the reported errors assume those indentation values exactly match the stiffness of the phantoms actually imaged.","fun_headline_variants_meta":{"raw":{"variants":["Fiber scope + deep learning yields real-time stiffness maps","AI endoscope measures tissue stiffness without wave direction","Mini probe with deep nets cuts elastography error fourfold","2.2 mm scope uses AI for real-time tissue stiffness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000236,"raw_usage":{"total_tokens":1383,"prompt_tokens":826,"completion_tokens":557,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":491}},"tokens_in":570,"tokens_out":557,"duration_ms":6958,"temperature":1.0,"reasoning_tokens":491,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T11:06:01.450480+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Indent the exact phantom regions imaged by the endoscope immediately after scanning and compare the co-located Young's modulus with the network's predictions; if the mean absolute error against those directly measured values substantially exceeds the reported 4.48±3.63 kPa, the phantom-to-phantom transfer assumption fails.","supporting_citations":[{"cited_title":"Deep learning for high speed optical coherence elastog- raphy,","cited_arxiv_id":null,"evidence_quote":"Establishes the 2D+t deep-learning baseline for OCT elastography that this work extends to pseudo-3D scanning."},{"cited_title":"4D deep learning for real-time volumetric optical co- herence elastography,","cited_arxiv_id":null,"evidence_quote":"Shows that 4D spatio-temporal deep learning can estimate elasticity in volumetric OCT data, motivating the 3D+t design."},{"cited_title":"Ultrasound shear wave elasticity imaging with spatio-temporal deep learning,","cited_arxiv_id":null,"evidence_quote":"Provides the spatio-temporal deep-learning approach and the indentation-test protocol used to label phantom stiffness."},{"cited_title":"Forward-viewing resonant fiber-optic scanning endoscope of appropriate scanning speed for 3d OCT imaging,","cited_arxiv_id":null,"evidence_quote":"Demonstrates forward-viewing resonant fiber-optic scanning endoscope designs that enable high-speed FSE OCT."},{"cited_title":"High-speed fiber scanning endoscope for volumetric multi-megahertz optical coherence tomography,","cited_arxiv_id":null,"evidence_quote":"Supplies the FSE design and the spatial-calibration procedure whose imperfections the deep network learns to absorb."},{"cited_title":"Arterial stiffness estimation by shear wave elastography: Validation in phantoms with mechanical testing,","cited_arxiv_id":null,"evidence_quote":"Defines the conventional FFT-based phase-velocity estimation from space-time maps that serves as the comparison baseline."},{"cited_title":"Diffuse shear wave spectroscopy for soft tissue viscoelastic characterization,","cited_arxiv_id":null,"evidence_quote":"Provides the diffuse shear-wave spectroscopy method relevant to estimating velocity from multi-directional complex wave fields."},{"cited_title":"Densely connected convolutional networks,","cited_arxiv_id":null,"evidence_quote":"Introduces the densely connected convolutional network architecture on which the spatio-temporal network is based."}],"review_version":1}