{"id":"6f42e827-6878-4319-8fea-a879d3a3fc84","arxiv_id":"2504.13711","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Frequency-domain tests show self-mixing laser sensing beats a microphone for contact classification under targeted ambient noise, while broadband-noise advantages shrink and motor noise remains a limitation.","lead":"Two robot fingertips are compared for detecting contact vibrations: one using a tiny self-mixing laser, one using a microphone. The laser classifies cup contents more reliably than the microphone when loud, targeted noises come from another cup, but motor noise from the robot's own wrist still degrades it.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Targeted-disturbance claim rests on one test scenario; paired statistical testing and a sensor-matched feature extractor are needed before the 'clear winner' conclusion is robust.","rationale":"The reader identified the AST-as-evaluator assumption as the weakest premise, and I agree that it is the central methodological soft spot. However, I would push further: the more concrete, testable concern is not only the pretrained model's audio bias but the absence of paired significance testing over the 40 targeted trials, where per-condition n=10 and standard deviations are large. The reader's verdict of CONDITIONAL is appropriate; my concern does not overturn it, but it sharpens the condition. The paper has independent support: the time-domain SNR analysis (Fig. 4e) shows a qualitative isolation advantage that does not depend on the classifier, and the authors disclose the motor-noise limitation and the modality-specific classifier bias in Section V. I would keep the verdict at CONDITIONAL, requiring the sensor-matched classifier check and significance testing before full acceptance, which is exactly what the reader asked for.","tokens_in":5721,"tokens_out":1249,"duration_ms":10681,"concrete_test":"Re-run the targeted-disturbance comparison with a sensor-matched classifier: e.g., train separate small CNNs on raw SMI waveforms and raw microphone waveforms (or on each sensor's native spectral representation), and report paired per-trial accuracy differences across the 40 targeted trials with a permutation test. If the laser advantage persists when the microphone is given features matched to its physical response, the central claim is robust; if the ranking flips or the difference becomes insignificant, the 'clear winner' claim is an artifact of the audio-pretrained AST evaluator.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim ('SMI is still the clear winner' for targeted disturbances) is supported by Table I: laser 0.92±0.15 vs microphone 0.34±0.42. However, the comparison depends on a single downstream classifier (AST pretrained on AudioSet) and on mel-spectrogram features designed for audio. The microphone's near-zero accuracy on 'Bolts, mimicking robot' and 'Bolts, shaking' suggests the microphone features miss the relevant cue or the classifier fails to separate the classes, but the authors acknowledge in Section V that 'microphone spectrograms are more suited towards distinguishing different events.' The key unidentified weakness is whether sensor-specific features or a sensor-matched classifier could change the ranking. The claim also aggregates 40 trials over four targeted scenarios with high variance (e.g., microphone 0.34±0.42, laser 0.92±0.15, with playdough-shaking giving microphone 0.38±0.08 and laser 1.00±0.00) without paired significance testing. Given the small per-condition n=10, the conclusion may be sensitive to a few trials.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an experimental comparison of two robotic fingertip sensing modalities for extrinsic-contact sensing: self-mixing laser interferometry (SMI) and a bottom-port microphone, extending the authors' prior work [8] from the time domain to frequency-domain validation. Section III reports time-domain signal-to-noise measurements: silicone drops under white-noise playback (laser better by 13.5 dB) and bolt drops into a second cup held by a person (laser achieves an additional 6.4 dB of in-band isolation relative to the microphone, with claimed perfect separation). Section IV introduces a three-class classification task (empty cup / bolts / playdough) using an AudioSet-pretrained Audio Spectrogram Transformer fine-tuned per sensor, trained on 150 clean trials and tested on eight 10-trial disturbance sets: four ambient (music/white noise) and four targeted (a person shaking a nearby cup containing bolts or playdough, with the robot cup empty). Main reported results: validation accuracy 1.00 for both sensors; ambient disturbances 0.91±0.10 (microphone) vs 0.95±0.03 (laser); targeted disturbances 0.34±0.42 (microphone) vs 0.92±0.15 (laser), with strong per-scenario variation (microphone 0.00 on both bolts scenarios and 1.00 on playdough-mimicking; laser 0.86 on playdough-mimicking and 1.00 on playdough-shaking). Section V discusses motor noise, which floods the SMI spectrum, supporting a recommendation for higher SMI readout rates.","tokens_in":5932,"tokens_out":18489,"duration_ms":169857,"significance":"If the results hold, this is a useful and unusually well-documented sensor comparison. Strengths: the test sets are held out and contain disturbance conditions never seen during training; both sensors record simultaneously, enabling a paired comparison; the time-domain SNR measurements are independent of any learned model and corroborate the direction of the classification results; identical preprocessing, architecture, and pretraining for both sensors make the classifier comparison controlled; and the audio-native pretraining plausibly biases the evaluation against the SMI modality, making the laser's performance a conservative lower bound. The paper's honest discussion of the reduced broadband advantage and of the laser's motor-noise vulnerability is another strength.","major_comments":[{"comment":"The targeted-disturbance claim aggregates four 10-trial scenarios into 0.92±0.15 vs 0.34±0.42, but the reported standard deviation is computed over the five cross-validation models, which all score the same ten test trials, so it is not a trial-level sampling uncertainty, and no significance test is reported. With n=10 per scenario, binomial uncertainty is large: in the 'Playdough, mimicking robot' row the laser's 0.86±0.22 is statistically indistinguishable from the microphone's 1.00±0.00, and this row—the closest embodiment of the abstract's 'multiple robots simultaneously collecting data for the same task' analogy—goes against the aggregate direction. The aggregate is driven almost entirely by the two bolt rows, where the microphone scores 0.00. The same lack of inferential statistics affects the broadband claim: 0.95±0.03 vs 0.91±0.10 have overlapping ranges, and the white-noise row favors the microphone (1.00 vs 0.97). The authors should report per-scenario and aggregate trial-level inference (e.g., McNemar tests on the 10 paired trials, since both sensors record simultaneously, or bootstrap confidence intervals over trials), state the scenario mix that defines 'targeted disturbances', and either support or qualify the abstract's 'clear winner' and the conclusion's 'significantly outperforms' accordingly.","section":"Table I and Section IV-C (also abstract and conclusion)"},{"comment":"The entire frequency-domain comparison uses one feature/classifier pipeline—an AudioSet-pretrained Audio Spectrogram Transformer operating on 128-bin log-mel spectrograms—applied identically to both sensors' signals. Because the pretraining distribution and the mel filterbank are designed for human audio, the head-to-head accuracies may reflect classifier/feature fit rather than sensor physics; the bias could plausibly run in either direction, since the microphone receives an initialization advantage from audio-native pretraining while the laser's interferometric fringe signatures may be poorly represented by mel features, or, conversely, the mel smoothing may discard microphone-relevant fine structure. The authors' own Section V statement that 'microphone spectrograms are more suited towards distinguishing different events' makes this dependence explicit. Because the conclusions in Section V are about the sensors' noise resilience rather than about this particular classifier, the authors should add a robustness check with a sensor-matched or feature-agnostic control—for example, training the same architecture from scratch per sensor on the same features, or a second representation such as a raw-waveform CNN or fringe-rate spectral features for SMI—and show that the ranking and per-scenario accuracies are stable; otherwise the central comparison is confounded with the feature pipeline.","section":"Section IV-B"},{"comment":"The time-domain SNR comparisons—laser better by 13.5 dB under white noise, and 4.8 vs 11.2 dB separation with 'perfect separation' for the laser in the bolt experiment—are reported as single point estimates. The number of drops per condition (M), the number of independent measurement sessions, and any run-to-run variability are not given, and no confidence intervals or error bars appear in Fig. 4(c)-(e). These values are used in Section V as supporting evidence for the frequency-domain interpretation, so the comparative claims in this section cannot currently be assessed for precision; the authors should provide repetition statistics or clearly label the figures as single-session measurements.","section":"Section III-B, Eqs. (1)-(2)"}],"minor_comments":[{"comment":"The paper states that the model is trained with a 'binary cross-entropy loss' for a three-class task; please specify whether categorical cross-entropy or multi-label binary cross-entropy was used.","section":"Section IV-B"},{"comment":"The paper should state explicitly that in the four targeted test sets the robot-held cup is always empty (the disturbance comes from the person-held cup) and that accuracy is scored against the robot cup's true content; this is implied but never stated, and it is essential for interpreting the confusion matrices.","section":"Section IV-A"},{"comment":"The explanation for the laser's 0.86 on 'Playdough, mimicking robot' invokes 'an accidental, uncontrolled disturbance or the non-stationary noise floor'; since the non-stationary noise floor is a known property from [8], it should be treated as a quantifiable sensor limitation (e.g., measured during the test session) rather than an ad-hoc explanation.","section":"Section V"},{"comment":"Equation (2) does not define p_i; clarify that these are peak-normalized event amplitudes and specify how peaks are detected and over what window the noise power P_noise is estimated.","section":"Eq. (2)"},{"comment":"Reference [8] has an inconsistent citation year ('in 2023 IEEE International Conference on Robotics and Automation (ICRA), 2025'); the author name 'L. V. den Stockt' likely should be 'L. Van den Stockt'.","section":"References"},{"comment":"Fig. 7 should label both axes with units (time, frequency) and state the color-map normalization per panel; as printed, the contrast between the microphone and laser spectrograms is difficult to compare quantitatively.","section":"Fig. 7"},{"comment":"The sentence 'microphone spectrograms are more suited towards distinguishing different events' reads as contradicting the accuracy tables; rephrase to state the two-sided conclusion explicitly (the microphone's spectrograms carry more event-discriminative content but also more disturbance content, which is why its accuracy degrades when disturbances are present).","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"Two risks to communicate to the editor. First, Table I's standard deviations are computed over the five CV models rather than over trials, so the headline 'clear winner' claim currently lacks trial-level statistical support; the revision should be checked for per-scenario paired statistics (McNemar or bootstrap over the 10 paired trials). Second, the requested sensor-matched/from-scratch classifier control is a substantive experiment: if it changed the per-scenario ranking, the paper would need substantial reframing, so this should be requested as a condition of acceptance rather than merely suggested. I also disagree mildly with one characterization of the stress-test note: the microphone's 0.00 scores on the bolt scenarios are not best read as the features 'missing the cue'—the paper's own time-domain data (Fig. 4e) show the microphone strongly hears the nearby bolt cup, so the failure is the microphone faithfully detecting a disturbance that the task labels as noise. That reading actually supports the paper's physical narrative; the real weaknesses are the aggregation without statistics and the single feature/classifier pipeline."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper extends the authors' previous ICRA work on self-mixing interferometry (SMI) for tactile sensing. What's actually new: a frequency-domain classification task (cup empty/bolts/playdough using an Audio Spectrogram Transformer), targeted noise experiments where a second cup is shaken beside the robot, and a spectral analysis of motor noise. The time-domain SNR measurements are direct and clean, and the authors are refreshingly honest about SMI's weaknesses, especially motor noise and the need for higher readout frequencies. Design and data files are on GitHub, which helps reproducibility.\n\nThe main empirical contribution is the comparison under targeted disturbances. The bolt conditions are striking: microphone accuracy drops to 0.00 while laser stays around 0.90-0.94. That is consistent with the time-domain result showing the laser isolates in-band noise by an extra 6.4 dB. So the direction of the effect is probably real, and the AST-pretrained classifier concern does not change that fundamental separation.\n\nSoft spots, in order of importance. First, the aggregate 'clear winner' line in the abstract is too strong. In the playdough mimicking-robot condition, the microphone actually scores 1.00 versus the laser's 0.86. The authors mention this in the discussion, but the abstract and conclusion ignore it. A more precise claim would say SMI wins specifically for bolt-like in-band noise. Second, there is no significance testing anywhere. With n=10 per condition and cross-validation splits of five models, the reported standard deviations overlap for the aggregate. A paired test across trials would help. Third, the frequency-domain comparison relies on one classifier (AST) with audio-derived mel features, which the authors concede favor microphone spectrograms. That makes the broadband-noise comparison (0.95 vs 0.91) less conclusive, though the time-domain SNR numbers support the general point.\n\nNone of these are fatal. The paper is a validation study, not a new sensing principle, and it is transparent about its limits. I'd send it to peer review. The revisions I'd want are modest: reword the 'clear winner' claim, add paired statistics, and discuss whether a sensor-specific feature extractor could close the broadband gap.\n\nWho's it for: robotics researchers working on tactile sensing or multi-robot cells where acoustic cross-talk is a problem. It won't reorganize any field, but it gives a solid engineering data point.","headline":"An honest, incremental validation study of SMI versus acoustic sensing for robot contact detection; the headline 'clear winner' claim is mostly supported but overgeneralized in the targeted-noise section.","tokens_in":6448,"tokens_out":3141,"would_cite":false,"duration_ms":31073,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In frequency-domain tests, a self-mixing laser fingertip clearly outperforms a microphone under targeted ambient noise while keeping a smaller edge under broadband noise.","keywords":["self-mixing interferometry","tactile sensing","acoustic sensing","extrinsic contact detection","noise resilience","audio spectrogram transformer","robotic manipulation","frequency-domain classification"],"falsifier":"Retrain the classifiers separately for each sensor with sensor-specific features and with targeted-noise examples added to the training set, then rerun the targeted-disturbance test sets; if the microphone accuracy rises to the laser's level, the claim that SMI is the clear winner under in-band noise would not hold.","tokens_in":5514,"feed_emoji":"📡","tokens_out":11523,"duration_ms":96393,"temperature":0.7,"pith_summary":"This paper asks whether a self-mixing laser interferometry (SMI) fingertip, which detects microvibrations through light reflected back into a laser cavity, can serve as an ambient-noise-resilient replacement for a microphone in robotic manipulation. It extends the authors' earlier time-domain comparison with a frequency-domain classification task: a robot shakes a cup and a classifier must tell whether the cup is empty, holds bolts, or holds playdough, while music, white noise, or the sound of another cup being shaken nearby are played. The paper finds that SMI still outperforms acoustic sensing under broadband ambient noise, though the margin is smaller than the time-domain analysis suggested, and that under targeted disturbances that mimic the desired signal the laser is the clear winner. It also reports that motor noise spreads across the SMI spectrum more than the microphone spectrum, and concludes that raising the SMI readout frequency is the main next step.","feed_headline":"Laser fingertip beats microphones when ambient noise mimics the signal","feed_subtitle":"The laser stays at 92 percent accuracy when targeted noise cuts the microphone to 34 percent.","key_machinery":"The load-bearing mechanism is the self-mixing fringe: every time the target moves $\\lambda/2$, the optical path difference changes by one wavelength and a discontinuity appears in the photodiode current, so a sinusoidal vibration becomes a signal with the same fundamental frequency plus extra higher harmonics. This frequency-spreading signature is what gives SMI its isolation from airborne ambient sound, and it is also why motor noise bleeds across the laser spectrum. The frequency-domain evaluation is carried by the Audio Spectrogram Transformer, a transformer classifier pretrained on audio spectrograms, applied identically to 128-bin log-mel spectrograms from both sensors.","core_discovery":"The paper's central claim is that for targeted noise disturbances, analogous to multiple robots collecting data for the same task in the same environment, self-mixing interferometry is still the clear winner over acoustic sensing. In the frequency-domain experiment the laser reaches $0.92 \\pm 0.15$ accuracy on targeted-disturbance test sets while the microphone falls to $0.34 \\pm 0.42$; on broadband ambient noise the laser scores $0.95 \\pm 0.03$ versus $0.91 \\pm 0.10$, so it outperforms but less decisively than the time-domain SNR numbers implied. The laser's dominant error under broadband noise is confusing bolts for playdough, while the microphone's dominant error is confusing an empty cup for playdough. The paper also establishes that motor noise from the robot wrist floods the entire SMI spectrum rather than staying in narrow bands, and that this frequency spreading is fundamental to the SMI fringe mechanism.","pith_inferences":["The paper does not test a sensor-specific classifier; because the laser's spectrogram is not natural audio, a different feature representation could change the ranking, especially on broadband noise where the margin is only a few accuracy points.","The motor-noise result suggests that simply raising the readout frequency may not suffice; modeling the laser's nonlinear fringe transfer function could be needed to separate contact-induced harmonics from motor noise.","A direct extension is to add targeted-noise examples to the training set; since the paper trains only on clean data, the microphone's collapse on targeted disturbances may partly reflect a classifier blind spot rather than a pure sensing limitation.","Because the two sensors show different error modes, a combined fingertip with both SMI and a microphone could classify contact events more reliably than either alone under mixed noise."],"forward_implications":["In multi-robot or shared environments where nearby agents produce sounds similar to the contact events being sensed, an SMI fingertip can keep a cup-content classifier near 90 percent accuracy, while a microphone-based classifier can fall to near zero on the same trials.","Even under broadband noise, a microphone can still classify events at roughly 90 percent accuracy, so the SMI advantage there is modest rather than decisive.","Because motor noise spreads across the SMI spectrum, the paper's stated priority of raising the readout frequency above the current 18 kHz is necessary for separating contact signals from robot self-noise.","The frequency-spreading effect implies that features distinguishing contact events may currently sit in unmeasured bands, so a higher sampling rate could reveal separations the present readout cannot see."],"supporting_citations":[{"why":"Explains the self-mixing interference effect that the fingertip sensor exploits.","marker":"[1]"},{"why":"Provides the behavioral model of the SMI signal used to reason about fringe shape and frequency spreading.","marker":"[2]"},{"why":"Supplies the cup-shaking task and the audio-cue classification approach that the frequency-domain experiment is modeled on.","marker":"[7]"},{"why":"Describes the two fingertip designs and the earlier time-domain noise-resilience results that this paper extends.","marker":"[8]"},{"why":"Supplies the Audio Spectrogram Transformer architecture used as the classifier for both sensors.","marker":"[9]"},{"why":"Provides the AudioSet pretraining data that makes the transformer classifier transferable to both sensors' spectrograms.","marker":"[11]"}],"fun_headline_variants":["Laser fingertip beats mic when noise mimics the signal","SMI laser wins contact sensing under targeted robot noise","Laser sensor outperforms acoustic in noisy robot tasks","Targeted noise? Laser fingertip still beats microphone"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that one audio-pretrained classifier, applied to the same kind of sound-frequency image for both sensors, judges the laser fairly even though the laser spreads frequencies differently than a microphone does.","fun_headline_variants_meta":{"raw":{"variants":["Laser fingertip beats mic when noise mimics the signal","SMI laser wins contact sensing under targeted robot noise","Laser sensor outperforms acoustic in noisy robot tasks","Targeted noise? Laser fingertip still beats microphone"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000219,"raw_usage":{"total_tokens":1436,"prompt_tokens":930,"completion_tokens":506,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":442}},"tokens_in":546,"tokens_out":506,"duration_ms":5253,"temperature":1.0,"reasoning_tokens":442,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:01:54.265007+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the classifiers separately for each sensor with sensor-specific features and with targeted-noise examples added to the training set, then rerun the targeted-disturbance test sets; if the microphone accuracy rises to the laser's level, the claim that SMI is the clear winner under in-band noise would not hold.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the behavioral model of the SMI signal used to reason about fringe shape and frequency spreading."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the cup-shaking task and the audio-cue classification approach that the frequency-domain experiment is modeled on."},{"cited_title":"Proesmans, W","cited_arxiv_id":null,"evidence_quote":"Supplies the Audio Spectrogram Transformer architecture used as the classifier for both sensors."}],"review_version":1}