{"id":"259cb6a0-dc78-4d6c-b6e7-c905df7fc1b6","arxiv_id":"2504.18004","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Off-the-shelf general-purpose audio foundation models, without fine-tuning, reach state-of-the-art or baseline performance on clean respiratory and heart sound tasks, while a respiratory-specific foundation model underperforms.","lead":"Researchers tested five pre-trained audio models, with frozen weights, on four heart and respiratory sound datasets. They found the models match or beat prior fine-tuned state-of-the-art results on clean recordings, but fall short on noisy recordings, and general-purpose models beat a respiratory-specific one.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"BMD-HS comparison in Section II-F is not controlled: recording-level splits risk patient identity leakage and the baseline's task/split are unverified, so the 'SOTA on clean data' claim is not yet supported.","rationale":"The paper's central practical claim is that frozen general-purpose audio encoders match SOTA on clean medical audio. That claim is empirically anchored by two tables: SPRS (Table II) and BMD-HS (Table IV). The SPRS result is 'comparable to SOTA', not surpassing, and the comparison set is limited to models from the same authors. BMD-HS is the only table where frozen models clearly beat a published baseline by a large margin. But that table is also the least controlled: the task is redefined as binary, the splits are non-standard, and there is no evidence of subject-disjoint splitting. In a dataset with 8 recordings per subject, leakage from random splits is a well-known failure mode and can fully explain high accuracy. This is not a disagreement with the field's consensus; it is an internal validity problem. The paper deserves credit for releasing code and using standard protocols for ICBHI, SPRS, and CirCor. A conditional verdict is appropriate: the authors should either provide a patient-disjoint, task-matched BMD-HS comparison or soften the SOTA claims. The reader's weakest assumption pointed to the comparability of the BMD-HS baseline; our concern is stronger and more specific, so agreement is partial.","tokens_in":8177,"tokens_out":8455,"duration_ms":84538,"concrete_test":"Re-run the BMD-HS evaluation with subject-disjoint splits, assigning all 8 recordings from each subject to either training or test while stratifying by murmur label, and compare frozen BEATs and AST against the same baseline [21] under identical splits and the same binary task. If the frozen-model advantage over the baseline collapses or the 0.952 accuracy drops substantially, the headline claim is an artifact of split leakage or task mismatch.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section II-F evaluates BMD-HS by converting a multi-disease dataset into binary murmur presence/absence, creating three random stratified splits (576/288) over 864 recordings from only 108 subjects, and comparing against the baseline of the original dataset paper [21]. Two unverified choices are load-bearing. First, the splits are described only as 'stratified'; with 8 recordings per subject, random recording-level splits almost certainly place the same subject in both training and test, enabling patient-identity shortcuts that inflate the reported BEATs accuracy (0.952) and F1 (0.875). Second, the paper never shows that the [21] baseline used the same binary task, same subject grouping, or the same evaluation metric. If either fails, Table IV does not support the abstract's claim that frozen models 'achieved SOTA performance' on clean data; the claim would rest on SPRS alone, where the frozen models are within statistical noise of, but not above, the authors' own fine-tuned M2D-X (88.49 vs 89.77 in Table II). The concluding SOTA claim is therefore not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates five audio foundation models (AST, BYOL-A, M2D, BEATs, and OPERA-CT) as frozen feature extractors across four medical sound tasks: ICBHI2017 and SPRS for respiratory sounds and CirCor and BMD-HS for heart sounds. For each task, a task-specific classifier (MLP or a four-block transformer encoder) is trained on the frozen features, and results are compared with published fine-tuned baselines and SOTA results. The authors report that general-purpose audio models consistently outperform the respiratory-sound-specific OPERA-CT, that frozen models are competitive on clean-data tasks (SPRS, BMD-HS) but not on noisy-data tasks (ICBHI2017, CirCor), and they release evaluation code. The central claim is that frozen general-purpose audio representations are practically useful for clean medical sound recordings.","tokens_in":8356,"tokens_out":6150,"duration_ms":58683,"significance":"If fully supported, the paper would provide useful practical guidance: frozen general-purpose audio encoders can serve as strong feature extractors for clean stethoscope recordings, and domain-specific pretraining on a narrow respiratory corpus is not automatically superior. The paper is strong in its use of public models and checkpoints, its adherence to established evaluation protocols for ICBHI2017, SPRS, and CirCor, and its release of code. The comparison across four tasks with consistent methodology is a valuable contribution. However, the BMD-HS comparison is not controlled for subject identity or task definition, and the abstract's 'achieved SOTA performance' wording overstates the SPRS results. With the BMD-HS issue fixed and the claims scaled back, the paper would still be a useful benchmark for the community.","major_comments":[{"comment":"The BMD-HS comparison is not controlled for subject identity or task definition. The 864 recordings come from only 108 subjects (8 recordings each), and the paper describes randomly stratified splits of 576/288 recordings without stating that subjects are kept disjoint. With recording-level splits, the same subject almost certainly appears in both training and test sets, enabling the classifier to exploit subject-specific acoustic patterns and inflating the reported metrics (e.g., BEATs accuracy 0.952, F1 0.875). In addition, the baseline from [21] is not shown to use the same binary murmur-presence task, the same subjects, or the same evaluation metric; the paper simply asserts comparability. To support the claim that frozen models outperform or achieve SOTA on BMD-HS, the authors should report results with subject-disjoint splits (e.g., group split by subject ID) and verify that the baseline used an identical task and partitioning scheme. As written, Table IV does not substantiate the 'SOTA performance' claim for clean data.","section":"II-F, Table IV"},{"comment":"The abstract and conclusion state that the models 'achieved SOTA performance on the other tasks with clean data.' This is not supported for SPRS: in Table II, the best frozen model (BEATs) reaches a score of 88.49±0.82, while the SOTA fine-tuned M2D-X scores 89.77±0.16, and M2D (16×4) scores 88.07±0.53. These results are close but below the SOTA value; the text itself says 'comparable.' For BMD-HS, the only comparison is against the dataset paper's baseline, which is not established as a SOTA system. The claim should be rephrased to say that frozen general-purpose features are competitive with, or in some cases exceed, published baselines, but they do not generally surpass current fine-tuned SOTA.","section":"Abstract and Table II"},{"comment":"The paper's central insight—that frozen features work well on 'clean' data but require fine-tuning on 'noisy' data—relies on a qualitative, unmeasured clean/noisy distinction. The authors characterize SPRS and BMD-HS as 'clean' and ICBHI2017 and CirCor as 'noisy' based on dataset descriptions and anecdotal comments, but no quantitative noise metric is provided. This distinction is load-bearing for the explanatory narrative. A quantitative indicator of noise (e.g., estimated SNR, proportion of non-target acoustic events, or recording-device diversity) would strengthen the claim and allow a falsifiable test of the correlation. As it stands, the clean/noisy attribution is a post-hoc interpretation.","section":"III Discussion, II-D, II-F"}],"minor_comments":[{"comment":"The phrase 'off-the-shelf' may mislead readers, because a task-specific network (MLP or transformer encoder) is trained on the frozen features. Suggest clarifying that 'off-the-shelf' refers to the fixed feature extractor only, and that a trainable downstream classifier is used in all experiments.","section":"I and II-B"},{"comment":"Please clarify how the binary murmur-presence label is derived from the multi-disease labels in BMD-HS, and report the class distribution in the training and test splits. This information is necessary for assessing the comparability with the baseline.","section":"II-F"},{"comment":"The large standard deviation for M2D (16×4) specificity (7.25) indicates instability across the five attempts; consider reporting per-seed results or discussing this variance in the text.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"The BMD-HS comparison is the key weakness. If the authors cannot reproduce the original baseline's exact protocol, they should downgrade the claim to 'outperforms the published baseline under our split' or remove BMD-HS from the SOTA claim. The paper would benefit from an explicit subject-disjoint split even if it lowers performance; the current numbers look suspiciously high for frozen features. The abstract's 'achieved SOTA' is too strong and should be made precise in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth a look if you work on applied audio ML for physiological sounds. It evaluates four frozen general-purpose audio encoders on four respiratory and heart sound tasks, releases code, and sticks to standard splits and metrics so the numbers line up with published SOTA. The genuinely useful result is that AudioSet-pretrained models consistently beat OPERA-CT, a domain-specific respiratory model, on both respiratory tasks. The SPRS table is the most solid: frozen BEATs and M2D land within noise of the fine-tuned M2D-X SOTA, which is a practical datapoint for anyone wanting to skip fine-tuning on clean data.\n\nThe soft spot is the BMD-HS section. The paper converts a multi-disease dataset into binary murmur presence/absence, creates random stratified splits over 864 recordings from only 108 subjects, and compares against the baseline of the original dataset paper. Two things are unverified: whether the baseline used the same binary task, same splits, and same subject grouping, and whether the splits are subject-disjoint. With roughly 8 recordings per subject, random recording-level splits almost certainly place the same patient in both training and test, which can inflate the reported accuracies (BEATs at 0.952). Without subject-disjoint splits or at least a demonstration that the baseline setup matches, the claim of SOTA performance on BMD-HS is not supported. The abstract and conclusion overstate the clean-data result: on SPRS the frozen models are comparable to, not above, the fine-tuned SOTA, and BMD-HS is the only place they clearly exceed a baseline. The headline should be dialed back to 'matched or beat the published baseline, pending a cleaner BMD-HS comparison.'\n\nThe clean/noisy distinction is also qualitative; they do not measure noise levels, but that is a minor issue. The paper ships code and uses public checkpoints, so the core SPRS and CirCor comparisons are reproducible.\n\nWho is this for? Practitioners who want to know whether fixed-weight encoders suffice for clean medical audio tasks, and researchers building frozen-encoder multimodal pipelines. It deserves peer review: the method is sound on three of the four tasks and the SPRS result is valuable, but the authors should be asked to fix the BMD-HS protocol (subject-disjoint splits, verify the baseline, and soften the SOTA claim).","headline":"Useful frozen-feature benchmark for medical audio, but the BMD-HS 'SOTA' claim is undermined by unverified baseline comparability and likely patient leakage.","tokens_in":8915,"tokens_out":2557,"would_cite":true,"duration_ms":25073,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Frozen audio foundation models, used without fine-tuning, reach state-of-the-art-level performance on clean respiratory and heart sound tasks but not on noisy ones.","keywords":["audio foundation models","frozen feature extractor","heart sound analysis","respiratory sound analysis","self-supervised pretraining","murmur detection","ICBHI2017","CirCor"],"falsifier":"Re-run the BMD-HS evaluation with the dataset authors' original multi-class task definition and original train/test split, using the same frozen models; if AST, BEATs, and M2D no longer exceed the baseline's accuracy and F1, the clean-data heart sound claim fails. A second check: apply a denoiser to ICBHI2017 before frozen M2D feature extraction; the noise-based explanation predicts the score should move toward 61.2.","tokens_in":7965,"feed_emoji":"🩺","tokens_out":6372,"duration_ms":58690,"temperature":0.7,"pith_summary":"This paper tries to establish that off-the-shelf audio foundation models, used with frozen weights and only a small trainable head, are practically effective for respiratory and heart sound analysis. Comparing four public benchmarks against published state-of-the-art (SOTA) fine-tuned results, it finds that frozen general-purpose encoders match or exceed the reference results on the two clean datasets (SPRS and BMD-HS) but fall short on the two noisy ones (ICBHI2017 and CirCor). It also claims that general-purpose models pretrained on a large, diverse audio corpus outperform OPERA-CT, a model pretrained specifically on respiratory sounds, in respiratory tasks. If right, this supports using fixed-weight audio encoders as drop-in feature extractors in medical audio systems, while identifying noise as the main obstacle that still requires fine-tuning or preprocessing.","feed_headline":"Frozen audio models match top results on clean heart and lung sounds","feed_subtitle":"Fixed-weight general-purpose audio encoders beat a respiratory-specific model; noisy recordings still need fine-tuning.","key_machinery":"The central object is the frozen feature extractor: a pretrained audio encoder with weights locked, feeding a small task-specific network (a multilayer perceptron or a 4-block transformer encoder) that is trained on each task's training set. Because the encoder weights are fixed, task performance is attributed to the representation quality of the foundation model itself. The comparison machinery is alignment: each task uses the same data splits, preprocessing, augmentation, and metrics as the prior state-of-the-art studies it is measured against, so the frozen-model scores are directly comparable to published fine-tuned baselines.","core_discovery":"The central claim is empirical: with weights frozen, today's general-purpose audio foundation models are already competitive with state-of-the-art fine-tuned results on clean auscultation data, and the quality of their frozen representations largely determines task performance. On SPRS, BEATs and M2D reach or approach the best published fine-tuned scores, and on BMD-HS, AST, BEATs, and M2D exceed the original baseline's accuracy and F1. On ICBHI2017 and CirCor, where recordings are noisy, none of the frozen models reaches the SOTA results, though M2D comes closest on CirCor. The paper further reports that OPERA-CT, a respiratory-sound-specific foundation model, underperforms general audio models on respiratory tasks, and attributes this to its smaller and less diverse pretraining corpus (140K monotonic respiratory sounds versus 2M AudioSet samples).","pith_inferences":["Extension: if the clean-versus-noisy pattern holds beyond these four datasets, recording quality, not domain match, is the main predictor of whether an audio foundation model can be used frozen; a universal denoising front-end could be the highest-leverage next component.","Extension: the BMD-HS result rests on a binary murmur-presence/absence redefinition with new stratified splits; the original baseline was not rerun under those exact conditions, so a strict head-to-head on identical splits is still an open check.","Extension: OPERA-CT's shortfall should not be read as evidence that domain-specific pretraining is useless at any scale; it only shows that 140K respiratory samples do not beat 2M diverse AudioSet samples with current methods.","Extension: a testable extension is to add denoising or sound separation before frozen encoding and measure whether ICBHI2017 and CirCor scores close the SOTA gap."],"forward_implications":["For clean stethoscope or microphone recordings, fixed-weight general-purpose audio encoders can be deployed without per-task fine-tuning and still match the best published results.","On noisy real-life recordings such as ICBHI2017 and CirCor, frozen features are not enough; fine-tuning or a denoising/target-sound-extraction front-end is needed.","Pretraining on a large, diverse general audio corpus is more valuable for respiratory tasks than pretraining on a narrow respiratory-only corpus of similar scale.","Recent self-supervised pretraining methods (M2D, BEATs) provide more transferable representations than earlier ones (AST, BYOL-A) in these medical audio tasks.","The released evaluation code offers a common benchmark for future audio foundation models to report frozen-feature performance."],"supporting_citations":[{"why":"Establishes the fine-tuned SOTA baselines for foundation models on heart murmur detection and supplies the codebase and experimental setup reused for CirCor.","marker":"[6]"},{"why":"Defines OPERA-CT, the respiratory-specific foundation model and benchmark that the paper compares against general audio models.","marker":"[2]"},{"why":"Supplies the ICBHI2017 respiratory sound dataset, official splits, and score metric for the first task.","marker":"[5]"},{"why":"Provides the baseline fine-tuned respiratory representations (CE and SCL) that frozen models are compared against on ICBHI2017 and SPRS.","marker":"[14]"},{"why":"Provides the M2D foundation model, the SPRS codebase and settings, and the fine-tuned SOTA results on both respiratory tasks.","marker":"[11]"},{"why":"Supplies BEATs, one of the evaluated frozen general-purpose models, and its AudioSet-based pretraining method.","marker":"[12]"},{"why":"Defines the CirCor dataset and the heart murmur classification task with weighted accuracy as the official metric.","marker":"[17]"},{"why":"Supplies the BMD-HS dataset and its baseline accuracy, sensitivity, specificity, and F1 that frozen models outperform.","marker":"[21]"}],"fun_headline_variants":["Frozen audio models match SOTA on clean heart and lung sounds","Off-the-shelf audio models beat respiratory-specific model on clean data","Frozen general-purpose audio beats fine-tuned respiratory model on clean auscultation","Clean data: frozen audio models rival fine-tuning; noisy data needs tuning","General audio encoders outperform respiratory-specific model on heart and lung tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the BMD-HS comparison is valid: the paper converts a multi-disease dataset into a binary murmur-presence/absence task with new stratified splits and treats the resulting scores as comparable to the original baseline, without showing that baseline used the same task definition and splits; if that comparability fails, the clean-data heart sound result is not established.","fun_headline_variants_meta":{"raw":{"variants":["Frozen audio models match SOTA on clean heart and lung sounds","Off-the-shelf audio models beat respiratory-specific model on clean data","Frozen general-purpose audio beats fine-tuned respiratory model on clean auscultation","Clean data: frozen audio models rival fine-tuning; noisy data needs tuning","General audio encoders outperform respiratory-specific model on heart and lung tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000863,"raw_usage":{"total_tokens":3710,"prompt_tokens":882,"completion_tokens":2828,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":2733}},"tokens_in":498,"tokens_out":2828,"duration_ms":18131,"temperature":1.0,"reasoning_tokens":2733,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:26:50.655310+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the BMD-HS evaluation with the dataset authors' original multi-class task definition and original train/test split, using the same frozen models; if AST, BEATs, and M2D no longer exceed the baseline's accuracy and F1, the clean-data heart sound claim fails. A second check: apply a denoiser to ICBHI2017 before frozen M2D feature extraction; the noise-based explanation predicts the score should move toward 61.2.","supporting_citations":[{"cited_title":"Exploring pre-trained general-purpose audio representations for heart murmur detection,","cited_arxiv_id":null,"evidence_quote":"Establishes the fine-tuned SOTA baselines for foundation models on heart murmur detection and supplies the codebase and experimental setup reused for CirCor."},{"cited_title":"An open access database for the evaluation of respiratory sound classification algorithms,","cited_arxiv_id":null,"evidence_quote":"Supplies the ICBHI2017 respiratory sound dataset, official splits, and score metric for the first task."},{"cited_title":"Pretraining Respiratory Sound Rep- resentations using Metadata and Contrastive Learning,","cited_arxiv_id":null,"evidence_quote":"Provides the baseline fine-tuned respiratory representations (CE and SCL) that frozen models are compared against on ICBHI2017 and SPRS."},{"cited_title":"Masked Modeling Duo: Towards a Universal Audio Pre- Training Framework,","cited_arxiv_id":null,"evidence_quote":"Provides the M2D foundation model, the SPRS codebase and settings, and the fine-tuned SOTA results on both respiratory tasks."},{"cited_title":"Heart murmur detection from phonocardiogram recordings: The George B. Moody PhysioNet Challenge 2022,","cited_arxiv_id":null,"evidence_quote":"Defines the CirCor dataset and the heart murmur classification task with weighted accuracy as the official metric."}],"review_version":1}