Pith. sign in

REVIEW 2 cited by

BTS: Bridging Text and Sound Modalities for Metadata-Aided Respiratory Sound Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.06786 v2 pith:542T6VWT submitted 2024-06-10 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords soundmetadatarespiratorymodelperformancerecordingclassificationmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Respiratory sound classification (RSC) is challenging due to varied acoustic signatures, primarily influenced by patient demographics and recording environments. To address this issue, we introduce a text-audio multimodal model that utilizes metadata of respiratory sounds, which provides useful complementary information for RSC. Specifically, we fine-tune a pretrained text-audio multimodal model using free-text descriptions derived from the sound samples' metadata which includes the gender and age of patients, type of recording devices, and recording location on the patient's body. Our method achieves state-of-the-art performance on the ICBHI dataset, surpassing the previous best result by a notable margin of 1.17%. This result validates the effectiveness of leveraging metadata and respiratory sound samples in enhancing RSC performance. Additionally, we investigate the model performance in the case where metadata is partially unavailable, which may occur in real-world clinical setting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Disentangling Dual-Encoder Masked Autoencoder for Respiratory Sound Classification

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A dual-encoder masked autoencoder with time-shuffle Siamese and vCLUB mutual-information losses reaches an ICBHI average score of 61.50 and the highest sensitivity among compared respiratory sound classifiers.

  2. Lungmix: A Mixup-Based Strategy for Generalization in Respiratory Sound Classification

    cs.SD 2024-12 conditional novelty 5.0 of 10

    Lungmix, a mixup variant with loudness masks and semantic OR label interpolation, improves some cross-dataset respiratory sound classification scores by up to 3.55 points.

Pith tools