Pith. sign in

REVIEW 1 cited by

Improving Weakly Supervised Sound Event Detection with Causal Intervention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.05678 v1 pith:IYHHKIUY submitted 2023-03-10 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords causaleventsoundclip-levelco-occurrenceconfounderscontextdetection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing weakly supervised sound event detection (WSSED) work has not explored both types of co-occurrences simultaneously, i.e., some sound events often co-occur, and their occurrences are usually accompanied by specific background sounds, so they would be inevitably entangled, causing misclassification and biased localization results with only clip-level supervision. To tackle this issue, we first establish a structural causal model (SCM) to reveal that the context is the main cause of co-occurrence confounders that mislead the model to learn spurious correlations between frames and clip-level labels. Based on the causal analysis, we propose a causal intervention (CI) method for WSSED to remove the negative impact of co-occurrence confounders by iteratively accumulating every possible context of each class and then re-projecting the contexts to the frame-level features for making the event boundary clearer. Experiments show that our method effectively improves the performance on multiple datasets and can generalize to various baseline models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Normal Patterns in Musical Loops

    cs.SD 2025-05 reject novelty 4.0 of 10

    A Deep SVDD model using HTS-AT and feature fusion learns normal patterns in variable-length bass and guitar loops, with residual connections improving the learned latent space.

Pith tools