Pith. sign in

REVIEW 1 cited by

MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.13886 v2 pith:YZ7RCBKN submitted 2020-10-26 eess.AS cs.SD

classification eess.AScs.SD
keywords marblenetnetworkactivitydeepdetectionneuralseparabletime-channel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present MarbleNet, an end-to-end neural network for Voice Activity Detection (VAD). MarbleNet is a deep residual network composed from blocks of 1D time-channel separable convolution, batch-normalization, ReLU and dropout layers. When compared to a state-of-the-art VAD model, MarbleNet is able to achieve similar performance with roughly 1/10-th the parameter cost. We further conduct extensive ablation studies on different training methods and choices of parameters in order to study the robustness of MarbleNet in real-world VAD tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment

    eess.AS 2026-07 conditional novelty 6.0 of 10

    kiloVAD, a 2.1k-parameter causal CNN VAD on standard Mel features, reaches 0.850 AUC on AVA-Speech and beats standard QAT by 1–4% at INT4 via angle-based self-distillation.

Pith tools