LeBenchmark 2.0: a Standardized, Replicable and Enhanced Frame- work for Self-supervised Representations of French Speech,

· 2024 · arXiv 2309.05472

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

representative citing papers

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users

eess.AS · 2026-06-19 · unverdicted · novelty 6.0

An end-to-end SLU architecture with frozen SSL acoustic encoder, LSTM classification head, and cross-modal distillation achieves 93% accuracy on simple commands and 82% on spontaneous speech at 7 ms latency on the new VoiceStick corpus, outperforming cascade baselines.

Assessing the Energy and Carbon Emissions of Neural Speaker Verification Model in Training and Inference

cs.SD · 2026-06-06 · unverdicted · novelty 3.0

Empirical study finds diminishing accuracy returns against steep energy growth for deeper and wider ResNet speaker verification models on VoxCeleb2.

citing papers explorer

Showing 2 of 2 citing papers.

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users eess.AS · 2026-06-19 · unverdicted · none · ref 45
An end-to-end SLU architecture with frozen SSL acoustic encoder, LSTM classification head, and cross-modal distillation achieves 93% accuracy on simple commands and 82% on spontaneous speech at 7 ms latency on the new VoiceStick corpus, outperforming cascade baselines.
Assessing the Energy and Carbon Emissions of Neural Speaker Verification Model in Training and Inference cs.SD · 2026-06-06 · unverdicted · none · ref 19
Empirical study finds diminishing accuracy returns against steep energy growth for deeper and wider ResNet speaker verification models on VoxCeleb2.

LeBenchmark 2.0: a Standardized, Replicable and Enhanced Frame- work for Self-supervised Representations of French Speech,

fields

years

verdicts

representative citing papers

citing papers explorer