PolySpeech-100 is a new benchmark for native-level speech comprehension across 110 linguistic variants that evaluates 22 models and reports E2E advantages on dialects, robustness gaps on low-resource languages, and degradation from Chain-of-Thought prompting.
Small-footprint keyword spotting using deep neural networks
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
On chemistry SMILES with a fixed 165-token base, BPE and Unigram-LM produce near-disjoint vocabularies (Jaccard ≤0.161) and Unigram-LM emits 29–41% more tokens across 22 matched conditions.
A distribution-alignment framework for unpaired cross-modal knowledge distillation with theoretical guarantees on feature and label alignment.
FedMChain improves multimodal federated learning by chaining modality-wise optimization phases with error-compensated regularization and sparse sign-guided aggregation to mitigate modality competition and cut communication overhead.
APEX generates four types of prototype-based explanations for pre-trained audio classifiers that preserve output invariance and target acoustic properties better than gradient methods applied to spectrograms.
A framework using covariance-based spectral signatures and TreeSHAP attributions on AASIST3 branches identifies four operational archetypes and a flawed specialization mode that explains high error rates on specific spoofing attacks.
A new open-source pipeline creates a 5,078-hour Russian speech dataset with prosody annotations, and VITS and SEMamba models trained on it outperform those trained on existing Russian corpora under equalized budgets.
GS-NFS accelerates dynamic 3DGS encoding and decoding by 1-2 orders of magnitude on GPU while maintaining competitive compression ratios and rendering quality.
BMRUs enable analog recurrent neural network hardware via discrete outputs that suppress noise 20-fold, with one-to-one parameter-to-circuit mapping and linear power scaling for recurrence.
citing papers explorer
-
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects
PolySpeech-100 is a new benchmark for native-level speech comprehension across 110 linguistic variants that evaluates 22 models and reports E2E advantages on dialects, robustness gaps on low-resource languages, and degradation from Chain-of-Thought prompting.
-
Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES
On chemistry SMILES with a fixed 165-token base, BPE and Unigram-LM produce near-disjoint vocabularies (Jaccard ≤0.161) and Unigram-LM emits 29–41% more tokens across 22 matched conditions.
-
Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and Algorithm
A distribution-alignment framework for unpaired cross-modal knowledge distillation with theoretical guarantees on feature and label alignment.
-
Boosting Multimodal Federated Learning via Chained Modality Optimization
FedMChain improves multimodal federated learning by chaining modality-wise optimization phases with error-compensated regularization and sparse sign-guided aggregation to mitigate modality competition and cut communication overhead.
-
APEX: Audio Prototype EXplanations for Classification Tasks
APEX generates four types of prototype-based explanations for pre-trained audio classifiers that preserve output invariance and target acoustic properties better than gradient methods applied to spectrograms.
-
Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance
A framework using covariance-based spectral signatures and TreeSHAP attributions on AASIST3 branches identifies four operational archetypes and a flawed specialization mode that explains high error rates on specific spoofing attacks.
-
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
A new open-source pipeline creates a 5,078-hour Russian speech dataset with prosody annotations, and VITS and SEMamba models trained on it outperform those trained on existing Russian corpora under equalized budgets.
-
GS-NFS: Bandwidth-adaptive Streaming of Dynamic Gaussian Splats and Point Clouds
GS-NFS accelerates dynamic 3DGS encoding and decoding by 1-2 orders of magnitude on GPU while maintaining competitive compression ratios and rendering quality.
-
Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations
BMRUs enable analog recurrent neural network hardware via discrete outputs that suppress noise 20-fold, with one-to-one parameter-to-circuit mapping and linear power scaling for recurrence.