NAC adapts multi-scale RVQGAN audio codecs with kinematic-specific losses to produce ordered action tokens that yield lower reconstruction error and higher task success than prior tokenizers in VLA models.
hub Canonical reference
Singing voice graph modeling for singfake detection
Canonical reference. 71% of citing Pith papers cite this work as background.
hub tools
citation-role summary
citation-polarity summary
roles
background 7representative citing papers
MMDG-Bench provides unified protocols and ten baselines for multimodal domain generalization, showing structured DG-MML combinations often outperform prior methods with insights on framework choice and backbone effects.
MusicDET models the distribution of real music features with frequency-guided normalizing flows to detect AI-generated music as out-of-distribution samples in a zero-shot setting.
ProtoSSL discovers generalizable prototypes from unlabeled time-series via self-supervision and assigns them to new tasks for interpretable predictions, outperforming supervised baselines in low-data regimes on ECG datasets.
X-VC achieves zero-shot streaming voice conversion via one-step codec-space conversion with dual-conditioning acoustic converter and role-assignment training on generated paired data.
SignRecGAN trains on separate sign and speech datasets via adversarial and reconstruction objectives to inject sign-derived prosody into TTS output using the S2PFormer model.
AVTok is a unified tokenizer that converts audio-video pairs into a compact 1D latent representation via dual-stream transformer and hierarchical training for improved reconstruction and cross-modal generation.
MJ EPA applies a single shared ViT encoder and one predictive objective within and across audio-visual modalities, reporting >6.8 mAP gains on AudioSet-20K and competitive video results with 10x less data.
Emotional prosody in LM-TTS localizes to the x-vector by elimination, where centroid arithmetic on speaker embeddings yields training-free cross-lingual gains of +0.29 emotion2vec cosine on English and +0.09 on Brazilian Portuguese while preserving identity.
SpeakerLLM unifies speaker profiling, recording-condition understanding, and structured verification reasoning in an audio-LLM via a hierarchical tokenizer and decision traces.
Alethia is a pretrained audio encoder using continuous embedding prediction and generative flow-matching reconstruction that outperforms existing speech foundation models on voice deepfake tasks with better robustness and zero-shot generalization.
PsyGAT structures conversations as dynamic temporal graphs with Psychological Expression Units and persona augmentation to reach state-of-the-art Macro F1 scores of 89.99 and 71.37 on DAIC-WoZ and E-DAIC while adding causal interpretability.
SenSE adds language-model semantic guidance to flow-matching generative speech enhancement via a dual-path masked conditioning strategy and reports SOTA results on distorted speech.
FAST applies discrete cosine transform to robot action sequences for efficient tokenization, enabling autoregressive VLAs to succeed on high-frequency dexterous tasks and scale to 10k hours of data while matching diffusion VLA performance with up to 5x faster training.
MRAF framework uses missing-token prompting and reliability-aware cross-attention fusion to achieve 100% accuracy on some POLY-SIM 2026 tasks and competitive results on missing-face cases.
HATS supplies human side-by-side preference judgments on ASR transcripts to measure correlation with lexical and embedding-based evaluation metrics.
Woosh is a new publicly released foundation model optimized for high-quality sound effect generation from text or video, showing competitive or better results than open alternatives like Stable Audio Open.
A factorized log-linear model (FLiP) recovers 75-80% of lexical content from sentence embeddings and reveals English-centric biases in SONAR, LaBSE, and Gemini encoders.
A framework using hearing-loss simulation in young listeners and the GESI objective measure shows that central and cognitive factors contribute to speech intelligibility differences in older adults beyond peripheral hearing loss.
IPA use disrupts content-generation writing tasks more than copying tasks because they share more cognitive resources.
Merging one audio and three video deepfake detectors with voting-based fusion gives ~70–73% accuracy on FakeAVCeleb, with the audio branch performing at chance in the wild.
AT-ADD introduces standardized tracks and datasets for evaluating audio deepfake detectors on speech under real-world conditions and on diverse unknown audio types to promote generalization beyond speech-centric methods.
The survey reviews the evolution of accent conversion from early DSP approaches to neural models, situating them in linguistic foundations and highlighting constraints, datasets, evaluations, and future directions.
citing papers explorer
-
NAC: Neural Action Codec for Vision-Language-Action Models
NAC adapts multi-scale RVQGAN audio codecs with kinematic-specific losses to produce ordered action tokens that yield lower reconstruction error and higher task success than prior tokenizers in VLA models.
-
MMDG-Bench: A Benchmark for Multimodal Domain Generalization
MMDG-Bench provides unified protocols and ten baselines for multimodal domain generalization, showing structured DG-MML combinations often outperform prior methods with insights on framework choice and backbone effects.
-
MusicDET: Zero-Shot AI-Generated Music Detection
MusicDET models the distribution of real music features with frequency-guided normalizing flows to detect AI-generated music as out-of-distribution samples in a zero-shot setting.
-
ProtoSSL: Interpretable Prototype Learning from Unlabeled Time-Series Data
ProtoSSL discovers generalizable prototypes from unlabeled time-series via self-supervision and assigns them to new tasks for interpretable predictions, outperforming supervised baselines in low-data regimes on ECG datasets.
-
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
X-VC achieves zero-shot streaming voice conversion via one-step codec-space conversion with dual-conditioning acoustic converter and role-assignment training on generated paired data.
-
Sign-to-Speech Prosody Transfer via Sign Reconstruction-based GAN
SignRecGAN trains on separate sign and speech datasets via adversarial and reconstruction objectives to inject sign-derived prosody into TTS output using the S2PFormer model.
-
AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation
AVTok is a unified tokenizer that converts audio-video pairs into a compact 1D latent representation via dual-stream transformer and hierarchical training for improved reconstruction and cross-modal generation.
-
MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning
MJ EPA applies a single shared ViT encoder and one predictive objective within and across audio-visual modalities, reporting >6.8 mAP gains on AudioSet-20K and competitive video results with 10x less data.
-
Task-Vector Arithmetic for Emotional Expressivity Control in Language-Model-Based Text-to-Speech
Emotional prosody in LM-TTS localizes to the x-vector by elimination, where centroid arithmetic on speaker embeddings yields training-free cross-lingual gains of +0.29 emotion2vec cosine on English and +0.09 on Brazilian Portuguese while preserving identity.
-
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
SpeakerLLM unifies speaker profiling, recording-condition understanding, and structured verification reasoning in an audio-LLM via a hierarchical tokenizer and decision traces.
-
Alethia: A Foundational Encoder for Voice Deepfakes
Alethia is a pretrained audio encoder using continuous embedding prediction and generative flow-matching reconstruction that outperforms existing speech foundation models on voice deepfake tasks with better robustness and zero-shot generalization.
-
Psychologically-Grounded Graph Modeling for Interpretable Depression Detection
PsyGAT structures conversations as dynamic temporal graphs with Psychological Expression Units and persona augmentation to reach state-of-the-art Macro F1 scores of 89.99 and 71.37 on DAIC-WoZ and E-DAIC while adding causal interpretability.
-
SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement
SenSE adds language-model semantic guidance to flow-matching generative speech enhancement via a dual-path masked conditioning strategy and reports SOTA results on distorted speech.
-
FAST: Efficient Action Tokenization for Vision-Language-Action Models
FAST applies discrete cosine transform to robot action sequences for efficient tokenization, enabling autoregressive VLAs to succeed on high-frequency dexterous tasks and scale to 10k hours of data while matching diffusion VLA performance with up to 5x faster training.
-
Missing-Token Prompted Reliability-Aware Fusion for Robust Polyglot Speaker Identification
MRAF framework uses missing-token prompting and reliability-aware cross-attention fusion to achieve 100% accuracy on some POLY-SIM 2026 tasks and competitive results on missing-face cases.
-
HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics
HATS supplies human side-by-side preference judgments on ASR transcripts to measure correlation with lexical and embedding-based evaluation metrics.
-
Woosh: A Sound Effects Foundation Model
Woosh is a new publicly released foundation model optimized for high-quality sound effect generation from text or video, showing competitive or better results than open alternatives like Stable Audio Open.
-
FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings
A factorized log-linear model (FLiP) recovers 75-80% of lexical content from sentence embeddings and reveals English-centric biases in SONAR, LaBSE, and Gemini encoders.
-
Disentangling peripheral hearing loss from central and cognitive effects on speech intelligibility in older adults
A framework using hearing-loss simulation in young listeners and the GESI objective measure shows that central and cognitive factors contribute to speech intelligibility differences in older adults beyond peripheral hearing loss.
-
Multitasking with Alexa Multitasking with Alexa: How Using Intelligent Personal Assistants Impacts Language-based Primary Task Performance
IPA use disrupts content-generation writing tasks more than copying tasks because they share more cognitive resources.
-
Ensemble Deep Learning Approaches for AI-Altered Video Detection
Merging one audio and three video deepfake detectors with voting-based fusion gives ~70–73% accuracy on FakeAVCeleb, with the audio branch performing at chance in the wild.
-
AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan
AT-ADD introduces standardized tracks and datasets for evaluating audio deepfake detectors on speech under real-world conditions and on diverse unknown audio types to promote generalization beyond speech-centric methods.
-
Accent Conversion: A Problem-Driven Survey of Sociolinguistic and Technical Constraints
The survey reviews the evolution of accent conversion from early DSP approaches to neural models, situating them in linguistic foundations and highlighting constraints, datasets, evaluations, and future directions.