JAVEdit-100k is the first large-scale dataset for instruction-guided joint audio-visual video editing, accompanied by JAVEditBench and the JAVEdit model that outperforms baselines on five of six metrics.
The T05 System for the VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
ASR self-verification via best-of-N sampling eliminates observed catastrophic failures in multiple neural-codec TTS models, with distillation transferring most of the robustness to single-shot decoding.
Emotional prosody in LM-TTS localizes to the x-vector by elimination, where centroid arithmetic on speaker embeddings yields training-free cross-lingual gains of +0.29 emotion2vec cosine on English and +0.09 on Brazilian Portuguese while preserving identity.
citing papers explorer
-
JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation
JAVEdit-100k is the first large-scale dataset for instruction-guided joint audio-visual video editing, accompanied by JAVEditBench and the JAVEdit model that outperforms baselines on five of six metrics.
-
Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs
ASR self-verification via best-of-N sampling eliminates observed catastrophic failures in multiple neural-codec TTS models, with distillation transferring most of the robustness to single-shot decoding.
-
Task-Vector Arithmetic for Emotional Expressivity Control in Language-Model-Based Text-to-Speech
Emotional prosody in LM-TTS localizes to the x-vector by elimination, where centroid arithmetic on speaker embeddings yields training-free cross-lingual gains of +0.29 emotion2vec cosine on English and +0.09 on Brazilian Portuguese while preserving identity.