Introduces MAD2 benchmark and shows dialogue context improves multimodal spoken claim verification, with conversational structure mattering more than framing.
WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
5 Pith papers cite this work, alongside 239 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 5verdicts
UNVERDICTED 5roles
method 2polarities
use method 2representative citing papers
SuperMemory-VQA provides 4,853 human-verified QA pairs from 52.9 hours of egocentric AI glasses recordings to benchmark AI systems on realistic long-horizon memory tasks including an unanswerable option.
Transformer-derived sentiment features from therapy sessions correlate with emotional-valence components of the OQ-45 and differ significantly between patients identified as at risk of deterioration by rational and empirical outcome models.
WorldSpeech supplies 65k hours of multilingual aligned speech data across 76 languages and delivers 63.5% average relative WER reduction after fine-tuning ASR models on 11 typologically diverse languages.
A cascaded SimulST system using Parakeet and Qwen 3.5 with adaptive black-box policies and RAG context achieves +5.82 XCOMET-XL improvement on En→De for IWSLT 2026.
citing papers explorer
-
Context-Aware Multimodal Claim Verification in Spoken Dialogues
Introduces MAD2 benchmark and shows dialogue context improves multimodal spoken claim verification, with conversational structure mattering more than framing.
-
SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory
SuperMemory-VQA provides 4,853 human-verified QA pairs from 52.9 hours of egocentric AI glasses recordings to benchmark AI systems on realistic long-horizon memory tasks including an unanswerable option.
-
The Association of Transformer-based Sentiment Analysis with Symptom Distress and Deterioration in Routine Psychotherapy Care
Transformer-derived sentiment features from therapy sessions correlate with emotional-valence components of the OQ-45 and differ significantly between patients identified as at risk of deterioration by rational and empirical outcome models.
-
WorldSpeech: A Multilingual Speech Corpus from Around the World
WorldSpeech supplies 65k hours of multilingual aligned speech data across 76 languages and delivers 63.5% average relative WER reduction after fine-tuning ASR models on 11 typologically diverse languages.
-
MLLP-VRAIN UPV system for the IWSLT 2026 Simultaneous Speech Translation task
A cascaded SimulST system using Parakeet and Qwen 3.5 with adaptive black-box policies and RAG context achieves +5.82 XCOMET-XL improvement on En→De for IWSLT 2026.