ORCA splits Q-Former queries into orthogonally constrained groups, reversing directional collapse and speaker-indistinguishability in audio-LLM connectors and gaining 26.4 points on SAKURA multi-hop reasoning.
Robust speech recognition via large-scale weak supervision
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
TimePro-RL interleaves timestamp embeddings in audio sequences and applies RL post-SFT to boost temporal alignment in LALMs, yielding gains on grounding, event detection, and dense captioning.
A multimodal model fuses Whisper acoustic embeddings with LLM-extracted linguistic features via gated fusion to achieve F1 scores of 89.47% and 90.14% on ADReSS and ADReSSo dementia detection benchmarks.
citing papers explorer
-
Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models
ORCA splits Q-Former queries into orthogonally constrained groups, reversing directional collapse and speaker-indistinguishability in audio-LLM connectors and gaining 26.4 points on SAKURA multi-hop reasoning.
-
Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt
TimePro-RL interleaves timestamp embeddings in audio sequences and applies RL post-SFT to boost temporal alignment in LALMs, yielding gains on grounding, event detection, and dense captioning.
-
Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection
A multimodal model fuses Whisper acoustic embeddings with LLM-extracted linguistic features via gated fusion to achieve F1 scores of 89.47% and 90.14% on ADReSS and ADReSSo dementia detection benchmarks.