Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:46:12.269019Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 1 inbound Pith citation observation for arXiv:2502.10447.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:46:12.269019Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:34.111978Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T15:39:34.251185Z
84 of 84 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cb00bbcc-96c2-42e8-b779-f4adb89613c5 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a90e81-c9e9-459c-87f8-006a926ba1dc · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d4d530-6232-44cd-a438-e5cf4c89d29e · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition S., Senior, A., Vinyals, O., and Zisserman, A
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09235db7-3652-4033-92ad-731095d6de81 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e695b4c-6200-41a6-b137-6a12fb193dd7 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6caeebfe-07b7-4133-9379-0b036559bba2 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d56858b-ac09-4ac0-ad31-d73f45facc6c · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Xls-r: Self-supervised cross-lingual speech representation learning at scale
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3dbff9a2-7450-40d3-b556-096fbdcda2bd · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition wav2vec 2.0: A framework for self-supervised learning of speech representations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f50549-e5f2-4082-84c2-aa4a81d73ecb · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b7e83e-c38c-4e3a-8a4a-ccc9451e40c4 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Timofte, R
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d21efac-daea-48f0-a0b4-23f51b2e50e2 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Large Language Models are Strong Audio-Visual Speech Recognition Learners
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd483a05-4471-44fb-91fb-0d045342d569 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 94c28bb5-c138-44f7-9c7c-6d1910553bdd · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0313e3cd-1a74-4706-b120-33ba3729349d · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mixtures of experts for audio-visual learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3bee3c5-68f5-4dab-b490-218a46cd56b7 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Self-supervised learning with random-projection quantizer for speech recognition
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8b5cf09a-2a9a-4c10-b50a-b937abbd56ee · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition J., Kim, M., and Ro, Y
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 10622f6a-273b-455e-9a13-874523981cac · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition S., Nagrani, A., and Zisserman, A
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e64d104-7cd0-4ae7-8d51-06a432a874d7 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unified scaling laws for routed language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5dd41cc2-e5c3-4328-aee0-e1070ad4bbc5 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Stablemoe: Stable routing strategy for mixture of experts
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0405de2f-f8a3-485f-94b7-c584ad2fa1ec · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A study of dropout-induced modality bias on robustness to missing video frames for audio-visual speech recognition
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5fa66c23-3834-481f-b695-ae71b99e2503 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Luettin, J
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 54992c07-7d83-4e18-8e1b-8fd801d9fcd0 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., and Matt, P
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eb738fe9-acd8-45a5-8078-6693ed7466bd · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b727f062-e417-4fa9-9016-14e32e64f9ff · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Boosting speech recognition robustness to modality-distortion with contrast-augmented prompts
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 96a8fece-b9ff-4f10-ac66-9350ea481c97 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Conformer: Convolution-augmented transformer for speech recognition
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0386b6f4-63c0-4444-8e36-a72ec75d3e59 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f54ce98-8b96-4709-a4c5-b81514cd5cea · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Jointly learning visual and auditory speech representations from raw data
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 55a73f4d-4585-42da-9190-fea994ad50c4 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Braven: Improving self-supervised pre-training for visual and auditory speech recognition
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d65721a7-5e49-4b3c-9e82-243a295f819e · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition XLAVS - R : Cross-lingual audio-visual speech representation learning for noise-robust speech perception
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c4f2e9f-e739-4545-a79d-fa40926bb734 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1dd6c8d1-c13f-468b-987a-f6e4568ca8da · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3875f7e9-9ea6-48d8-bd22-2b184af8a91f · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Shi, B
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac8f7845-fe02-42d2-83cf-0fa56b0f684d · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6de3fe-b46c-4d9f-9d0c-4bcc01d00591 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition N., Zhang, Y., and Beaufays, F
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 976eb8dc-4005-4619-8d18-2970d5df3951 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8bddd93f-c46e-47c0-93bc-95c37467a604 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3cc7802d-eaff-464b-b122-a172b1f036fb · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 83791601-00d5-48e6-b597-166a6fcf1351 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A., Jordan, M
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c769899-7176-4dfd-a205-9ec91ae2553b · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mixtral of Experts
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b155f64-0f2f-48f0-9979-835e89202665 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3c730175-c11c-4cf6-8f99-b4b2550e4145 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Scaling Laws for Neural Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9ba568c-e35b-46a9-a66d-5c63c2c32c62 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning video temporal dynamics with cross-modal attention for robust audio-visual speech recognition
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eb92f129-9c51-4e93-9443-b911ae612b24 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Multi-task corrupted prediction for learning robust audio-visual speech representation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e50f65bf-d38b-4a7a-95f0-414b568881b1 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Adam: A Method for Stochastic Optimization
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d557243-3fe3-4ea7-a398-006108a0781c · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Moai: Mixture of all intelligence for large language and vision models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba99a6f1-d267-4102-9903-c77f037e8c06 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition \ GS \ hard: Scaling giant models with conditional computation and automatic sharding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ad298ad-4c4a-4eb7-b2e3-256862aa31fd · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unified cross-modal attention: Robust audio-visual speech recognition and beyond
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 982ffbce-c590-48c5-981d-0ec2a7725ecf · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Pace: Unified multi-modal dialogue pre-training with progressive and compositional experts
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7ee4fd34-f493-45cf-93ec-be559e2e05e9 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e1e6b9f-0e75-4527-98d2-32b203226ecd · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4d5394ea-e862-476a-9d68-7db199733b31 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76554e89-2cc7-415b-9f11-c51a54917429 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., and Pantic, M
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation decb22d6-5180-42cb-b35a-4c68ea23fd0b · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition End-to-end audio-visual speech recognition with conformers
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b6e30b98-4524-45fd-ba64-55921ab1de61 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a4dd371-89e0-49d0-8b75-82ee9dfb957c · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Recurrent neural network transducer for audio-visual speech recognition
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 40f518dc-ba49-4251-969c-5dbce41495e0 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mm1: methods, analysis and insights from multimodal llm pre-training
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 736a4236-703c-4be2-8c3b-53dfb9d93e9c · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Multimodal contrastive learning with limoe: the language-image mixture of experts
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 262ab514-50ef-47a2-abb3-201a5b8f83c8 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition G., and Ogata, T
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a554eb4e-ae2a-4261-95fa-43ec6a0329ce · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3a623aaa-0e13-431f-8564-fb3387ff6617 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Bleu: a method for automatic evaluation of machine translation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c91ae05-3531-44c0-99b7-4a1e6ee84503 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A call for clarity in reporting bleu scores
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation faf4f2ff-e926-49b8-9ebe-e85bfd014e25 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Lipsound2: Self-supervised pre-training for lip-to-speech reconstruction and lip reading
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b11b7af-4db4-4e2b-92e6-e9953ef6fad7 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e558eeb-03d6-4dde-9b12-1e0885aa2506 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning from the master: Distilling cross-modal advanced knowledge for lip reading
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e5f9b8f4-3539-49df-ba88-92dbcfee486b · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition wav2vec: Unsupervised pre-training for speech recognition
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6fe91de-707b-404d-b25d-68d2f3a1f931 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition H., Nagrani, A., and Schmid, C
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93c99fbf-3002-4c1e-bd85-c855be4c7cd0 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 698d2898-de93-452b-a09a-31664d927a91 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Scaling vision-language models with sparse mixture of experts
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 761244b3-694b-40ea-b43e-498b303e25b9 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning audio-visual speech representation by masked multimodal cluster prediction
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6455672e-f8dc-4faa-beec-aee8ca7f209d · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Robust self-supervised audio-visual speech recognition
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e469726c-a996-4160-b4cf-4052afa9dfe9 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition MUSAN: A Music, Speech, and Noise Corpus
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea26ca4a-a117-4a82-87d1-9e3eb8a279eb · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e6dd0a-b44f-498d-b5c3-5266d235f667 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition N., Kaiser, ., and Polosukhin, I
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683256ca-aa77-47d4-9a5f-4f8c754714e0 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition T., and Li, H
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e95febc-e8b0-4589-b53d-916f1e2398be · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Language-routing mixture of experts for multilingual and code-switching speech recognition
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 555f1a1d-5da3-4299-9a20-b307e07dc4fe · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition R., and Hayashi, T
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 962a535d-9a14-443b-9c67-34f771193ca1 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Robust audiovisual speech recognition models with mixture-of-experts
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3584682c-30ff-4622-9c96-728984916f11 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d8f1c0d4-9dd2-4fcb-81c8-501a463e55d6 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Speechmoe2: Mixture-of-experts model with improved routing
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 08b33f54-e9a5-4b79-b0ed-562a83643f94 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Visual hallucination elevates speech recognition
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5b13026c-fa99-4c0d-970f-b9f68f4aff78 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Self-supervised audio-visual speech representations learning by multimodal self-distillation
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3515f5d7-0b53-4273-bafd-96cd12610c8b · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Uni-perceiver-moe: Learning sparse generalist models with conditional moes
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0a2625e8-12f8-4dad-8059-86cde6accefa · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3cea8a1-ca54-4a16-b6fc-d1420e753587 · outbound
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition ST-MoE: Designing Stable and Transferable Sparse Expert Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10189419-0fe1-4ad1-9f36-c209cea88af3 · inbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.