Pith. sign in

Paper Citation Record · LEDGER

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

As of 9 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 1 inbound Pith citation observation for arXiv:2502.10447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10447 v2

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:46:12.269019Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:34.111978Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:39:34.251185Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cb00bbcc-96c2-42e8-b779-f4adb89613c5 · outbound

This paper cites write newline.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.941722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.941722Z digest=sha256:bcc03e2eeddc3f2de35e998a7f0a0abfc6b051c594245d1bab417723c444b0e2

Observation f8a90e81-c9e9-459c-87f8-006a926ba1dc · outbound

This paper cites GPT-4 Technical Report.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.947979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.947979Z digest=sha256:bf7af2f5ad777d8b16c8b3de09f4e7faa25f9349bb23371f5a6fc878ae2d32e7

Observation 78d4d530-6232-44cd-a438-e5cf4c89d29e · outbound

This paper cites S., Senior, A., Vinyals, O., and Zisserman, A.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition S., Senior, A., Vinyals, O., and Zisserman, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.952950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.952950Z digest=sha256:aced09dff8289bf8437814b4005e098a429a78c507bf77070dca68ceb5acbd30

Observation 09235db7-3652-4033-92ad-731095d6de81 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.957891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.957891Z digest=sha256:23fca58f978cd79317d4411347c4e7c75c21a8c11a625ef59cea97e3c75a22a0

Observation 2e695b4c-6200-41a6-b137-6a12fb193dd7 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.963258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.963258Z digest=sha256:9d3f78b307d6b440c44ad53bc99412bdca9b3631fc54a0f24ca1bf3c574fa620

Observation 6caeebfe-07b7-4133-9379-0b036559bba2 · outbound

This paper cites Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.967831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.967831Z digest=sha256:140f8bace54a8e7f397eba431701fc3764ca5fb9eed33f1cc5a3acf502f16e71

Observation 0d56858b-ac09-4ac0-ad31-d73f45facc6c · outbound

This paper cites Xls-r: Self-supervised cross-lingual speech representation learning at scale.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Xls-r: Self-supervised cross-lingual speech representation learning at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.105633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.972414Z digest=sha256:33146667c30a6bd78c8cb47a3c2d5a7c674fb9b9bf8f8d96169915f4d38fbabd

Observation 3dbff9a2-7450-40d3-b556-096fbdcda2bd · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.977100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.977100Z digest=sha256:859568669ba4804faf2098c8e354470e7090014fa4decfd015ddf3997ac373b4

Observation 62f50549-e5f2-4082-84c2-aa4a81d73ecb · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.981186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.981186Z digest=sha256:8241dfaa41c3e2c8dd2da6e16a6a14431401325276377574db148121bc4342e6

Observation a4b7e83e-c38c-4e3a-8a4a-ccc9451e40c4 · outbound

This paper cites and Timofte, R.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Timofte, R

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.089504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.985691Z digest=sha256:703af1c9dbbc5338ff08c28c01b220e2fdac8a56549f3e860764149d0756bfb4

Observation 9d21efac-daea-48f0-a0b4-23f51b2e50e2 · outbound

This paper cites Large Language Models are Strong Audio-Visual Speech Recognition Learners.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Large Language Models are Strong Audio-Visual Speech Recognition Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.989857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.989857Z digest=sha256:5b2eb8156e0ac2b7a5ae25a0e7ed80cbe7b816346e2775de0852e049dfcd64d2

Observation cd483a05-4471-44fb-91fb-0d045342d569 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:13.080054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.994730Z digest=sha256:985094ad526437a66814512927586dc2ed93a61dcea590d293c9b2a29015a2ee

Observation 94c28bb5-c138-44f7-9c7c-6d1910553bdd · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.070272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.998955Z digest=sha256:7ebab1987b14f0ada36ab41edcd5e73e041c32848e1efe372d926b3653394372

Observation 0313e3cd-1a74-4706-b120-33ba3729349d · outbound

This paper cites Mixtures of experts for audio-visual learning.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mixtures of experts for audio-visual learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.060255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.003079Z digest=sha256:e8a13093a1712b4fb1cb61278b86e38e04bfe7a13f5fe0d3561435f347caa5aa

Observation d3bee3c5-68f5-4dab-b490-218a46cd56b7 · outbound

This paper cites Self-supervised learning with random-projection quantizer for speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Self-supervised learning with random-projection quantizer for speech recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.050389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.007423Z digest=sha256:74a2f54b0b07b9509f737d95c348b8317cb95854b34ab4d12ef3c5802f23bb34

Observation 8b5cf09a-2a9a-4c10-b50a-b937abbd56ee · outbound

This paper cites J., Kim, M., and Ro, Y.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition J., Kim, M., and Ro, Y

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.040422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.011353Z digest=sha256:92c4001677ecaeed4adc0ac5a823b19a9331011bd8331aa313741b6d26431371

Observation 10622f6a-273b-455e-9a13-874523981cac · outbound

This paper cites S., Nagrani, A., and Zisserman, A.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition S., Nagrani, A., and Zisserman, A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.030425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.015208Z digest=sha256:7343d987b15aa36855a91959bb02d1b0c603272f69fcee7077a6677ddb7154ff

Observation 5e64d104-7cd0-4ae7-8d51-06a432a874d7 · outbound

This paper cites Unified scaling laws for routed language models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unified scaling laws for routed language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.019361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.019021Z digest=sha256:a436e9a9aa5c33be8dbe5464136691489a3309ee2d78f597ee6587fd374c01d6

Observation 5dd41cc2-e5c3-4328-aee0-e1070ad4bbc5 · outbound

This paper cites Stablemoe: Stable routing strategy for mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Stablemoe: Stable routing strategy for mixture of experts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.008786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.022925Z digest=sha256:3cb7066c7acd5167f5e6857deb5f69f2897e8bb127bab97be64ae8792508b1ab

Observation 0405de2f-f8a3-485f-94b7-c584ad2fa1ec · outbound

This paper cites A study of dropout-induced modality bias on robustness to missing video frames for audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A study of dropout-induced modality bias on robustness to missing video frames for audio-visual speech recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.998388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.027308Z digest=sha256:b9854add51f9650d192cbabaa2fcfb7ad7a93a237e186c169f9939a6b67bdb17

Observation 5fa66c23-3834-481f-b695-ae71b99e2503 · outbound

This paper cites and Luettin, J.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Luettin, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.987504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.031328Z digest=sha256:a38e5881fe1c0da867c211f24f058470af53893830cfdf20ecac48918705c05d

Observation 54992c07-7d83-4e18-8e1b-8fd801d9fcd0 · outbound

This paper cites W., and Matt, P.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., and Matt, P

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.977185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.035295Z digest=sha256:55b3f8c3004fc2f8c62b2ae682d8a8159292ae5cb9aef32fecb15aad5d671abe

Observation eb738fe9-acd8-45a5-8078-6693ed7466bd · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.039182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.039182Z digest=sha256:78351220b0c728cf84d79e0fc1093a9147fd7dc6592fbd7cb77cef03835fcaa6

Observation b727f062-e417-4fa9-9016-14e32e64f9ff · outbound

This paper cites Boosting speech recognition robustness to modality-distortion with contrast-augmented prompts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Boosting speech recognition robustness to modality-distortion with contrast-augmented prompts

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.959944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.043299Z digest=sha256:20f9f27a7187b09c793c47a807418551d34a09c06381ae8c115fae4c051cb52d

Observation 96a8fece-b9ff-4f10-ac66-9350ea481c97 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Conformer: Convolution-augmented transformer for speech recognition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.949631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.047254Z digest=sha256:76f60363afcf8e2c000083c2fb46df6bebb54643ea18fd13cbdef30b2ee202da

Observation 0386b6f4-63c0-4444-8e36-a72ec75d3e59 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.051209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.051209Z digest=sha256:b98a1b0bbf346a4d2a4557d07552e121d8a73606746c6bcf9763ed5bde6431bd

Observation 6f54ce98-8b96-4709-a4c5-b81514cd5cea · outbound

This paper cites Jointly learning visual and auditory speech representations from raw data.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Jointly learning visual and auditory speech representations from raw data

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.938761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.055256Z digest=sha256:c7b8bc7ee16b2bcf47cc60877a66af038e4d7ea5f5fc420c4e11e28bd5c1bc98

Observation 55a73f4d-4585-42da-9190-fea994ad50c4 · outbound

This paper cites Braven: Improving self-supervised pre-training for visual and auditory speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Braven: Improving self-supervised pre-training for visual and auditory speech recognition

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.927405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.059701Z digest=sha256:11cfd9a1661396d2159b89c7144e2601c931d2782ffe4764fa486d321cf825c6

Observation d65721a7-5e49-4b3c-9e82-243a295f819e · outbound

This paper cites XLAVS - R : Cross-lingual audio-visual speech representation learning for noise-robust speech perception.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition XLAVS - R : Cross-lingual audio-visual speech representation learning for noise-robust speech perception

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.917268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.063242Z digest=sha256:566aa301040596f6fe5f1565133d4d433c3d787b597e36dcd0714687c715b7bb

Observation 9c4f2e9f-e739-4545-a79d-fa40926bb734 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.906730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.067010Z digest=sha256:e58c0a0bfce090a1875f8de04e8355cb18c8588668f408e63112fdd90103b78a

Observation 1dd6c8d1-c13f-468b-987a-f6e4568ca8da · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.896616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.071089Z digest=sha256:54707c9dfd56f5e1971d48620d0bf065a30e4551f8500bdc03b2406c85aae532

Observation 3875f7e9-9ea6-48d8-bd22-2b184af8a91f · outbound

This paper cites and Shi, B.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Shi, B

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.886581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.074951Z digest=sha256:983f269622bd42bfbe12df6c686a80b00113bc9521b69ab6d8a02802b3d9304e

Observation ac8f7845-fe02-42d2-83cf-0fa56b0f684d · outbound

This paper cites H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.078580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.078580Z digest=sha256:ebc6f8948f99062e72f5ea9cb5602eda2c4573f8f88f353600d07461f3df5b42

Observation 6f6de3fe-b46c-4d9f-9d0c-4bcc01d00591 · outbound

This paper cites N., Zhang, Y., and Beaufays, F.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition N., Zhang, Y., and Beaufays, F

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.869924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.082306Z digest=sha256:87b18fd1a98a58fa2cd7133391954fea980ea28a119669d85f3b33ca7c46577e

Observation 976eb8dc-4005-4619-8d18-2970d5df3951 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.859640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.085973Z digest=sha256:6c302c2fa3616482da17837e3c72a7634a9e71ac6ca5961d40b0820c51e6626c

Observation 8bddd93f-c46e-47c0-93bc-95c37467a604 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.849510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.089684Z digest=sha256:a282856fdf35fc8b913ba05fdcdf445fa987ad6ad4f073630393bd55ae68461c

Observation 3cc7802d-eaff-464b-b122-a172b1f036fb · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.839686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.093433Z digest=sha256:1c9366d43d377e6adad5364efa06eb40cbd3299ad040edae8954d6e6cb8e7331

Observation 83791601-00d5-48e6-b597-166a6fcf1351 · outbound

This paper cites A., Jordan, M.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A., Jordan, M

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.097176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.097176Z digest=sha256:eda806c87fc1d38c052fc0ad337e299e34b722ed874401b60c510124a2195261

Observation 9c769899-7176-4dfd-a205-9ec91ae2553b · outbound

This paper cites Mixtral of Experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mixtral of Experts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.100811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.100811Z digest=sha256:67dfdb46a6dd116da7fd252f23331bbb4b23ff8a71df7cd0a7499063b9c58b6a

Observation 1b155f64-0f2f-48f0-9979-835e89202665 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.824050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.104768Z digest=sha256:ebf8f5b35514baec1d6adb10da57c1d21ce8a56bbb2cd508cfa47a794629da34

Observation 3c730175-c11c-4cf6-8f99-b4b2550e4145 · outbound

This paper cites Scaling Laws for Neural Language Models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Scaling Laws for Neural Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.108410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.108410Z digest=sha256:b27658f63aa0e2f0f576bfb1e3a1f9db403424c3cd2665521e2e8d01bb640b0b

Observation f9ba568c-e35b-46a9-a66d-5c63c2c32c62 · outbound

This paper cites Learning video temporal dynamics with cross-modal attention for robust audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning video temporal dynamics with cross-modal attention for robust audio-visual speech recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.814744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.112122Z digest=sha256:659f6056b820b36e888f51b663919c24d1db87756a6f96ceaa0c9fc0f323ddab

Observation eb92f129-9c51-4e93-9443-b911ae612b24 · outbound

This paper cites Multi-task corrupted prediction for learning robust audio-visual speech representation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Multi-task corrupted prediction for learning robust audio-visual speech representation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.804393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.115892Z digest=sha256:aa3d5cf427d0809fc99f33622b26c184e616fde2f3f45ccefad63b46546443b1

Observation e50f65bf-d38b-4a7a-95f0-414b568881b1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Adam: A Method for Stochastic Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.119735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.119735Z digest=sha256:91fc95fc1278725c21a524211902f88c7c371c3c4734e5c8e9e77cb08d6e97e7

Observation 9d557243-3fe3-4ea7-a398-006108a0781c · outbound

This paper cites Moai: Mixture of all intelligence for large language and vision models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Moai: Mixture of all intelligence for large language and vision models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.794162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.123375Z digest=sha256:50cc1bff2450c82a6ee26e371289c0a4a92635216230a62391dbb046146a10e9

Observation ba99a6f1-d267-4102-9903-c77f037e8c06 · outbound

This paper cites \ GS \ hard: Scaling giant models with conditional computation and automatic sharding.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition \ GS \ hard: Scaling giant models with conditional computation and automatic sharding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.783255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.127007Z digest=sha256:5128317a1f3a3aad0d64bb5f040478c79befc2840488e34ebc4670ccaa498d81

Observation 3ad298ad-4c4a-4eb7-b2e3-256862aa31fd · outbound

This paper cites Unified cross-modal attention: Robust audio-visual speech recognition and beyond.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unified cross-modal attention: Robust audio-visual speech recognition and beyond

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.772919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.130698Z digest=sha256:3ad0926dcf2a6007f385924347cafa43540ef6e2cccb05aea062d7d329d82503

Observation 982ffbce-c590-48c5-981d-0ec2a7725ecf · outbound

This paper cites Pace: Unified multi-modal dialogue pre-training with progressive and compositional experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Pace: Unified multi-modal dialogue pre-training with progressive and compositional experts

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.762175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.134232Z digest=sha256:c78604fa03cc7b6c74e69ad1b0737a33440183e6f4f17154ec99fca63dc9fd63

Observation 7ee4fd34-f493-45cf-93ec-be559e2e05e9 · outbound

This paper cites Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.137924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.137924Z digest=sha256:9714e1fd01e32e1fbd87ac27a34ce0a52e74ae204b2b794ca2751535f7a8650e

Observation 2e1e6b9f-0e75-4527-98d2-32b203226ecd · outbound

This paper cites Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.750980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.141733Z digest=sha256:6bdcc779e8c6fdd7a16fdba62dc059df069c0cc14f9dd26a63317a5dc8dfda12

Observation 4d5394ea-e862-476a-9d68-7db199733b31 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.145504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.145504Z digest=sha256:3c1d065db4cbf21074b85a24cc8acf19ac6de47104f83f25b6d9a06b65ab846a

Observation 76554e89-2cc7-415b-9f11-c51a54917429 · outbound

This paper cites W., and Pantic, M.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., and Pantic, M

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.740438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.149611Z digest=sha256:cff76c3d316f34eb91773689e0053f88736e33f1e49581665ba5c6ff2ed273d3

Observation decb22d6-5180-42cb-b35a-4c68ea23fd0b · outbound

This paper cites End-to-end audio-visual speech recognition with conformers.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition End-to-end audio-visual speech recognition with conformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.730261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.153350Z digest=sha256:f2999b86dcf60951b7d971cde939dd882d51c7dcd08418c4b03a8825ed17ea29

Observation b6e30b98-4524-45fd-ba64-55921ab1de61 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.719069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.156880Z digest=sha256:2e0c90f387075f702fe37c860852aee8b7be09710a902d339a66be6064cb3620

Observation 7a4dd371-89e0-49d0-8b75-82ee9dfb957c · outbound

This paper cites Recurrent neural network transducer for audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Recurrent neural network transducer for audio-visual speech recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.707947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.160528Z digest=sha256:ee7bb745ca56f818bdff5cc770b0f9bad5693759932f7f1d69e15aa23ee756b4

Observation 40f518dc-ba49-4251-969c-5dbce41495e0 · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mm1: methods, analysis and insights from multimodal llm pre-training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.696792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.164169Z digest=sha256:d6c8ef1b88f96657c13401b66b0f034a0a529c2880a0bb674f696d34fe877b6a

Observation 736a4236-703c-4be2-8c3b-53dfb9d93e9c · outbound

This paper cites Multimodal contrastive learning with limoe: the language-image mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Multimodal contrastive learning with limoe: the language-image mixture of experts

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.686160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.167622Z digest=sha256:8eeec2b05a4226f8a00113710ba805b51ad09a0f67e47a03698f06e14f6262f3

Observation 262ab514-50ef-47a2-abb3-201a5b8f83c8 · outbound

This paper cites G., and Ogata, T.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition G., and Ogata, T

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.675558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.171272Z digest=sha256:f7ac0e6f1d3dfae11f77e582fc11e554e5b584ee22dc2d7a77a8f57a66618217

Observation a554eb4e-ae2a-4261-95fa-43ec6a0329ce · outbound

This paper cites Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.665127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.174871Z digest=sha256:d4b8cd0708209877a49de4202775f3289727b957358874791565aa758778faf8

Observation 3a623aaa-0e13-431f-8564-fb3387ff6617 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Bleu: a method for automatic evaluation of machine translation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.179726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.179726Z digest=sha256:adb719e046388763504ea478e533ad3bbb921c3e53cc499793de942fc9c279f9

Observation 7c91ae05-3531-44c0-99b7-4a1e6ee84503 · outbound

This paper cites A call for clarity in reporting bleu scores.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A call for clarity in reporting bleu scores

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.648936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.183249Z digest=sha256:148c2c0243bf8824d86a285763743724d339af59adada1ba685562119896fbd9

Observation faf4f2ff-e926-49b8-9ebe-e85bfd014e25 · outbound

This paper cites Lipsound2: Self-supervised pre-training for lip-to-speech reconstruction and lip reading.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Lipsound2: Self-supervised pre-training for lip-to-speech reconstruction and lip reading

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.638468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.186969Z digest=sha256:dd9b9da4b783caa6525a77ba4e201e974a211453cfec691dea459918121fcd03

Observation 1b11b7af-4db4-4e2b-92e6-e9953ef6fad7 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.190741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.190741Z digest=sha256:9a75b676aba08bbb73edf96c9836d0c03543594694a4bb4963656ec8c8b366ac

Observation 9e558eeb-03d6-4dde-9b12-1e0885aa2506 · outbound

This paper cites Learning from the master: Distilling cross-modal advanced knowledge for lip reading.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning from the master: Distilling cross-modal advanced knowledge for lip reading

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.622031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.194385Z digest=sha256:c5c09f58bf111732beee7167ba6a3626553e63992a17dc15d78c004f4fce96dd

Observation e5f9b8f4-3539-49df-ba88-92dbcfee486b · outbound

This paper cites wav2vec: Unsupervised pre-training for speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition wav2vec: Unsupervised pre-training for speech recognition

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.611683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.197809Z digest=sha256:aa5df5df90e09954db71e4ac597c61486cbb66a4809ca6a9674fe94162dcf091

Observation e6fe91de-707b-404d-b25d-68d2f3a1f931 · outbound

This paper cites H., Nagrani, A., and Schmid, C.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition H., Nagrani, A., and Schmid, C

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.600769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.201418Z digest=sha256:21118627388c21a133709cfd18aeef3fd8050c3e989b10c8bdc59edc789a3f3a

Observation 93c99fbf-3002-4c1e-bd85-c855be4c7cd0 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.589682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.204937Z digest=sha256:2c63b514b38f16c69938bd8560cd5d25b09b1766c192035a8b52dd8b1fbd9409

Observation 698d2898-de93-452b-a09a-31664d927a91 · outbound

This paper cites Scaling vision-language models with sparse mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Scaling vision-language models with sparse mixture of experts

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.579332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.208750Z digest=sha256:fd85fb2fb66cccbf7135a78e8f810a8d31b64133db3a3ab5f976ebd252abd71a

Observation 761244b3-694b-40ea-b43e-498b303e25b9 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning audio-visual speech representation by masked multimodal cluster prediction

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.567388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.212277Z digest=sha256:f9d9f3453a45299904bfa1e838020570de32503d112cdbbf5b7d32edeee67c3c

Observation 6455672e-f8dc-4faa-beec-aee8ca7f209d · outbound

This paper cites Robust self-supervised audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Robust self-supervised audio-visual speech recognition

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.556400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.215914Z digest=sha256:fbfe20ad783ce4b02d65e5d2a01871e8b93e03d3a8b0fb6dafba596662103ff2

Observation e469726c-a996-4160-b4cf-4052afa9dfe9 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition MUSAN: A Music, Speech, and Noise Corpus

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.219928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.219928Z digest=sha256:64cdb40a9e27d2543225cd7f00763d79127d09d0479cbcd7b5f1edb72fb67259

Observation ea26ca4a-a117-4a82-87d1-9e3eb8a279eb · outbound

This paper cites The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.223984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.223984Z digest=sha256:eb6b9fdf1765f8e5c8b356ad49dfe3f27c749227f85de304a4c93c1b4af61024

Observation 70e6dd0a-b44f-498d-b5c3-5266d235f667 · outbound

This paper cites N., Kaiser, ., and Polosukhin, I.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition N., Kaiser, ., and Polosukhin, I

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.227915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.227915Z digest=sha256:b8b25d7772e28cdb41f6cbdd6999626980d49fac8a28991bb0daea1cbd912e86

Observation 683256ca-aa77-47d4-9a5f-4f8c754714e0 · outbound

This paper cites T., and Li, H.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition T., and Li, H

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.534091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.231696Z digest=sha256:c80a9f4774baeba1027a3d895442ffb554399b83c89adf61fb6ae3a3a18e284f

Observation 1e95febc-e8b0-4589-b53d-916f1e2398be · outbound

This paper cites Language-routing mixture of experts for multilingual and code-switching speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Language-routing mixture of experts for multilingual and code-switching speech recognition

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.524068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.235459Z digest=sha256:4632c065d8b8f75a663ace67d58e39bb0d1f860f8a84008c1daaad7955d68f1d

Observation 555f1a1d-5da3-4299-9a20-b307e07dc4fe · outbound

This paper cites R., and Hayashi, T.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition R., and Hayashi, T

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.513677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.239453Z digest=sha256:821fb889864d57e44d525e4d3c691fad05ca8ff0f856dc52fd78a59eb3fec3e6

Observation 962a535d-9a14-443b-9c67-34f771193ca1 · outbound

This paper cites Robust audiovisual speech recognition models with mixture-of-experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Robust audiovisual speech recognition models with mixture-of-experts

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.503834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.243307Z digest=sha256:69e92e82888a798cae9efac65b322672d3046f93222c5f47709df45634f46357

Observation 3584682c-30ff-4622-9c96-728984916f11 · outbound

This paper cites Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.492396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.246899Z digest=sha256:1076395cfb6ab19f2f30cba57e525d67219aa56328db9596c16b1648fb75ba93

Observation d8f1c0d4-9dd2-4fcb-81c8-501a463e55d6 · outbound

This paper cites Speechmoe2: Mixture-of-experts model with improved routing.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Speechmoe2: Mixture-of-experts model with improved routing

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.477707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.250437Z digest=sha256:7fcd35414708317dfed80a436f4b312f572f14446b78858d9937f59ce1f6b8f2

Observation 08b33f54-e9a5-4b79-b0ed-562a83643f94 · outbound

This paper cites Visual hallucination elevates speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Visual hallucination elevates speech recognition

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.464533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.253960Z digest=sha256:fb6c30eb36d30e0b3806acd2d819da2cd1af7c9e1c0a79e78ccd0a5cb22f7e77

Observation 5b13026c-fa99-4c0d-970f-b9f68f4aff78 · outbound

This paper cites Self-supervised audio-visual speech representations learning by multimodal self-distillation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Self-supervised audio-visual speech representations learning by multimodal self-distillation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.453215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.257656Z digest=sha256:9d4cd921ec733980cec3fd911d3d8df06b125293bbf2ffba54a7aa43912e292a

Observation 3515f5d7-0b53-4273-bafd-96cd12610c8b · outbound

This paper cites Uni-perceiver-moe: Learning sparse generalist models with conditional moes.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Uni-perceiver-moe: Learning sparse generalist models with conditional moes

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.440286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.261199Z digest=sha256:140c6f41104a4b42c2eeb09c4f2e39de0eb2bb8f4b9955deabe4ae2f96c25413

Observation 0a2625e8-12f8-4dad-8059-86cde6accefa · outbound

This paper cites Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.425112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.264765Z digest=sha256:8ff9280cfd6a0e5367e8de77cb33e88d6d84e540223d283d1b66a01eae592094

Observation b3cea8a1-ca54-4a16-b6fc-d1420e753587 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.269019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.269019Z digest=sha256:702bdbc20bd1bd1a3ab921dd458f4d42994ec588e1d9f226e6f45fb2b68711f4

Pith citing papers

Observation 10189419-0fe1-4ad1-9f36-c209cea88af3 · inbound

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach cites this paper.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:39:34.345367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:39:34.111978Z digest=sha256:0253b5a2694ac6a5590b611bbcc8c10ae7fc7fd7eba3293a70f51c043bd1ec98