Pith. sign in

Paper Citation Record · LEDGER

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

As of 19 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 1 inbound Pith citation observation for arXiv:2502.10447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10447 v2

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:46:12.269019Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:34.111978Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:39:34.251185Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cb00bbcc-96c2-42e8-b779-f4adb89613c5 · outbound

This paper cites write newline.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.941722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.941722Z digest=sha256:d4f2f82936509a78e18d2116dcdb43fde7e3d9eac8ee5f4e31656c006993cd26

Observation f8a90e81-c9e9-459c-87f8-006a926ba1dc · outbound

This paper cites GPT-4 Technical Report.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.947979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.947979Z digest=sha256:08365c0853d17186f771caaa0f300c179f6c7db719ac926c57962ff864072130

Observation 78d4d530-6232-44cd-a438-e5cf4c89d29e · outbound

This paper cites S., Senior, A., Vinyals, O., and Zisserman, A.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition S., Senior, A., Vinyals, O., and Zisserman, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.952950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.952950Z digest=sha256:1fd10a1ea9a17504ceea053342c7e08f1d470ea8dc2906bbbe0b75b94b1bb680

Observation 09235db7-3652-4033-92ad-731095d6de81 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.957891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.957891Z digest=sha256:52a7029bfa144360c6f00595697a7af51eb3a29d6e599c313e80196378b1bd6f

Observation 2e695b4c-6200-41a6-b137-6a12fb193dd7 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.963258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.963258Z digest=sha256:255e4c32c88378da6959ca88f85d1123686c6be00dd5de358b41c68b5cd2e2cc

Observation 6caeebfe-07b7-4133-9379-0b036559bba2 · outbound

This paper cites Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.967831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.967831Z digest=sha256:43b98f280ffc4b2edb6f407f2012db7c1b856d0339fd3ebb583e2a12f0dbe2cc

Observation 0d56858b-ac09-4ac0-ad31-d73f45facc6c · outbound

This paper cites Xls-r: Self-supervised cross-lingual speech representation learning at scale.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Xls-r: Self-supervised cross-lingual speech representation learning at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.105633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.972414Z digest=sha256:e6836c1ba77fd27857030fec65938864bbd2dd568b8ac4062adb2ad44de27a5b

Observation 3dbff9a2-7450-40d3-b556-096fbdcda2bd · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.977100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.977100Z digest=sha256:581f0f1ab4060c40a6604d8544ecc1c7c3e29986c51f831a5019cc4225b2b85e

Observation 62f50549-e5f2-4082-84c2-aa4a81d73ecb · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.981186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.981186Z digest=sha256:a6a994912017fc9da9cc3062815504ceccac821eec33a6bd7d45b2933ffa3ec7

Observation a4b7e83e-c38c-4e3a-8a4a-ccc9451e40c4 · outbound

This paper cites and Timofte, R.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Timofte, R

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.089504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.985691Z digest=sha256:81d7f7db664027f3b4b0ace118db115f29a462593bed4d27025748af9f8f5b7b

Observation 9d21efac-daea-48f0-a0b4-23f51b2e50e2 · outbound

This paper cites Large Language Models are Strong Audio-Visual Speech Recognition Learners.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Large Language Models are Strong Audio-Visual Speech Recognition Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.989857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.989857Z digest=sha256:baa13d065e4f95bce5cbd20ec759f7ca51f420a22bd90a603b49b7bb3edb0959

Observation cd483a05-4471-44fb-91fb-0d045342d569 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:13.080054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.994730Z digest=sha256:10cba05028b03b7ccebae3a321e882538915866bfc7d702aed3c8ebc15e8c24f

Observation 94c28bb5-c138-44f7-9c7c-6d1910553bdd · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.070272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.998955Z digest=sha256:461ef736c4ce4190b0b6ce8a7953eaa59320b6eeed8bec2fb8d833b1baaf5939

Observation 0313e3cd-1a74-4706-b120-33ba3729349d · outbound

This paper cites Mixtures of experts for audio-visual learning.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mixtures of experts for audio-visual learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.060255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.003079Z digest=sha256:d980f8c7bf7e47cb719c44a63c52edf271d632b2451411c0a3c8b93063637f29

Observation d3bee3c5-68f5-4dab-b490-218a46cd56b7 · outbound

This paper cites Self-supervised learning with random-projection quantizer for speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Self-supervised learning with random-projection quantizer for speech recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.050389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.007423Z digest=sha256:583e96e43e993c7e12ac98004af2350c88329ed85f13dc32edf9f942ad2e54d7

Observation 8b5cf09a-2a9a-4c10-b50a-b937abbd56ee · outbound

This paper cites J., Kim, M., and Ro, Y.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition J., Kim, M., and Ro, Y

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.040422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.011353Z digest=sha256:5224080b2085d59be9a21b3ece417d887550ea44b656e35cd82418b0d400c070

Observation 10622f6a-273b-455e-9a13-874523981cac · outbound

This paper cites S., Nagrani, A., and Zisserman, A.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition S., Nagrani, A., and Zisserman, A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.030425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.015208Z digest=sha256:0cfbc733455b3886b7a657ab393364ac6afca9c703baa5674dac2fdeb817a990

Observation 5e64d104-7cd0-4ae7-8d51-06a432a874d7 · outbound

This paper cites Unified scaling laws for routed language models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unified scaling laws for routed language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.019361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.019021Z digest=sha256:8b580e424c71f4b27072457c0b147a9972167701e3dc2e4bb2c1223555fb8820

Observation 5dd41cc2-e5c3-4328-aee0-e1070ad4bbc5 · outbound

This paper cites Stablemoe: Stable routing strategy for mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Stablemoe: Stable routing strategy for mixture of experts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.008786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.022925Z digest=sha256:648778febda0c95db32be93b59750af2bcf7b166c0351ad8e7d0b1f5900d4964

Observation 0405de2f-f8a3-485f-94b7-c584ad2fa1ec · outbound

This paper cites A study of dropout-induced modality bias on robustness to missing video frames for audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A study of dropout-induced modality bias on robustness to missing video frames for audio-visual speech recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.998388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.027308Z digest=sha256:db7c9d3e64ef1d1e9ede28907d5b45ce4f8cb0d2a1f184fe23e1db2f431fa600

Observation 5fa66c23-3834-481f-b695-ae71b99e2503 · outbound

This paper cites and Luettin, J.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Luettin, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.987504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.031328Z digest=sha256:fd5be4410d9532966f10ef7c07c6026f0d4167c715d03dd8c196c7ed3f7d50e8

Observation 54992c07-7d83-4e18-8e1b-8fd801d9fcd0 · outbound

This paper cites W., and Matt, P.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., and Matt, P

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.977185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.035295Z digest=sha256:5477e07ee66b67a5ce0bcfae19ba0d0753b6782a832797de1fae7bd9204322fb

Observation eb738fe9-acd8-45a5-8078-6693ed7466bd · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.039182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.039182Z digest=sha256:6d2c9480746d6957152bc7f1b796ce518ff589d1eb196d499822175c5bc05e03

Observation b727f062-e417-4fa9-9016-14e32e64f9ff · outbound

This paper cites Boosting speech recognition robustness to modality-distortion with contrast-augmented prompts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Boosting speech recognition robustness to modality-distortion with contrast-augmented prompts

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.959944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.043299Z digest=sha256:406a8fd067a6e6ef410966d941d2a24b263d16f53748b11dc33ee738e662fcdf

Observation 96a8fece-b9ff-4f10-ac66-9350ea481c97 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Conformer: Convolution-augmented transformer for speech recognition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.949631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.047254Z digest=sha256:c5b8fcb82605d1b29867c2db771e115de70461c2528ced1d04d3d4ac2cb8c999

Observation 0386b6f4-63c0-4444-8e36-a72ec75d3e59 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.051209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.051209Z digest=sha256:a674581a34775d609fbe23879275f00fb3f49c67c06bfae9ed4c2d8e0d96bdff

Observation 6f54ce98-8b96-4709-a4c5-b81514cd5cea · outbound

This paper cites Jointly learning visual and auditory speech representations from raw data.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Jointly learning visual and auditory speech representations from raw data

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.938761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.055256Z digest=sha256:a2d7131b8d6e9c487a98bfd9c96f4ef48182f678f3d05105ab82b6f636de282f

Observation 55a73f4d-4585-42da-9190-fea994ad50c4 · outbound

This paper cites Braven: Improving self-supervised pre-training for visual and auditory speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Braven: Improving self-supervised pre-training for visual and auditory speech recognition

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.927405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.059701Z digest=sha256:25519baf758040c33d5bca1102066a297f6aa6d767b35f0303ffc7aa83df740b

Observation d65721a7-5e49-4b3c-9e82-243a295f819e · outbound

This paper cites XLAVS - R : Cross-lingual audio-visual speech representation learning for noise-robust speech perception.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition XLAVS - R : Cross-lingual audio-visual speech representation learning for noise-robust speech perception

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.917268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.063242Z digest=sha256:1c788916254ee12fc38d084760dedb42205f8827a626d65517d4980160a082ee

Observation 9c4f2e9f-e739-4545-a79d-fa40926bb734 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.906730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.067010Z digest=sha256:464f9c35a5c70497959cbb4f854bca68eafa2f098c2a809d47579f14b1f070eb

Observation 1dd6c8d1-c13f-468b-987a-f6e4568ca8da · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.896616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.071089Z digest=sha256:07487349f72356268fd7896bf4a3cc18a921979dd9b39cc6749c870ac678eb21

Observation 3875f7e9-9ea6-48d8-bd22-2b184af8a91f · outbound

This paper cites and Shi, B.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Shi, B

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.886581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.074951Z digest=sha256:a33e544144ca8beee67df9440a6c2105befa83516a5c7667e20a59e6e7afa8fa

Observation ac8f7845-fe02-42d2-83cf-0fa56b0f684d · outbound

This paper cites H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.078580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.078580Z digest=sha256:9d5eb4b6b966049c4cae6ac5c50c39780638ca2643f18c1ef46a85ce9164c2da

Observation 6f6de3fe-b46c-4d9f-9d0c-4bcc01d00591 · outbound

This paper cites N., Zhang, Y., and Beaufays, F.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition N., Zhang, Y., and Beaufays, F

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.869924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.082306Z digest=sha256:4cef97e29890b95b715da2bc96b3bdd5527de3df5a1265596da243fcd62624bc

Observation 976eb8dc-4005-4619-8d18-2970d5df3951 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.859640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.085973Z digest=sha256:459a20b980dfaa33ec64c45241296290f9c72e75327b482bdf25c29c9ac51206

Observation 8bddd93f-c46e-47c0-93bc-95c37467a604 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.849510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.089684Z digest=sha256:7946cb053f9d834b33e8194091437036d0eee61f25af948bbea60a56d740a113

Observation 3cc7802d-eaff-464b-b122-a172b1f036fb · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.839686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.093433Z digest=sha256:23c1fe69e9f15178902266e674a0a6b285244a4bff56b3a0c97bce0e7c596923

Observation 83791601-00d5-48e6-b597-166a6fcf1351 · outbound

This paper cites A., Jordan, M.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A., Jordan, M

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.097176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.097176Z digest=sha256:9cce60db49764841c8ffcae99b3697017ebaff2b4d5565170a0187df8997eefd

Observation 9c769899-7176-4dfd-a205-9ec91ae2553b · outbound

This paper cites Mixtral of Experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mixtral of Experts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.100811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.100811Z digest=sha256:b631aa2fae1155380775ed3fcd91cf3e328027956ce97051b1c671c1af330396

Observation 1b155f64-0f2f-48f0-9979-835e89202665 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.824050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.104768Z digest=sha256:2a1156fac8320b8d79a58ade2aea37cfdeb1a34b8de6b55c3b69d00dbaa20fd4

Observation 3c730175-c11c-4cf6-8f99-b4b2550e4145 · outbound

This paper cites Scaling Laws for Neural Language Models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Scaling Laws for Neural Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.108410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.108410Z digest=sha256:dc2d8912a684a379503f0befa122a74cebf45222dee477ba81292f3fbc5f82ae

Observation f9ba568c-e35b-46a9-a66d-5c63c2c32c62 · outbound

This paper cites Learning video temporal dynamics with cross-modal attention for robust audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning video temporal dynamics with cross-modal attention for robust audio-visual speech recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.814744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.112122Z digest=sha256:b45cb55c6e831e7895d137bf8a540b5761b7d0abb9f32d4960b8b2bab2c519d7

Observation eb92f129-9c51-4e93-9443-b911ae612b24 · outbound

This paper cites Multi-task corrupted prediction for learning robust audio-visual speech representation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Multi-task corrupted prediction for learning robust audio-visual speech representation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.804393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.115892Z digest=sha256:2f503a825f2fcd869927a56a87ccc25bb77d922deed449513a4907b175fdafb3

Observation e50f65bf-d38b-4a7a-95f0-414b568881b1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Adam: A Method for Stochastic Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.119735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.119735Z digest=sha256:05052eced5c0e58c687e8bf5487a626ab82c827c8685f7e1b54cd4cbe18451d6

Observation 9d557243-3fe3-4ea7-a398-006108a0781c · outbound

This paper cites Moai: Mixture of all intelligence for large language and vision models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Moai: Mixture of all intelligence for large language and vision models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.794162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.123375Z digest=sha256:c80e70a31812b12c103570675e484b48d992279bc6db196615bd31345de84c09

Observation ba99a6f1-d267-4102-9903-c77f037e8c06 · outbound

This paper cites \ GS \ hard: Scaling giant models with conditional computation and automatic sharding.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition \ GS \ hard: Scaling giant models with conditional computation and automatic sharding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.783255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.127007Z digest=sha256:e312ae2891336f97e722ed05862a983320e04ed9ee7c6638070e056ddb16ed97

Observation 3ad298ad-4c4a-4eb7-b2e3-256862aa31fd · outbound

This paper cites Unified cross-modal attention: Robust audio-visual speech recognition and beyond.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unified cross-modal attention: Robust audio-visual speech recognition and beyond

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.772919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.130698Z digest=sha256:1c7a99ab886a5bf0fdce1bbcdc097f587a9725294769ce3242413354922b71ea

Observation 982ffbce-c590-48c5-981d-0ec2a7725ecf · outbound

This paper cites Pace: Unified multi-modal dialogue pre-training with progressive and compositional experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Pace: Unified multi-modal dialogue pre-training with progressive and compositional experts

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.762175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.134232Z digest=sha256:2007d8fd4b9c2ed90112efcd63db0ac9746929f885332ce0cadeba27a75d393c

Observation 7ee4fd34-f493-45cf-93ec-be559e2e05e9 · outbound

This paper cites Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.137924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.137924Z digest=sha256:01f5a282acd9fc7adf99194df7c0514028a4ff4ee471e569c84578b71c8e63eb

Observation 2e1e6b9f-0e75-4527-98d2-32b203226ecd · outbound

This paper cites Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.750980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.141733Z digest=sha256:652b5fd6960ed3d2791feb62f7ab22064171c5007bf9bb2a882edf31f3c8519b

Observation 4d5394ea-e862-476a-9d68-7db199733b31 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.145504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.145504Z digest=sha256:6be9105eafdf9722f8b1ed63de479955fb8d7e94bf69c67285ef66dc41212265

Observation 76554e89-2cc7-415b-9f11-c51a54917429 · outbound

This paper cites W., and Pantic, M.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., and Pantic, M

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.740438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.149611Z digest=sha256:ab59c7571202d74e806ce6bd8dcaa65d323fecc4992aedf91f5ed4b28b366fd6

Observation decb22d6-5180-42cb-b35a-4c68ea23fd0b · outbound

This paper cites End-to-end audio-visual speech recognition with conformers.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition End-to-end audio-visual speech recognition with conformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.730261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.153350Z digest=sha256:bea38a46000358c589695cf7495c0821b54dc7e953d799f3d42318b7bed10958

Observation b6e30b98-4524-45fd-ba64-55921ab1de61 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.719069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.156880Z digest=sha256:46c3b53c19e59c27e1e6a977e74734875a1bc225570524134ceeca569c998319

Observation 7a4dd371-89e0-49d0-8b75-82ee9dfb957c · outbound

This paper cites Recurrent neural network transducer for audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Recurrent neural network transducer for audio-visual speech recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.707947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.160528Z digest=sha256:f35784caf9c1d41dde5c941b741f46366d5c85d856ab280c334483f88247ebad

Observation 40f518dc-ba49-4251-969c-5dbce41495e0 · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mm1: methods, analysis and insights from multimodal llm pre-training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.696792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.164169Z digest=sha256:211e7f758d7061657d485d0992896a4a8abdac384efe15954601106bc580d00f

Observation 736a4236-703c-4be2-8c3b-53dfb9d93e9c · outbound

This paper cites Multimodal contrastive learning with limoe: the language-image mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Multimodal contrastive learning with limoe: the language-image mixture of experts

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.686160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.167622Z digest=sha256:9b499be01b3a5ad492b5643b375907afde38393e8a7b7af077e90c81018c4ad0

Observation 262ab514-50ef-47a2-abb3-201a5b8f83c8 · outbound

This paper cites G., and Ogata, T.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition G., and Ogata, T

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.675558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.171272Z digest=sha256:15ebf2707621459f734b811a1a948a305cc75f7340fdd9cfeb954de2776aad60

Observation a554eb4e-ae2a-4261-95fa-43ec6a0329ce · outbound

This paper cites Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.665127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.174871Z digest=sha256:e9ed168f812ed543fe0bf465cde55ed5b5306df20b05fb11dd2a12e6907b1356

Observation 3a623aaa-0e13-431f-8564-fb3387ff6617 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Bleu: a method for automatic evaluation of machine translation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.179726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.179726Z digest=sha256:06b85af75afbfac9b995fab1fd4a23890bd7b18003325b9645db33de699fd69b

Observation 7c91ae05-3531-44c0-99b7-4a1e6ee84503 · outbound

This paper cites A call for clarity in reporting bleu scores.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A call for clarity in reporting bleu scores

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.648936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.183249Z digest=sha256:e9a6a1023325d3f0b40d65c48a2f8be147198cf7a6163322205c972056510447

Observation faf4f2ff-e926-49b8-9ebe-e85bfd014e25 · outbound

This paper cites Lipsound2: Self-supervised pre-training for lip-to-speech reconstruction and lip reading.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Lipsound2: Self-supervised pre-training for lip-to-speech reconstruction and lip reading

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.638468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.186969Z digest=sha256:ab4be49e07706bbe76d547b8ff2c652e61f907a309ad7778017a42a78231e8c7

Observation 1b11b7af-4db4-4e2b-92e6-e9953ef6fad7 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.190741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.190741Z digest=sha256:28407f86f9c94870888c9f14278ca0d9880aee06cc379a71b316260ccc467f9d

Observation 9e558eeb-03d6-4dde-9b12-1e0885aa2506 · outbound

This paper cites Learning from the master: Distilling cross-modal advanced knowledge for lip reading.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning from the master: Distilling cross-modal advanced knowledge for lip reading

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.622031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.194385Z digest=sha256:08fed245cc7c83e46df7fccc5cbf9de51f54fce80512ae0ad9d361facf84f2f8

Observation e5f9b8f4-3539-49df-ba88-92dbcfee486b · outbound

This paper cites wav2vec: Unsupervised pre-training for speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition wav2vec: Unsupervised pre-training for speech recognition

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.611683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.197809Z digest=sha256:05c49900f029a53c6c8a4bb59be2c358874dd50658f08c0bdcc9f3a4aef63af6

Observation e6fe91de-707b-404d-b25d-68d2f3a1f931 · outbound

This paper cites H., Nagrani, A., and Schmid, C.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition H., Nagrani, A., and Schmid, C

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.600769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.201418Z digest=sha256:99b2e69f9db225e319447cccd03538a493fdb2fe99fc4d5c1edb9ee94469424f

Observation 93c99fbf-3002-4c1e-bd85-c855be4c7cd0 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.589682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.204937Z digest=sha256:24c203c67cd37e31e17c53255c6b2af07753ddfcca12ead10e108d86701c2dbb

Observation 698d2898-de93-452b-a09a-31664d927a91 · outbound

This paper cites Scaling vision-language models with sparse mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Scaling vision-language models with sparse mixture of experts

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.579332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.208750Z digest=sha256:6c0da1c6a2420b712446efd2294ed5dc474cd801b3fd84b6a6e0cd4064a17443

Observation 761244b3-694b-40ea-b43e-498b303e25b9 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning audio-visual speech representation by masked multimodal cluster prediction

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.567388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.212277Z digest=sha256:ae7368c524f9308501c7ccf85dfb59bb0ce2d7568fbd88d725a23986e8c2f33f

Observation 6455672e-f8dc-4faa-beec-aee8ca7f209d · outbound

This paper cites Robust self-supervised audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Robust self-supervised audio-visual speech recognition

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.556400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.215914Z digest=sha256:2ddd7d88e22e836adc49caff81c7e0495d5d9949a1a9a4e24829a801e9bce68c

Observation e469726c-a996-4160-b4cf-4052afa9dfe9 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition MUSAN: A Music, Speech, and Noise Corpus

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.219928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.219928Z digest=sha256:f381e765bc911338833784c0d2297d4d7dac11808b0fe70c2f8a94cc2531b8e3

Observation ea26ca4a-a117-4a82-87d1-9e3eb8a279eb · outbound

This paper cites The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.223984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.223984Z digest=sha256:b9d0d77f5c96d47ff5c5a906ff47608b2b91ce8181435a863f61ed818a2c8b91

Observation 70e6dd0a-b44f-498d-b5c3-5266d235f667 · outbound

This paper cites N., Kaiser, ., and Polosukhin, I.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition N., Kaiser, ., and Polosukhin, I

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.227915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.227915Z digest=sha256:be181aec1a820ffabd4a5c63d9200caee1dedb8d34ff96a863b0450f96ca7e56

Observation 683256ca-aa77-47d4-9a5f-4f8c754714e0 · outbound

This paper cites T., and Li, H.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition T., and Li, H

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.534091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.231696Z digest=sha256:9377d122c8edaa27c2501f6088a39bbbc3a725cf85b83ef326937582ac1c666a

Observation 1e95febc-e8b0-4589-b53d-916f1e2398be · outbound

This paper cites Language-routing mixture of experts for multilingual and code-switching speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Language-routing mixture of experts for multilingual and code-switching speech recognition

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.524068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.235459Z digest=sha256:a1e6f4036fc60da9c227c8bd131972a5bf7fd80510eb45c8e7ce4cb2796322cf

Observation 555f1a1d-5da3-4299-9a20-b307e07dc4fe · outbound

This paper cites R., and Hayashi, T.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition R., and Hayashi, T

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.513677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.239453Z digest=sha256:6878cbb12e90c35a6d55e7e00f644cd74a9ae47cefccf76086d8d3db94b1bd4c

Observation 962a535d-9a14-443b-9c67-34f771193ca1 · outbound

This paper cites Robust audiovisual speech recognition models with mixture-of-experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Robust audiovisual speech recognition models with mixture-of-experts

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.503834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.243307Z digest=sha256:821840c367a655fc5bdc41ec7ed1d58ae9dbc875854ccfd31be6cd02e26e4c61

Observation 3584682c-30ff-4622-9c96-728984916f11 · outbound

This paper cites Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.492396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.246899Z digest=sha256:707246c5154597696553e74274e8ef6e47ea4d161b636b0fe9ea15223c9ce9c0

Observation d8f1c0d4-9dd2-4fcb-81c8-501a463e55d6 · outbound

This paper cites Speechmoe2: Mixture-of-experts model with improved routing.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Speechmoe2: Mixture-of-experts model with improved routing

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.477707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.250437Z digest=sha256:570dac5caa8698b47453ebb9c987a97712783c1e7c75b366f4ff32020511be9d

Observation 08b33f54-e9a5-4b79-b0ed-562a83643f94 · outbound

This paper cites Visual hallucination elevates speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Visual hallucination elevates speech recognition

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.464533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.253960Z digest=sha256:4f71ca8ab512d5e5309843e497721f209c9ae4b2044d0af16a9831e3ed32a522

Observation 5b13026c-fa99-4c0d-970f-b9f68f4aff78 · outbound

This paper cites Self-supervised audio-visual speech representations learning by multimodal self-distillation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Self-supervised audio-visual speech representations learning by multimodal self-distillation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.453215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.257656Z digest=sha256:05103ccc44b81dfbe275b2c2047fdd67a503ee3310e0a2e9d0769c91a60977b8

Observation 3515f5d7-0b53-4273-bafd-96cd12610c8b · outbound

This paper cites Uni-perceiver-moe: Learning sparse generalist models with conditional moes.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Uni-perceiver-moe: Learning sparse generalist models with conditional moes

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.440286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.261199Z digest=sha256:08427bdafc28eedb0407c753f06789daf34c7956f084b522ff5acee2fed31ba6

Observation 0a2625e8-12f8-4dad-8059-86cde6accefa · outbound

This paper cites Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.425112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.264765Z digest=sha256:fd786cea050e091db24c201c8a6bd6e416d613800f99241df84cb05af6181730

Observation b3cea8a1-ca54-4a16-b6fc-d1420e753587 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.269019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.269019Z digest=sha256:a84eacbce774ac468537d830576f867313f97a7a60efbf93bcf5e5d10307859b

Pith citing papers

Observation 10189419-0fe1-4ad1-9f36-c209cea88af3 · inbound

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach cites this paper.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:39:34.345367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:39:34.111978Z digest=sha256:f3beddd91278387cd097b6d86d87a6bc9fd1b0119b1a415b63eee1689503916a