Pith. sign in

Paper Citation Record · LEDGER

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

As of 15 August 2026, this Paper Citation Record lists 100 of 123 outbound references and 0 inbound Pith citation observations for arXiv:2607.16107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16107 v1

Coverage vector

measured 100 of 123 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:22:50.302973Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 123 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14fb054c-6582-4ac5-b322-e69b0d8aa1f0 · outbound

This paper cites Advances in neural information processing systems , volume=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Advances in neural information processing systems , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:40.437297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:40.437297Z digest=sha256:43db3b3febfc3c27fa45d065d14be8e9fd542eb0943bb04434131890f8927f71

Observation 1f2d56e0-1e4e-4816-a437-e5f4d5a5421b · outbound

This paper cites International conference on machine learning , pages=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos International conference on machine learning , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:40.537286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:40.537286Z digest=sha256:116f6d6cf8ac0dc57e9941af624bdf4270f55327d3afe534075a36a2fc2c6885

Observation 212efb0f-8fe4-419e-b877-dfe28fa0303f · outbound

This paper cites Advances in neural information processing systems , volume=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Advances in neural information processing systems , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:40.614003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:40.614003Z digest=sha256:ae87c5d28ab40497f0350ebeb7040de135e1d2fffb911d934eb237a9ef71388a

Observation e1550466-f35c-4afc-858b-2eddcd710cf0 · outbound

This paper cites OMCAT: Omni Context Aware Transformer.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos OMCAT: Omni Context Aware Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:40.689615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:40.689615Z digest=sha256:8dc5a2d3d03069dc7598d9e51cbe3d82756341e34d258570e88a39ef39ac0d7b

Observation e9c2704b-7436-4211-8579-6aca9c607716 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:40.797149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:40.797149Z digest=sha256:41578a71182d066d15114f900f0348501083585e597b05256210462123f059c4

Observation 38d2d651-2e57-4ac8-9475-2ff65d45f4f2 · outbound

This paper cites Qwen3-VL Technical Report.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Qwen3-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:40.901843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:40.901843Z digest=sha256:53757892e2db4b25fe61c909e71ac425658044aa7ba535b101a4487a047b1d81

Observation e61542aa-ca77-444c-a21c-c24e12eed8de · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.000381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.000381Z digest=sha256:89f212b382e5d029a66a60bbf0504940de31779db6a51d332eccd73215e26134

Observation 7dc6ea97-3063-45f7-aae0-def63612fbea · outbound

This paper cites Qwen3.5: Accelerating Productivity with Native Multimodal Agents , url =.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Qwen3.5: Accelerating Productivity with Native Multimodal Agents , url =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.090879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.090879Z digest=sha256:fbf9ee7884f794af8d00e23bcfa1ab6f88f4b9261955a758b7e9e28b616dd39c

Observation e780d499-6b41-4c88-aca9-5acb343089f9 · outbound

This paper cites Listen, Think, and Understand.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Listen, Think, and Understand

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.155465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.155465Z digest=sha256:66b5b7aa723622ef0bb96179c562528a29a1eb11fc490aedbfa2dffef012985b

Observation 3c6d3b1b-44e7-4df3-830c-b9c5dbce2a02 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.231325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.231325Z digest=sha256:0a2d70409cfb092a19c579a37d0999a1e632766fcdea189e83a36d435e6428e8

Observation 3e6c9454-80b9-4e57-8cc1-2b06661cce35 · outbound

This paper cites Qwen2-Audio Technical Report.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Qwen2-Audio Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.305763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.305763Z digest=sha256:264099aacde0eeb80033d6efb94298f7331002ca5f0284d71c2371f0936b44b2

Observation bb551276-05be-4acd-a8ad-5fc8880eebb3 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.376265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.376265Z digest=sha256:cb35cd707404cff3fa5ff29fd175c4d053f85a3aa8a0a84870718a26c425b03b

Observation a7372958-5c1e-4aeb-9e5a-23c64111fc5c · outbound

This paper cites Proceedings of the 2023 conference on empirical methods in natural language processing: system demonstrations , pages=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Proceedings of the 2023 conference on empirical methods in natural language processing: system demonstrations , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.446777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.446777Z digest=sha256:36da978a84dbd106f490addb6b1551b3b0f612f0da22dcbe3415787004496150

Observation bbad4e9f-0853-418a-966d-4f70df750658 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.513991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.513991Z digest=sha256:7acbf2f9c3f469295c3ce87ebd9ad229d1d3fa2ff5ebaa439a0d5e8530f7534c

Observation 81ebca45-58da-47b9-9c4a-b1bb9cbe8de0 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.603969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.603969Z digest=sha256:656b35a8ba2fa6713fef45024332d7aaaad8e46a98502f80bda80785276020db

Observation 20af0184-3800-45b3-b05c-27ec19a6ce3b · outbound

This paper cites arXiv preprint arXiv:2506.15220 , year=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos arXiv preprint arXiv:2506.15220 , year=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.707564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.707564Z digest=sha256:6c1f2a7a9c6cd6a9e73fde44d36b1711746c95756ac20a9c29aeacd659cd51fe

Observation 88feba05-0ed9-4955-99ce-cf8ee90383c6 · outbound

This paper cites arXiv preprint arXiv:2505.18110 , year=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos arXiv preprint arXiv:2505.18110 , year=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.799529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.799529Z digest=sha256:e893038a90c63d5eae3ff4c78db207547470e6583995e6e67b96df6d883b00ff

Observation fa56da0f-20c9-4e01-b3ac-3fe6c8c65e4b · outbound

This paper cites GPT-4 Technical Report.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.875668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.875668Z digest=sha256:117a7eda328fc826f6a4caff44a1486fa0a3cdfb9c0dea93ada855f17906482b

Observation 9566dfce-d775-4b18-98a6-47cc75d1a9c9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.949877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.949877Z digest=sha256:0c7cf2326e69a9bf12bf75bf9cdcc4ebce943f6bffa500992795500d687499b9

Observation 8720fd63-fcdb-424d-838d-55f5e703c9ff · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.018175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.018175Z digest=sha256:90dc517ca4257d84a4c0e6e70f0c51ec325f8b39ff94ec89be55c6f55c841e55

Observation 00795037-c8fb-46e5-8510-11d543ba2b4f · outbound

This paper cites Qwen3-Omni Technical Report.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Qwen3-Omni Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.090172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.090172Z digest=sha256:3821a7886cad26f81430b443750d0cf613deb4bc1ea13ed0003f2464f39762a4

Observation 42b117b3-9681-4ea0-87ee-540ea41917cf · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.159270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.159270Z digest=sha256:2d5b231080c57891bff627dc1f06d86e59239c72fb4ce923de6f350e44cf70b7

Observation 62bb7db4-accb-46de-8675-b296ba13e2e4 · outbound

This paper cites arXiv preprint arXiv:2510.15870 , year=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos arXiv preprint arXiv:2510.15870 , year=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.235632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.235632Z digest=sha256:a46b092b2c84b1c347fd57b910b304fb6d6462f987eba6263873adbfacf4a443

Observation 82b24654-7de2-487f-a41a-b7a01fcb4b2a · outbound

This paper cites European Conference on Computer Vision , pages=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos European Conference on Computer Vision , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.321906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.321906Z digest=sha256:6f5c3ef2214d0a30642ec5a28ddaf04696fa3634438fbbec6559e46078d30b41

Observation 1a431ab1-ec56-4fda-8e91-21f43e848455 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.392444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.392444Z digest=sha256:96b9a8fdeec4be3d626b4952cd4d5dfd7f97a6bad242e7237e1b150176917ec3

Observation c97dc29e-8fc5-4ad4-86e4-5d6e0b14bfbf · outbound

This paper cites MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.469251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.469251Z digest=sha256:fe49c6e9e3c74cbb57a5e277af84b2011fc5f0ca2f35b0e2a1ca2de57ecbd153

Observation b668cce2-0fc5-4baf-b068-67c779f4af65 · outbound

This paper cites arXiv preprint arXiv:2510.20579 , year=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos arXiv preprint arXiv:2510.20579 , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.561452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.561452Z digest=sha256:bac6ab074e04567200d3e9177cf17c1ef1400126396bb598c7befee144f3fe29

Observation dd087e89-9926-4d0f-ab5e-8a5110d19796 · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.659092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.659092Z digest=sha256:384b3b1f32de13d1e490db6c5ba3d3a48948dc141fa1f775e5625ccf72fa6935

Observation 0cc3648a-60ca-40b9-8d7b-722b913a3c4b · outbound

This paper cites Proceedings of the ACM Web Conference 2024 , pages=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Proceedings of the ACM Web Conference 2024 , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.758210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.758210Z digest=sha256:2c993d1696b0411516ce9e2aa49362329bfbd5c72112089202da42efdb3ab941

Observation be4b30a9-dbd6-4ef6-80e1-276ccb3bb29e · outbound

This paper cites Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.836941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.836941Z digest=sha256:419396f14aacb9fefc981e3789552c080381bd9226a8a864b73ebf8897d2f928

Observation 1b580945-e9cd-485e-94c7-0b3bb827eaf8 · outbound

This paper cites Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.917716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.917716Z digest=sha256:36811c725cf04836f3322e1e20394e782da7316215c1928547b62479ce161880

Observation 4b996d6c-e279-4777-b4cb-428b4e22de25 · outbound

This paper cites arXiv preprint arXiv:2511.15848 , year=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos arXiv preprint arXiv:2511.15848 , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.009396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.009396Z digest=sha256:09b5e8f1696eb7ad1252ce9c668288faaa9da15127ece86af5b2b2a413d0c6f1

Observation 8c2dbadf-30ff-4138-85e7-4f09009d0f0d · outbound

This paper cites European Conference on Computer Vision , pages=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos European Conference on Computer Vision , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.107901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.107901Z digest=sha256:1da8b0d60dc51f152b8144ed858a933dc77be642fd5347d0c95b69900607a896

Observation 2797f0e1-d57f-4a6f-8f08-150a2911757a · outbound

This paper cites 2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) , pages=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) , pages=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.209064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.209064Z digest=sha256:ae87f77c208e623f77df4432082c36015e16cfe04b9ef32128e757e421e3a34b

Observation b126fabe-ca66-4027-8f5f-65efd45d83be · outbound

This paper cites Qwen3.5-Omni: Scaling Up, Toward Native Omni-Modal AGI , url =.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Qwen3.5-Omni: Scaling Up, Toward Native Omni-Modal AGI , url =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.299890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.299890Z digest=sha256:b275aeec84e15ad287a2ad9de3c99d8b4a36e412736be2b8733090bc5a65ce6b

Observation 69a30c48-55a7-4e38-b11e-07600fb82e9a · outbound

This paper cites Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.362060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.362060Z digest=sha256:42ec93409c8c7aaa2fb7b4eedfa929a1987c483892ef9096f57b6be073285750

Observation 0c08b50a-da78-4480-9812-f013054d2b8d · outbound

This paper cites GPT-4o System Card.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos GPT-4o System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.444670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.444670Z digest=sha256:e692fddf813f499bb3aba81a30487f841eb93df16ac95a613f0d16c7a9cb2aca

Observation 4da9fdf6-5955-401c-b3f3-4b64f8f548c8 · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.538706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.538706Z digest=sha256:aae881bc29bd7606bd22ee0528f1482d9f67e810a48905f97e45d8701b892544

Observation c7fc28df-1016-4d27-a67f-bab84fb3303c · outbound

This paper cites Long Context Transfer from Language to Vision.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Long Context Transfer from Language to Vision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.630184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.630184Z digest=sha256:2ea5780b0de2db97552da5e3b6a873f0ec4c9a66e53fbf1a3111b4fe30838526

Observation d422a69f-90a0-4bc3-98f1-3403dc3fa754 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.720306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.720306Z digest=sha256:c80b38ae6cfbecec13c3ee6463e79e79f887fd4b388a0bc3b2f127fbe052a38a

Observation 5b48f25d-3cd4-4a78-96a5-bc02eb2089fb · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.783767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.783767Z digest=sha256:a518ce6455a2a036d6602feca2e1f56570918a3196d21c83cf1985ccb23ffb94

Observation 5798ff05-7636-42d1-a667-0b82db648e16 · outbound

This paper cites arXiv preprint arXiv:2511.10289 , year=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos arXiv preprint arXiv:2511.10289 , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.839298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.839298Z digest=sha256:493ece13c741b0ba1f5001d1a350642bd1b88be311cc35221bbb80de199c7542

Observation efffeae0-65a2-4a10-845f-16b9b7c89b5a · outbound

This paper cites an unresolved cited work.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.897431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.897431Z digest=sha256:8ec5648e4fc70e64a73d04fdf8abd3059850fb0f3e167c6d05d1d0c1a3b0b8be

Observation 2a8502f4-d1c2-4928-bfc5-81c09289f904 · outbound

This paper cites The Second.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos The Second

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.984329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.984329Z digest=sha256:1b4ab2bd833ab024bd337a28150641a417ea6502a74b36016179d3561a39cec3

Observation 49f84baf-e103-45eb-ad37-01437851bf8b · outbound

This paper cites 2024 IEEE Spoken Language Technology Workshop (SLT) , pages=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 IEEE Spoken Language Technology Workshop (SLT) , pages=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.045812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.045812Z digest=sha256:b8da5cf92a0b08fab4b2577d770cb71aa54542d4c7b1a712c9a6ea9babad086d

Observation 46ad318d-c52b-4bec-84b1-4dde077421cb · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.107511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.107511Z digest=sha256:ae92e36d2003cb1a246b5b0a57727b33295dbe2cbd89ef045f14004459a9f8c4

Observation aae0847c-5ace-490b-b222-6749d1ae8137 · outbound

This paper cites Computer speech & language , volume=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Computer speech & language , volume=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.171672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.171672Z digest=sha256:2d2f3acfb3a08fde4e981528bf0522b06891ffe693f09bb92d5ba77e2381aea0

Observation 9e293a0b-0cb4-4840-914a-3c64e96a547b · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.232624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.232624Z digest=sha256:c0c79bda3e4a9baef1c16c579ba43aeb5ae8b0a47f813a98d6a1d5292d687d46

Observation 05cbc8c8-da41-4b9c-a717-82222585d635 · outbound

This paper cites Oriental COCOSDA 2017 , year=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Oriental COCOSDA 2017 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.292757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.292757Z digest=sha256:56c58d75918528bd800fb74028c19da62e04e628f0389034bbd7732e47378c05

Observation 41ff913c-4e78-4724-8f72-e8144eac96ab · outbound

This paper cites 2018 , howpublished=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2018 , howpublished=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.378133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.378133Z digest=sha256:dcc2fe7af25a1ae4623421414c4cd1a4619e02aa0c1049d94914dcc23f1aec72

Observation 4c300cbe-48c2-464d-b347-962bb006cff6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.467470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.467470Z digest=sha256:b6bffb5c8235016d20198ac5a300bda77ab38369c9d1b96948b73dc10621e761

Observation f118a2ad-9d96-4d42-ae46-6ccaee531d15 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.534118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.534118Z digest=sha256:5ba443002d36a7244b7c54afe4a47fce4b75facfbffd3b153e1543a5a8c4998a

Observation 8cdae003-f0b8-4083-8e72-274d2a119585 · outbound

This paper cites 2016 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2016 , eprint=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.592091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.592091Z digest=sha256:b9c32b3f0af72e742fa478bcb1da3836d8cf3edcdedb9aa883e57902abfcd7e0

Observation 55484920-d3df-4176-80d9-1493c33d750c · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.683400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.683400Z digest=sha256:0fab76723b697ba9623963398b37db47b8fee47e3bae8d64cc8cc87e6caa6b52

Observation b0b0c5bc-6cc3-4134-b9fb-d3a2511a4d5e · outbound

This paper cites 2022 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2022 , eprint=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.769541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.769541Z digest=sha256:ea0bdd09968e4972b9043087d32db1f0f7113670a9c3250be409cdceee3f737b

Observation cb268a2f-c898-4e38-ba78-d62e7fc9b42f · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.859583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.859583Z digest=sha256:078ce5e36dcf0785c2c490c046bfc60cdd3b911053b3f7001f1c096b5cdc9660

Observation c77d002d-fb4d-4e8c-a5de-3c0f2ec0f28a · outbound

This paper cites 2022 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2022 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:44.946062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:44.946062Z digest=sha256:cb71d473e76a65e745757423fc82e860b5b8e80ac42a93ee942c57d55b8fe836

Observation 8c2a1083-c05f-440d-be4d-63897827ef4b · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.009092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.009092Z digest=sha256:6ad3e6b288fb013cb584d8a5f087abe6539058f5f7c333b14a8b6fd7c37ce6ee

Observation 96317226-48b5-4547-8fd2-0670fbb0196d · outbound

This paper cites 2023 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2023 , eprint=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.082704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.082704Z digest=sha256:2822a2f39ee88f951e793b7b279d0f4607f5a31501b754c7e4ce3eb75667b4b6

Observation bc54388a-5216-4ac3-a44c-c8d17e240de1 · outbound

This paper cites an unresolved cited work.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.145377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.145377Z digest=sha256:2014aebfd1263ceed99355db5fe09ec729889fef774e5169cfed91292b19a59c

Observation a90ffa87-d61c-4411-b4c8-5627d63b2d05 · outbound

This paper cites and Ellis, Daniel P.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos and Ellis, Daniel P

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.209967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.209967Z digest=sha256:4c8f190e9864a924631b94bf45e783b0bb9b816d5803b0888d086de6c47668b2

Observation 87141554-f6fb-4326-9ab1-e1830a397d5a · outbound

This paper cites 2023 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2023 , eprint=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.281555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.281555Z digest=sha256:a780dec27a14479be6f861d4dec17a0d3f659a732b0071c1dffb5d60c1998f91

Observation abee88e4-6156-40ad-bc02-041525c323f7 · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.364742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.364742Z digest=sha256:7e3ee5cac5712e018b1d0180b3590dc7bf4d7fa9a68d739587ecb715f8892132

Observation 4795bd32-d9b0-4ff3-9ce9-3f4068d7a821 · outbound

This paper cites 1993 , howpublished =.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 1993 , howpublished =

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.424894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.424894Z digest=sha256:f4e62873aa351ca7a54fface042c70ef1bede510b202f9ee83fdead650af2117

Observation 74ae0bd6-beb4-4d4b-9c04-8eac09ef98c5 · outbound

This paper cites 2023 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2023 , eprint=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.472290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.472290Z digest=sha256:7463a581907a1e02f81c21794215ca111ad229af02168aed34cb9eae7c4d0562

Observation a821ec62-5c5f-457a-a95f-89cb3ccabd49 · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.559470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.559470Z digest=sha256:b066725d936b30f2ad47643e224823afa2146c7fcf6aba599c82b94207fee7ca

Observation 923eb1b6-3088-44ce-a88e-d911f8792df5 · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.616519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.616519Z digest=sha256:bc108455963f3df5e185a84b8f7817cf2ae316b9bf640900675bdb3bf2b03dc6

Observation 93b0697e-ef99-40c9-a76d-4fbbf350dec7 · outbound

This paper cites 2026 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2026 , eprint=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.702118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.702118Z digest=sha256:dee76a763ae5ef3371f85f8a95a84bce81904ef4d11453d748642b664202f966

Observation 9d2fec60-061a-425f-b866-b61bd84377ec · outbound

This paper cites 2019 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2019 , eprint=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.761624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.761624Z digest=sha256:446a3715cbe2266dd738b156bd122af71c603c135e1464427afa8ecce23e58d9

Observation 38f5dbcf-7489-4cd0-a4b7-e407b2a81ca1 · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.846104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.846104Z digest=sha256:68aa3b6304eab7aa9c324435b2fb5f5dd92b1a9dfc9774ad4a9cb42391e3a3b3

Observation fb56dfb5-0815-48bf-bac9-41b1a072b9ff · outbound

This paper cites 2026 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2026 , eprint=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:45.937648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:45.937648Z digest=sha256:2eee5e116963cc29f4fd888a8bb81e43d443449002a8c7dd588472cb8ec1eb75

Observation 95dd7b8e-426b-4df5-a651-3366a78072b6 · outbound

This paper cites 2026 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2026 , eprint=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:46.024191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:46.024191Z digest=sha256:597d00e5d124da49283ed44543a5ab311f8da1508bc4303d6d5b0dff612f1b0e

Observation 26c99fd9-73a0-4e17-b117-54f1eb66f30c · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:46.111036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:46.111036Z digest=sha256:3557c46abee266aa965290aebe47d85d561fe033604249f81a24fda8e779a82c

Observation 41121929-a1b2-483b-bec2-e4395bd05015 · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:46.171684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:46.171684Z digest=sha256:550f4f4faad727570592cc0b384169d5e046a4978b5e4dd0381906eb9db8bb3e

Observation 90a1066a-c69d-401a-85a6-b0b9cfbc3d25 · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:46.277747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:46.277747Z digest=sha256:22301c3e0d347afb51cc4a0c274489c4481f9d3341f5e60dc1f204c3db67a6c8

Observation dc2538a6-b8d4-4e69-9772-ac541b8629a3 · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:46.445898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:46.445898Z digest=sha256:76cee6da3f820ba739c02286335441191778a3d89f606c2f316e12883d8a6a3b

Observation 7b77b87d-a174-4afa-b4bd-9e8645427cb3 · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:46.620323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:46.620323Z digest=sha256:e06d27e213b41e3772c977eee63ce781533448db6175d89c067c28952810e46f

Observation bb1b1134-9667-4379-ab2f-c82d1a8e444d · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:46.713585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:46.713585Z digest=sha256:b1989ea57cbe7dd67fde7373cb0c6efa796a6f1eff8b2c2e97c2f45af02f64a0

Observation 9af27be9-a66c-46ec-807a-ac2d6fa845eb · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:46.816255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:46.816255Z digest=sha256:532d5f1a3ef3c4654b6e83ed396b1aaebd2ac7a48b18549d75c24b436edae7ea

Observation 747537bd-1831-49cd-9a3a-063077568257 · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:47.016118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:47.016118Z digest=sha256:ccd2fdac805fbe4dd4b0b68609d4eb8cc87cba7dc4f8cc47b4324103776aa52a

Observation 48a0ef75-764a-4f0c-aac2-29abcb18e807 · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:47.128477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:47.128477Z digest=sha256:5cde1d92afb9d23743f7862493455a795c6500cb16c2605a84a45cd96b39f2a4

Observation 92a158d4-2ce6-4f4b-98b4-aeab81bafd32 · outbound

This paper cites 2023 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2023 , eprint=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:47.266624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:47.266624Z digest=sha256:b3a58f7c15a98e2b420035dc3fe7be2a4e124bf8e162cd0cf7ae3b3a0c020c60

Observation a5ed575d-fac3-4920-b397-22f58c0e0c17 · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:47.382179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:47.382179Z digest=sha256:01f1d1909f0d9ceb03c5084c679b6fe312b3b4b3eeb7753190abd067869b99b0

Observation bae5297c-fa4a-4f60-92ad-31438f0f95c4 · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:47.556138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:47.556138Z digest=sha256:adffe446527b5d8dade4677b8f6c5355daa84434efb6a07a1b98f2b3c4f0074d

Observation ec011631-a1fd-4f6f-b15c-6a5e86328e2a · outbound

This paper cites 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages =.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages =

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:47.829504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:47.829504Z digest=sha256:b39bb4e7d5d35ec779abfa7d5fb55c8e00a004c92531e3981493c6045cd48090

Observation 8dd6dc99-3b28-4aca-b7b1-57ff937b624f · outbound

This paper cites 2021 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2021 , eprint=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:47.996882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:47.996882Z digest=sha256:11f68e2154d4bdb235c3d9c700c8cd1dfae35a088ef15caefae8f791c824f01b

Observation 8087c9d4-4ae7-483f-8498-6bc733bab06b · outbound

This paper cites TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation , ISBN=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation , ISBN=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:48.154855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:48.154855Z digest=sha256:b78751b437d610b14ff78df2adcab1c14726c73684b32353e93fa50bdf20af85

Observation aac1044a-e4cd-47a0-a286-6cd07625d6f0 · outbound

This paper cites GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio , url=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio , url=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:48.332408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:48.332408Z digest=sha256:25fbee8145f00ec20fe1f6bf9809d8f3bba8053231ee09acfb81ef5aae318ab2

Observation 22294c93-2c8f-4ebe-bbcc-026e02346527 · outbound

This paper cites 2020 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2020 , eprint=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:48.511479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:48.511479Z digest=sha256:54eb7cc9bb07e4258331d90ee309160645074d0d8783773b2e99dbed567312ee

Observation b4fd878e-699c-4e92-a96f-65d40b7122b5 · outbound

This paper cites 2026 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2026 , eprint=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:48.688062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:48.688062Z digest=sha256:38c0246110c1904d66a38e9e06fa43c262b6e7c512802f0214cd4063c4873741

Observation e86bdea4-523f-4b6b-8dc5-e99f2d9c31ef · outbound

This paper cites 2026 , month = apr, note =.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2026 , month = apr, note =

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:48.860270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:48.860270Z digest=sha256:709be7b6f0bcebbbcfa1b2ac8bb13f82d59ab78349120b75176c30b23a4b1027

Observation d368d464-f875-484a-b7ae-915e17dc4df7 · outbound

This paper cites The New York Times , year =.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos The New York Times , year =

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:49.004520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:49.004520Z digest=sha256:23aecb4856fa35f5cb3fc2faa60eac934cc8ec69dfa0b36a12dd93727248f875

Observation c6a0ac69-9de6-43d7-8ae7-c91326b3b2ed · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:49.190620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:49.190620Z digest=sha256:a5daf13c300c6b1f80cc9041ad19f8be62c3845b62949453b78ab4ee492d6f4a

Observation 1c6c62ab-426c-4e75-aa32-036571d639fc · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:49.328060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:49.328060Z digest=sha256:ca50c20c7d1685b26defb56b25f6b640761d9a5143af3d6c5526ac3301cdbcc3

Observation b2c56403-fc92-4fd2-a36f-5a6b49e17f9e · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:49.466034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:49.466034Z digest=sha256:9aae3845744ae389ea99177dae9e1c0f7d40904d2d156b22ae561c20157692f1

Observation 0da4c31a-aafd-488c-89d3-fb9f96dc2ad5 · outbound

This paper cites 2025 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2025 , eprint=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:49.630807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:49.630807Z digest=sha256:1a20bccd83197051d7cd6981fc9879bf162a65bbf708523d7e1a0bf03e6d6703

Observation 182a556f-9aaf-4bb1-890f-e7f4b798a4e0 · outbound

This paper cites 2021 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2021 , eprint=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:49.828306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:49.828306Z digest=sha256:c0f778074dfa46efcd6918e8086d2d0615c14f080ff7cf392a02029812ad6753

Observation 19310eec-35d7-4840-8a8a-fa0e767e5073 · outbound

This paper cites 2019 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2019 , eprint=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:49.999541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:49.999541Z digest=sha256:ddc491ae890dd991149f1e740611ccaf3f595691b7a1619dce00cf69890d23e6

Observation e4a5fd58-60f4-419a-ae91-ebe3d1e9d9dc · outbound

This paper cites 2022 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2022 , eprint=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:50.151413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:50.151413Z digest=sha256:8b6ad72b89c0fa349a263b25ab13ebae1bed5c40b3de80603743cbe9fd62f2f5

Observation a26cf0f7-ca21-4ff5-8627-43deb295b4c6 · outbound

This paper cites 2024 , eprint=.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2024 , eprint=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:50.302973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:50.302973Z digest=sha256:0cac49aa21ccb6785abc6af6795f2740e74d9b02cd1dbc95d82bfb5abb579c6c

Pith citing papers

No inbound Pith citation observations are available.