Pith. sign in

Paper Citation Record · LEDGER

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

As of 9 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 5 inbound Pith citation observations for arXiv:2506.13642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13642 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:19.591854Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T06:35:35.951554Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:38:39.602552Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact3
  • verified fuzzy30
  • unresolved44
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99f2d984-0e85-404f-af46-d47bad206dc3 · outbound

This paper cites Hello gpt-4o, 2024.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hello gpt-4o, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.616625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.211972Z digest=sha256:1ee6e3e477e190854c6876600b4c58527e2c507c9610919a08ae70da4b5f4e72

Observation 4ca0a7a8-44db-47f7-ba61-10986922dad3 · outbound

This paper cites Gpt-4v(ision) system card, 2024.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Gpt-4v(ision) system card, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.579132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.217907Z digest=sha256:b86e597d173123fd2cb735ec26851302b54f91603f2434d9d13168ca688bcb71

Observation cb87d4dc-dd24-4c03-98d4-d07b0c69ad42 · outbound

This paper cites Visual instruction tuning.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Visual instruction tuning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.534744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.222673Z digest=sha256:3c95af85cc57ae7b99896ce5e8ef3aa26989479c27e40e55cb1028c8f72b966e

Observation 491b9874-b198-45ae-8e7b-ab812f1ef63d · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.490643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.227705Z digest=sha256:602525e314f70599cd74c19aba0c66e4acb1189a307689ce07fe594bde9684ad

Observation 5c78f174-2af0-4e89-8c5c-5555031f56af · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.232753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.232753Z digest=sha256:558b960bee4b98411a74f932257151263c63aa67ec624b033ff4cbece0449016

Observation b9c087cb-4c3b-42f8-9a46-81de19f1e5a2 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.237792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.237792Z digest=sha256:26ac04b169a41e0aeb6ae37478b7084edaffcace7ec514621f14d6b303731810

Observation 1ab3b4fb-bae9-4a7b-a7d3-1e62ea7e4305 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.243615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.243615Z digest=sha256:874530370e70a0509bbeb8838fb4b223874b24bf203c76dc08bcbc2c4b3f84a4

Observation c1b7e80d-0a19-4bc4-9c34-91e24f446e9a · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.248709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.248709Z digest=sha256:030e461983c25c13202a2bcafccfbe4982284e3daf77c55d64700d0d7d7c09b7

Observation 7448f4ea-2857-4560-beb2-e3769993f9bd · outbound

This paper cites LLaMA- omni: Seamless speech interaction with large language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA- omni: Seamless speech interaction with large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.452144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.253606Z digest=sha256:4f2ddbbbdc953edc7dfa40aa238e8c9f8fd2464a4366ba4c93f01982b03ca21b

Observation e1373ac1-5099-4818-8963-e2ba1604d6b3 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Moshi: a speech-text foundation model for real-time dialogue

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.258298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.258298Z digest=sha256:c98101fe65d072ce959aeaf9097ea2572b59cc0acc5386362cb1362a12be75a8

Observation 171ac415-b9c8-47de-aab3-1bd0b16a5aec · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.263029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.263029Z digest=sha256:d5baf9c0ece4c1b62adf768d0987b24e781e654f97d89dfe57870333e0e9ad20

Observation 3c0249fb-3f19-4f57-bc5e-ef16e778e3a3 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.267903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.267903Z digest=sha256:4c2323eb9271b4af6a1dd96b273d931a2535cd8420211a8611dec22d666be848

Observation 00c9d13f-5346-4bb2-ad89-999189f95508 · outbound

This paper cites Baichuan-omni technical report, 2024.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Baichuan-omni technical report, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.273419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.273419Z digest=sha256:c889b618ae3ad075f6565dbd2c33dd2e7ea2f01fc6b34fee88b49fde207269d1

Observation 0adcf620-700a-46e1-a44e-26ca606ba35a · outbound

This paper cites Qwen2.5-Omni Technical Report.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Qwen2.5-Omni Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.277910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.277910Z digest=sha256:f7fa3142de2ce10a0da5d46585e65951a28bc146af7888371641cff5d0cd81cf

Observation 88f62f36-3081-4824-8222-4e2ba58ee3fe · outbound

This paper cites Connectionist tem- poral classification: labelling unsegmented sequence data with recurrent neural networks.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Connectionist tem- poral classification: labelling unsegmented sequence data with recurrent neural networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.282753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.282753Z digest=sha256:1fb427805614224be2b127daa364e47bce0fc0247e083fef43f50822b2cb7b0d

Observation 77b2176b-144e-4162-aa55-adcd1168d97c · outbound

This paper cites Learning transferable visual models from natural language supervision.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Learning transferable visual models from natural language supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.431137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.287814Z digest=sha256:a3414eb15796b82023d3127c50d5da5b519c800f9dc69fbb0c3c487d6f372d04

Observation e545fb63-b739-4379-885b-1e155e47775a · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.412866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.292479Z digest=sha256:2a16ac84ffdc233b45ae6ee4487be5f9bb20f56c1f846b6693915a45d091b2a9

Observation 162c306d-abff-46dd-a029-c6edbdfa953e · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.391181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.297381Z digest=sha256:08609abd90e0a0f4ab16143f0a56d887b49884e308d832bb43b0b69dbeb512c2

Observation 16c684df-b327-4f82-8bb9-7e17c22b4528 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.302280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.302280Z digest=sha256:e659752ff06d5cdba6739f288aa0a13b99c0b79987008c103d9683868d9ac1b3

Observation 2828e8b0-ff6c-4883-b663-885ba785378b · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.371372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.307062Z digest=sha256:45d1663766440c8138f31f42cb1769e6a3729e98a3af76ee5b1b7c718017d09a

Observation 90ff68e9-ab1c-4fff-90a3-784606db2304 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.311995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.311995Z digest=sha256:7c5ab3647f9228fe5200270af2909d2f478b231227c8e0a31a8eeb8895c45a8c

Observation bd4ad7ce-5871-44c9-9b2f-0bc008ee470e · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.316833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.316833Z digest=sha256:899cba75ff99e70fc2ae9a19178fe7beb54b7b17b13f82cf702b6d63a420c7b1

Observation aea73c1b-d24c-42f7-9447-41c286b980ea · outbound

This paper cites Video-ChatGPT: Towards detailed video understanding via large vision and language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Video-ChatGPT: Towards detailed video understanding via large vision and language models

Reference 23

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:33:21.348872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.321745Z digest=sha256:6fcac1d6e5f044722057441563b4093c30d024cbf8afa748fa4410f5b96d1406

Observation fa43ef72-8084-4d98-b812-6fccee3b5343 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.326679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.326679Z digest=sha256:336cff7f0731a879d8456043d31d8067a33ef4bf593198cb9f02a708c1a989e8

Observation d03987c5-bb12-41b9-896b-94b8880e4c21 · outbound

This paper cites Video-LLaV A: Learning united visual representation by alignment before projection.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Video-LLaV A: Learning united visual representation by alignment before projection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.331593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.331593Z digest=sha256:54feb5ca99246ae3823fd02dcc13e8e4e36b618cf72c3cbe611851bf8336b5a0

Observation e40c7e43-5d2a-4814-a945-19f891f788fb · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.341323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.341323Z digest=sha256:24ae459f8e4a5938bfc372f378191c3fb52aa6428c1ca55a6e423a65fccb70ff

Observation 09ba4a5e-31cf-4a00-89cd-cc489715c8a5 · outbound

This paper cites SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.346344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.346344Z digest=sha256:71a464c82bab55d62a5e7ba0bcce97c50f8425aee8e9bc76a430dd80f6e2a37b

Observation a75a3f1e-d172-4cc1-a04d-c8d055653d95 · outbound

This paper cites Slam-omni: Timbre-controllable voice interaction system with single-stage training,.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Slam-omni: Timbre-controllable voice interaction system with single-stage training,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.304020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.351418Z digest=sha256:d4b73930da7aa6e92beecac8d108cd715277caa88cb1d25251c0cd468c619486

Observation 7e87629b-a43e-499e-b0e3-1baac53bd90e · outbound

This paper cites Robust speech recognition via large-scale weak supervision, 2022.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Robust speech recognition via large-scale weak supervision, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.278885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.361775Z digest=sha256:b5649e3a7a7e6329cfd9faf5a3c074835122f85228f69571c865a5278e93d943

Observation 20ba2d0d-0fac-499b-a766-14073c675d5b · outbound

This paper cites SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.356530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.356530Z digest=sha256:7a00f1ea25cd12dd439888f93878e44fcf73620b1958d3844b0aec30446d61ac

Observation 76a7d486-85e6-46e4-a154-e4a7113a5e5d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.371514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.371514Z digest=sha256:1e8fe2fe7a6957fc4930d5c0694f234a860baca4727dec34b0fd94cccb888680

Observation 364e34da-2e2e-4130-ae6e-9ecf434d071f · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.366716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.366716Z digest=sha256:341b871d8b29967e6e8ced6e98e540237549f634e6303e65dcabf3efb0cee775

Observation 40a80281-7f94-4d0c-a57c-851ee32bd298 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.381385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.381385Z digest=sha256:c5e33411597e8c7523e5685e8d7021c5666a2de5d4fc5d50de17a153caa5d918

Observation 1393f3cc-df50-425f-a3c8-4714f50a3c8c · outbound

This paper cites Speechtokenizer: Uni- fied speech tokenizer for speech language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Speechtokenizer: Uni- fied speech tokenizer for speech language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.255218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.376928Z digest=sha256:9f610d4d2c9a7b943277b34c01d22e7d833249bae23c11f78ec2165c058b0f2b

Observation 49e0837c-b280-4bda-9abb-ed83e6a41216 · outbound

This paper cites an unresolved cited work.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.390752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.390752Z digest=sha256:dca5e56af1e9f28cd5c3ff2e7505c52df11c9e67ea4f7fc53988cc3455a18623

Observation 9b572f75-01bb-4bdd-8a94-1835bf8faf8e · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.212324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.386288Z digest=sha256:ca906661f015cf7ac168ee19c1b018520f0dfdb8dd1db51229c33f1850fc80be

Observation 6907a3a9-4e83-4aa4-a288-2d6d4560034d · outbound

This paper cites M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.400418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.400418Z digest=sha256:599605cdac47a2b3da30e7bf999c477c403ce363a1f642a122cfca930ccae0a9

Observation 79c360b9-12cc-47f9-8e9d-f0c67c6d7099 · outbound

This paper cites Megrez-omni technical report, 2025.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Megrez-omni technical report, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.191840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.395633Z digest=sha256:ea422b8bff693bd0ae47747ffe2f053204355c1b3a52e25d014c62715caac595

Observation 3f5707ab-84f4-4a61-aa79-9f94835d864c · outbound

This paper cites EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.410942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.410942Z digest=sha256:f2c96a61811b322c2b5f2ce12fd2a678db84fec35197a62b61b84b68b086b35e

Observation 49d2bc70-bf19-47eb-9c8f-432f1887f8fb · outbound

This paper cites Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:33:20.191604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.405533Z digest=sha256:cbfa8bc9d70506d87710ee84ee3670b0f0af2a83fa98be178382263d4e33c461

Observation 5f8eeb56-c015-44ca-b3ee-76677e515931 · outbound

This paper cites Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.420806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.420806Z digest=sha256:fe985408375b05b4e0d527fbe4856d8d13e41f965808826cc3e4dddd29cc115e

Observation 2337fda3-cc8c-4241-b296-8134f263461a · outbound

This paper cites Openomni: Advancing open-source omnimodal large language models with progressive multimodal alignment and real-time self-aware emotional speech synthesis, 2025.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Openomni: Advancing open-source omnimodal large language models with progressive multimodal alignment and real-time self-aware emotional speech synthesis, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.416428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.416428Z digest=sha256:3bde34e8b1f79785ef5096853c66969c9436e72b0aa817e2d4f0a54a35524e0d

Observation 9eef4a63-d46b-45e7-b69a-d26e6226ab37 · outbound

This paper cites Stream- Speech: Simultaneous speech-to-speech translation with multi-task learning.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Stream- Speech: Simultaneous speech-to-speech translation with multi-task learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.430154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.430154Z digest=sha256:5bf59130ea0e84608c91f64c4799f05bf95476f1d17660ceaffbf2a13f2954af

Observation aec48e8a-01c8-456d-98c8-0c62539af473 · outbound

This paper cites LLaV A-mini: Efficient image and video large multimodal models with one vision token.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaV A-mini: Efficient image and video large multimodal models with one vision token

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.172880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.425639Z digest=sha256:83c11a75b6d2e73494b687ccc36b6ed483978915dbd6db3800f3421f3f939668

Observation 7b3b7fc3-c782-4dd2-9b12-c42d5f2977ea · outbound

This paper cites Future-guided incremental transformer for si- multaneous translation.Proceedings of the AAAI Conference on Artificial Intelligence, 35 (16):14428–14436, May 2021.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Future-guided incremental transformer for si- multaneous translation.Proceedings of the AAAI Conference on Artificial Intelligence, 35 (16):14428–14436, May 2021

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.153594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.439367Z digest=sha256:bdae7b2599ba3b80f305e7c125b8aea7f6a1a1d2a6176578964acbc6eb2f0d40

Observation 0d4c36c6-dd8c-4162-a85d-e910384fbe28 · outbound

This paper cites STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.434764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.434764Z digest=sha256:720a2994f1f9cd9aecdfe32ad3b954cf18ce78bfc750d33222fd8bac5c08d43f

Observation b886f9d2-8782-415d-98a7-ca81cdb87e1f · outbound

This paper cites Information-transport-based policy for simultaneous trans- lation.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Information-transport-based policy for simultaneous trans- lation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.448868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.448868Z digest=sha256:61ee73952d6d249788648a88340055772390239be4e90b95dbb96a188a0d658d

Observation ed288ccc-6968-47fe-9e67-c6d5fdd2ceaa · outbound

This paper cites Universal simultaneous machine translation with mixture-of- experts wait-k policy.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Universal simultaneous machine translation with mixture-of- experts wait-k policy

Reference 48

Resolution
verified exact
doi, observed 2026-08-07T00:33:19.693942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.443755Z digest=sha256:c78e8a76a4bd6e980cc176fc16571117612d9b3591441a41334994cec03c29ce

Observation 030d498b-c61a-4151-8daf-91be67b987b6 · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Librispeech: An asr corpus based on public domain audio books

Reference 49

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:33:19.458294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.458294Z digest=sha256:ee4fb7faff33d502a6d9d45790d46a16cada4377068ee6cbec1a32f814a63e37

Observation cad4eaf6-cc97-4b89-922c-a8f33a02d995 · outbound

This paper cites End-to-end simultaneous speech translation with differentiable segmentation.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model End-to-end simultaneous speech translation with differentiable segmentation

Reference 50

Resolution
verified exact
doi, observed 2026-08-07T00:33:19.662287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.453507Z digest=sha256:6c4e0d33a7a2b44cc33ae50eec40341ab5b42816992babe03ccadbafd63dfa7e

Observation 9ddbca23-aa03-4346-b39c-71a5b8142f41 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.133409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.467939Z digest=sha256:562d3fa3dbb6a90a393a5b6234a78692f517928ab75e48a4d08b734a5ac6a75b

Observation 8b54d6d9-9aaf-4136-a244-4645ab435966 · outbound

This paper cites WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.462794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.462794Z digest=sha256:676895ba9e4da7050664ac6b4d43eb5dd091a37d09f38d0b146a0a1a33999de5

Observation 47529764-308a-4bec-98c9-0cf8a7c46da4 · outbound

This paper cites Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.088297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.477513Z digest=sha256:9799f1ed810cb8828d46addd6339cb6e5b57548b1dfe190452accbcdf828358a

Observation 15dbb48e-f8d2-4abf-a228-c45ec74ff014 · outbound

This paper cites Hudson and Christopher D.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hudson and Christopher D

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.109643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.472638Z digest=sha256:6d17209b2d71e9549db889fc0060287161fcf6b8495b36a298d623d9ca049a57

Observation b4715a3b-9719-4bfc-8ba2-92d0d69d98ad · outbound

This paper cites Towards vqa models that can read.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Towards vqa models that can read

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.044743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.486601Z digest=sha256:1f9cc2687a5d2e9559c7f35ba3fe759edca23646cca95f5f2c58f779a7cee735

Observation ce3a6731-d55c-4499-acf2-7f72fe21f344 · outbound

This paper cites Learn to explain: Multimodal rea- soning via thought chains for science question answering.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Learn to explain: Multimodal rea- soning via thought chains for science question answering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.065463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.482086Z digest=sha256:d8edc3e2bf67e79ef3e6b29dd1639ca9acefb53135a82a30fc266cbe2f017a49

Observation 9ab882af-001f-4b7d-9ed5-96fe76b00cae · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.496435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.496435Z digest=sha256:972608443158ae7dd07a65ddb9a4d006ec0e44343b5798ffb8c8a6a9b6127cd9

Observation f8475a16-8b23-45cb-a4c0-978f7933fbb5 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Evaluating object hallucination in large vision-language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.023280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.491754Z digest=sha256:ffab239027ddf0458010e17ecadc9df0203211f38386d26d6215e4e90af7f77b

Observation 6f30e2b9-f600-4094-b497-c145b6453040 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Seed-bench: Benchmarking multimodal large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.004498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.506706Z digest=sha256:4212397b4f6fad9dcb57d226e8d2ddbd30eb7425dfc2e095b38a40c67fedb6ec

Observation a170ddd9-daa1-4f37-8cb4-3e6b391b00ef · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MMBench: Is Your Multi-modal Model an All-around Player?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.501199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.501199Z digest=sha256:37c22900de8c18cf32f13e7600335f15fe5c9031cca9931c21ab53489d52d904

Observation 0e0579f0-f8fe-4228-ad87-283cd2471329 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities,.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Mm-vet: Evaluating large multimodal models for integrated capabilities,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.972598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.516384Z digest=sha256:cd12bc2e6d3a379fa74c6313bfd7f6a52b4f230c7c7a631e824781dd763f486f

Observation 783463fc-6429-43de-aa46-6cdf7ab10579 · outbound

This paper cites Visual instruction tuning.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Visual instruction tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.511589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.511589Z digest=sha256:552f8d7e2992c41c7080cfde18126ca378b1961ea6078218d1dcdc3a9cb1090a

Observation 5a224011-d38c-4c0b-8e36-661891981140 · outbound

This paper cites Semantic parsing on Freebase from question-answer pairs.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Semantic parsing on Freebase from question-answer pairs

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.935198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.531289Z digest=sha256:e0f45f716c85bd352579db64d096d2b03baacb5f1cb0db3adb2281fd1c8420b7

Observation 3af33f61-c3d4-48a9-bffb-067b3f7454fe · outbound

This paper cites Visit-bench: A dynamic benchmark for evaluating instruction-following vision-and-language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Visit-bench: A dynamic benchmark for evaluating instruction-following vision-and-language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.898083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.540625Z digest=sha256:eafa0e16871a2ef81617f1b2dda3cdc165f06b4e79d19f4e558d68afeb50e5b3

Observation f5487bc2-ab3e-4bc9-b9d6-0d435ab04a66 · outbound

This paper cites Spoken question answering and speech continuation using spectrogram-powered LLM.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Spoken question answering and speech continuation using spectrogram-powered LLM

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.953856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.526568Z digest=sha256:d825f8a2da58eed814de68e0d57f97f47334d98e919df2e87e3cecf7da2963b8

Observation cf89375b-0d49-4cba-8248-1ed91245a399 · outbound

This paper cites Improved baselines with visual instruction tuning.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Improved baselines with visual instruction tuning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.549650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.549650Z digest=sha256:6db12ff8d410908963b43a3ec71d9c3476f08d44ad4ab8c6580b24ef937cd924

Observation 40133eb9-41ff-4f21-979f-69e2e2a4f8ce · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model, 2023.URL https://huggingface.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Introducing idefics: An open reproduction of state-of-the-art visual language model, 2023.URL https://huggingface

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.854697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.554791Z digest=sha256:59041922953527e492f875c2b7fcd6ff6007745b7e150b1b89947f5e15271859

Observation f2e0cdb4-aef5-43c3-a9d0-eae9f60b6d8e · outbound

This paper cites Textually pretrained speech language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Textually pretrained speech language models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.835643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.559359Z digest=sha256:e064022c6f137aa496fadd1de237f691afea4df20112abd2792b6ee707963036

Observation abe7db92-287e-4465-9595-8a98817694c7 · outbound

This paper cites BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.545187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.545187Z digest=sha256:2e22ce63d6db4c8416915aceebb3274dbda6239d10a5cdd8a3d4b7154ee6ac7b

Observation 1cfecd7a-cb94-4f44-ba50-8e016548ba32 · outbound

This paper cites The Llama 3 Herd of Models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model The Llama 3 Herd of Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.568881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.568881Z digest=sha256:9cd055301d50fa3320d14f381bbfd7c93753fcb2c60b3700a4a998f44b4f50dc

Observation a119a8b9-f13b-4860-8699-792aabae8383 · outbound

This paper cites Sigmoid loss for language image pre-training.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Sigmoid loss for language image pre-training

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.573384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.573384Z digest=sha256:85cfb194674ff06f79f094b57a39dd02a2dab605b3558354625ab0931455319c

Observation f5e67540-2a28-4247-9b5c-61fabd5ceea5 · outbound

This paper cites LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.577773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.577773Z digest=sha256:f8280067f1f4d26f5f02138b97405fd18a7c4d9155f1d1c442ef7dd9ea992536

Observation 387faf8f-1424-435a-b860-dc09ef65783b · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.563774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.563774Z digest=sha256:10a9072b85afa71491a386dd1b64a246f8dbe6e361367db225c5e72f6bc7552f

Observation b913fec9-ae1a-4c18-bfd8-c3f2b8284f66 · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Enhancing chat language models by scaling high-quality instructional conversations

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.586908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.586908Z digest=sha256:50f57450f261772795c5cf06ae1000e4a03df0ef6500c4d38e3f493c56c91181

Observation 22b2724e-132d-4889-9947-a15400715054 · outbound

This paper cites does not allow traveling to the second floor.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model does not allow traveling to the second floor

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.591854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.591854Z digest=sha256:fa13a6779e921fb38f470bf365565a326d789f850cb15a6c938abbd53f9b8094

Observation 7d2c1a34-e4c4-48c0-b0cf-f4ee13d27e4e · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.582482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.582482Z digest=sha256:4513c56c4815180fe0a6129a74b26fc8d60f8484f632b12bce9fb16ebbb8e990

Observation 8c062e03-716e-412a-8997-ad76e25128a1 · outbound

This paper cites URL https://aclanthology.org/ D13-1160/.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model URL https://aclanthology.org/ D13-1160/

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.916950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:33:19.535826Z digest=sha256:c105e1562143778efd2bea505d0246c597cb4aad6e058e2229e6777ab8b4a89d

Observation 8ede076c-5894-4de9-8ed9-890ce438eb60 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.521264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.521264Z digest=sha256:f30ac82322ada4b5f1b37f5b5df2e7c4bf45eeb7b8fd467dd1ce92c1edd6fb2b

Observation cba64abf-8252-40a7-9f0b-0e0ebac5daca · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.342.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model doi: 10.18653/v1/2024.emnlp-main.342

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.336391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.336391Z digest=sha256:82b62cfd3f96995720c0244678c42f319ac3df74f09f6da836fb6ab84221c0cd

Pith citing papers

Observation 286ba852-3dfa-43a3-9a01-4abb28a5dfab · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.168111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:276338b9694a3114ac6ac771be585d1cb5b68703bfabf621d3cf6eaf4489c034

Observation 20d11b5b-424b-468b-98b5-45b0698792ac · inbound

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory cites this paper.

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:59.686225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:05:40.242925Z digest=sha256:d77c3e631bedb8b2015db4b6aa0fb73b6d25df342ac54e4c8ec29da177d6297c

Observation b3a8cf48-e086-43d3-bc8e-60ddc7414b0b · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.634543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:8f7465e4edac2de298dea930368ff4b0642cce7d61709e3e55a9729c009c8748

Observation 95af9d12-ec95-40ab-8db1-5aad1f844823 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.603859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:67a122ffe7666a07e49cf133b218edc055eda24f117a7277c731b19b0299194b

Observation f28a4588-daff-4cac-8c3e-8020f4f57764 · inbound

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory cites this paper.

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T06:35:35.951554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T06:35:35.951554Z digest=sha256:1ccf44007c29cef2c38eb49ba0b1605bccd3636c0111e30167846abb7be507ee