Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:56.898917Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 37 inbound Pith citation observations for arXiv:2501.06282.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:56.898917Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:52:07.794411Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
69 of 69 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 4d43ee89-7e06-48a8-bb4b-0180550ffa80 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc597275-f6ca-4605-b446-e077a9119a7a · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa3bcb17-216f-4dd9-9684-62e5fe3e175f · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7a839c8-45f7-4aad-af85-79b48f825d85 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Common Voice: A Massively-Multilingual Speech Corpus
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 263bfe31-5f2a-42e7-a89d-2f5b17f21708 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Semantic parsing on freebase from question-answer pairs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd9c724c-9031-4956-9436-be006d3f294f · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 064bb74c-9156-46c3-9694-4b800f688268 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Chang, Sungbok Lee, and Shrikanth S
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 890de6a2-6f05-43a3-98ae-2553ceb23e21 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Cooper, Michael K
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e46e7d-4f90-4a98-a79c-c4d5b847f5ac · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ddacd8e-8820-478a-b99d-c42fd61a151b · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Qwen2-Audio Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 766b2d98-420a-4132-b942-ebceb2672ddf · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction The fisher corpus: A resource for the next generations of speech-to-text
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c6b218b8-0e78-47f8-8fbe-fc8f75258d44 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901124cb-b118-472a-9d6b-7b1eca27ef75 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Fleurs: Few-shot learning evaluation of universal representations of speech
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19124383-203e-421b-b27e-3f5e71b8eccc · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Moshi: a speech-text foundation model for real-time dialogue
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe772ac-2aaa-4431-bb90-e43c1903b061 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd76858f-2a23-4fa6-9669-48c2d30d7140 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f70485c-6dd9-4382-a73c-ba0ec6bbc1ca · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b296ea0-6cfe-4575-8c7e-66e7714aa4ac · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d1b6baf-4f0e-4b93-b34f-a9f36ca54100 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19999553-e51c-43df-830e-a199bc48e067 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 946008fc-901e-4f96-90a0-471e7e034325 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction GPT-4o System Card
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee46379-d50a-4e51-893c-932f373aa885 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Textrolspeech: A text style control speech corpus with codec language text-to-speech models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9bc52bb4-df75-41c5-8be0-c758430ee32e · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76d2cc58-dd2b-42c4-b757-3eb3b5c08802 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6387e3a-4a47-4d00-a53e-77ceee7a840d · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Prompttts 2: Describing and generating voices with text prompt
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d0193d99-198c-42e4-9d4e-1182085b31cb · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Hashimoto
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7094f46e-75da-491a-b045-e36a3f40d431 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Schuller, and Jianhua Tao
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 889e0ee1-4e31-4af4-af29-f99ef5429cca · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Rouge: A package for automatic evaluation of summaries
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a60179a-80ad-419c-855e-fdad8c3e90c3 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Advancing large language models to capture varied speaking styles and respond properly in spoken conversations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2d5c28ff-e9a7-4c01-8b6b-0bbcb4b6157b · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Paralinguistics-enhanced large language modeling of spoken dialogue
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5ef7d7dd-6c2f-4802-bf93-294c488d2413 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Recording for eyes, not echoing to ears: Contextualized spoken-to-written conversion of ASR transcripts
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7e68416f-9e9a-4439-86ba-dfc0db082199 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cbeae967-33b4-41c2-b95d-50f8d569f215 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Language Model Can Listen While Speaking
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911251f6-75d4-4f90-8400-568096a12af4 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction The MSP-Conversation Corpus
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 84e2f0cb-fa15-4900-8389-cbfd6b2111dc · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction PSLM: parallel generation of text and speech with llms for low-latency spoken dialogue systems
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bff93507-2102-43b9-8186-a55e0555fe3d · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36fdad6d-380d-4d6f-8ff8-bb27e436d29e · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Generative Spoken Dialogue Language Modeling
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a36f4b3f-4457-4e60-b27f-03e504e37a83 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Librispeech: an asr corpus based on public domain audio books
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a72cac6a-5ebe-4d8e-8991-ff6c3cba4fa3 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction BLEU : a method for automatic evaluation of machine translation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d893e1d-033e-4fe3-a22f-a07f82f98528 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction MELD : A multimodal multi-party dataset for emotion recognition in conversations
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7435291b-23f5-46c8-9ac0-793c95dad7fe · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Robust speech recognition via large-scale weak supervision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb7bfb36-54af-4332-9c51-2136b9e466f3 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59465bfa-f433-45f8-aa6b-0d272310c90e · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Seaco-paraformer: A non-autoregressive asr system with flexible and effective hotword customization ability
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f5c666d7-4760-4326-84cd-bad81b94ae6f · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Prompttts++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ca76bd80-41f1-42f3-9797-533d3b1f3607 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction PandaGPT: One Model To Instruction-Follow Them All
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d99fce26-a44b-4516-8955-5b97bac8d132 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Moss: An open conversational large language model
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 25f289c4-ade9-4e54-a546-d8f560828858 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction SALMONN : Towards generic hearing abilities for large language models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d7b22c3e-de65-435a-a1c0-23efd78a4267 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Hashimoto
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d156f465-36ba-4816-97d5-944e6a7f4280 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Qwen2.5: A party of foundation models, September 2024
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a213b5d4-e9de-4abf-a876-3f9f2f6d79e8 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Peloquin, Bokai Yu, Hongyu Gong, and Shyamnath Gollakota
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c93fdb21-b21b-42c5-86c9-2728b3491a27 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Covost 2 and massively multilingual speech translation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e29f0643-f91d-499b-ba34-09985e2a4d06 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction A Full-duplex Speech Dialogue Scheme Based On Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81d59fda-cf2e-4c2e-88ed-5fc67522dd23 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c4f0230-ea80-4bd2-ac0e-75de7eed028d · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Next-gpt: Any-to-any multimodal LLM
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 56bc4ed6-4168-466e-944c-b4cccbe421fb · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac98cf7a-db57-4ed4-a003-6ba6518043f0 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a112542-b1cb-40a5-af00-25387a883c67 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 697b6272-1014-4899-b639-1b9f9f60678a · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c5cc91c-664b-4965-8b11-72aae04067f9 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction M2met: The icassp 2022 multi-channel multi-party meeting transcription challenge
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 82f5b4c4-359b-4b56-a715-6311a995534e · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f7c16d6-8c46-40dc-89dd-7c30043b047a · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7344f4fc-08be-4d16-9191-f9f9449a7afc · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Design of speech corpus for mandarin text to speech
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4884d303-fbbd-4329-8b8d-13bd2356acca · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea82a65-33ba-4ab6-9c0e-28d651428f87 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 412d4691-bfa3-4730-a4cd-b0560506266f · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Xing, Hao Zhang, Joseph E
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 00d21988-b6ed-4f4e-9c09-0e49937b5209 · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset, 2021
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bf0ef60b-364e-4c1c-b2b2-d313a8f6cf1e · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction @esa (Ref
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3486519a-7d85-4a2e-830d-771eefe1f76a · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Unresolved cited work
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cae7078-9d07-49f5-b004-9cd8c21c982b · outbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction WavChat: A Survey of Spoken Dialogue Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19db4261-2075-4f44-8b65-a0faa0eeb656 · inbound
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486a5b6a-ea91-4e3c-94d9-7d547a21d731 · inbound
LUCY: Linguistic Understanding and Control Yielding Early Stage of Her MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8366d02-27ad-4e39-8c1c-979a30e28278 · inbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cf9fb5dd-37db-4f48-bcff-41420203c652 · inbound
Qwen2.5-Omni Technical Report MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 289e2685-f9ab-426c-96da-39c594ffe397 · inbound
Kimi-Audio Technical Report MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2f5e0890-292c-4033-8655-acc12cb9abaf · inbound
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8de665c-47f6-46ec-bace-b58b78ad2f12 · inbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b70932-5755-41b0-90f4-6ac9cbb1fd62 · inbound
SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5802628b-d704-4ac7-be71-bddc3a1e7dcb · inbound
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d2aea3f9-24c4-499c-ab07-bcca87572d2a · inbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b35f5ec1-0f86-42ca-a4e5-1425f62ae257 · inbound
RoboEgo System Card: An Omnimodal Model with Native Full Duplexity MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c79f311-4e44-4233-a9bf-7396fc98ea25 · inbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a47b8e75-4b61-4b57-b1ed-d6e17c8c917d · inbound
SHNU Multilingual Conversational Speech Recognition System for INTERSPEECH 2025 MLC-SLM Challenge MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9690060e-f981-4df2-969d-edbf34cee12c · inbound
Differentiable Reward Optimization for LLM based TTS system MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6665ef24-3b40-480a-bf0d-30f4124a50d6 · inbound
Step-Audio 2 Technical Report MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4fe26263-f164-4659-a8fe-f347056a6859 · inbound
FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 835e4410-3f65-4d47-bf3a-bba964b3caf9 · inbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5052c7e-0f0d-4940-b654-d55c810e0ddf · inbound
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3111c3be-fa8c-4298-9f55-830136d93b89 · inbound
Game-Time: Evaluating Temporal Dynamics in Spoken Language Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d153856f-7e16-4451-80dd-62b539f890d7 · inbound
Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3084213f-457d-4427-a41e-431578c03d96 · inbound
A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f87ebdfc-9267-41f8-b353-9a15b04da1a1 · inbound
Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4bd9864a-415f-4561-9c89-e1e4e80f81fd · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6e833eaf-2b6b-4a9f-b76d-736e5db9cb9e · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f98c64bc-ba00-4f4c-8738-a43e47c1706f · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc6f7e98-2b30-46dc-9b13-9ca9073d695e · inbound
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation de3cdd73-cdc7-4589-bf53-8de655a8d4ad · inbound
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 261905a0-a092-4246-8f4c-c7f6c1ffe5ae · inbound
IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 45174887-b8b0-46f6-8832-7439fbf48ad6 · inbound
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8c92bdd4-85d7-4681-b4cc-64134ed8f291 · inbound
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d85abf04-d89f-4a3a-bf8d-9ff602672990 · inbound
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1c775997-0ae1-4ebb-bcca-133882e01dff · inbound
A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0083f0bf-be5f-402f-a7a8-fddd8bc5ad46 · inbound
COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1cf1750e-18f3-48bc-b49e-15a5843b5a95 · inbound
COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 32763740-303e-4d41-870f-d964311c304e · inbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation be3d8eac-f471-4e53-86b3-0a7ecd8971f1 · inbound
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b295e9-c9a7-4c43-bf82-ef8fd55297a9 · inbound
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.