Pith. sign in

Paper Citation Record · LEDGER

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

As of 23 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 37 inbound Pith citation observations for arXiv:2501.06282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06282 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:56.898917Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:52:07.794411Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4d43ee89-7e06-48a8-bb4b-0180550ffa80 · outbound

This paper cites write newline.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.598965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.598965Z digest=sha256:056609e7067fcf352f1e794d7ae0dae7a4402c2d91daccd8db7a4b29224b3228

Observation dc597275-f6ca-4605-b446-e077a9119a7a · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.605425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.605425Z digest=sha256:39eae360a64d588a15c800218ac3862e2cb2e18388b3162397c3fcccaaa29ec5

Observation fa3bcb17-216f-4dd9-9684-62e5fe3e175f · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.610163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.610163Z digest=sha256:beead82e985d126eda87b5aad916a1b9a6b03008a606f63611825152eb98b689

Observation c7a839c8-45f7-4aad-af85-79b48f825d85 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Common Voice: A Massively-Multilingual Speech Corpus

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.614639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.614639Z digest=sha256:cdc90ed09604466dc593c2318f62fb9bfb9beb7868b7fb1102eaf7d404d891bd

Observation 263bfe31-5f2a-42e7-a89d-2f5b17f21708 · outbound

This paper cites Semantic parsing on freebase from question-answer pairs.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Semantic parsing on freebase from question-answer pairs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.619567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.619567Z digest=sha256:94aedf050bbbd1b8e222fbba54486c01291c4c49160a4bd7adc8e380aa0e8e53

Observation fd9c724c-9031-4956-9436-be006d3f294f · outbound

This paper cites an unresolved cited work.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:10:57.843240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.623839Z digest=sha256:cdbb608b2b236f7a014a2674a79e318659f4e0ab99db577de297651c37273aba

Observation 064bb74c-9156-46c3-9694-4b800f688268 · outbound

This paper cites Chang, Sungbok Lee, and Shrikanth S.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Chang, Sungbok Lee, and Shrikanth S

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.829100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.627935Z digest=sha256:50178d3be91124363fb78080eba49d93e5c4892e0a11842f48cb068368e15999

Observation 890de6a2-6f05-43a3-98ae-2553ceb23e21 · outbound

This paper cites Cooper, Michael K.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Cooper, Michael K

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.632218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.632218Z digest=sha256:c624e7105794b517aa9cdaff6007bf4c957b481c29e201e01f5fa44b10b19a5d

Observation 40e46e7d-4f90-4a98-a79c-c4d5b847f5ac · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.636363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.636363Z digest=sha256:97c43b19f210befa240153fac8b5cc09b28f724e80acd2b28d855cca917461e3

Observation 0ddacd8e-8820-478a-b99d-c42fd61a151b · outbound

This paper cites Qwen2-Audio Technical Report.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Qwen2-Audio Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.640948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.640948Z digest=sha256:b6965fce94216095b0a893281236e4512be703bf0953c32d21c1d6854c968a7a

Observation 766b2d98-420a-4132-b942-ebceb2672ddf · outbound

This paper cites The fisher corpus: A resource for the next generations of speech-to-text.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction The fisher corpus: A resource for the next generations of speech-to-text

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.816239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.645866Z digest=sha256:88257b39206c4d1f4bde4e3535f66f6f2b35ac89025bd4c027218f27e4c30c21

Observation c6b218b8-0e78-47f8-8fbe-fc8f75258d44 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.649859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.649859Z digest=sha256:ef6eab8769408f4248933408ef0d34f9acb6aaf3d9750a21b14c212ab7b54be8

Observation 901124cb-b118-472a-9d6b-7b1eca27ef75 · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Fleurs: Few-shot learning evaluation of universal representations of speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.654222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.654222Z digest=sha256:43bf9c9a38be3500716c71882f4c4248eb9de1927b5c9fafc4903aa050b1ae32

Observation 19124383-203e-421b-b27e-3f5e71b8eccc · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Moshi: a speech-text foundation model for real-time dialogue

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.657967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.657967Z digest=sha256:958a180a3d20307e87722c60aefe0d190f98d88f92aaaff634c130818719a0fa

Observation 5fe772ac-2aaa-4431-bb90-e43c1903b061 · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.662418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.662418Z digest=sha256:02f78b82f15aa4c79301db81baad1a18ea7d5119bb0e3d7d9458f2e8595047c3

Observation fd76858f-2a23-4fa6-9669-48c2d30d7140 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.665976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.665976Z digest=sha256:1b587c696904bd4da22ad8851f90dfc064f4fc8fd1cccc833a50819eb20f050a

Observation 9f70485c-6dd9-4382-a73c-ba0ec6bbc1ca · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.670340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.670340Z digest=sha256:14e1d9913dfb29f4223a1a0a6b14dc7a8b1e1a8e164acd61c760c1e70ec29ecf

Observation 5b296ea0-6cfe-4575-8c7e-66e7714aa4ac · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.674083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.674083Z digest=sha256:66dbba78e59aca13ef9a534285e271763f920a319a004e82a4f873e28d86aedf

Observation 4d1b6baf-4f0e-4b93-b34f-a9f36ca54100 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.680203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.680203Z digest=sha256:c30e38a86ea6fd3fe3aeb2619e3052c8169e82ac4b28a2809c1ed9c956162327

Observation 19999553-e51c-43df-830e-a199bc48e067 · outbound

This paper cites Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.794215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.684862Z digest=sha256:b09bb11a8b82764177bb338604c67a0d40f753a943a641f23ed417a27603b574

Observation 946008fc-901e-4f96-90a0-471e7e034325 · outbound

This paper cites GPT-4o System Card.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.689629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.689629Z digest=sha256:b75f417b4d0591379ce500f64c6a4f5c1c490e7b1ed33bd44c4a9c5a2059d4ee

Observation fee46379-d50a-4e51-893c-932f373aa885 · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Textrolspeech: A text style control speech corpus with codec language text-to-speech models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.779903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.697906Z digest=sha256:e1784294e8859ac993561de420c4b5cb8abed13ed24b63c65c2714bbdf3d0c31

Observation 9bc52bb4-df75-41c5-8be0-c758430ee32e · outbound

This paper cites ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.701885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.701885Z digest=sha256:e1c780c8af9ca06ece3b46fc492ecc0583f17c8241594b082cbe48bdb3693d59

Observation 76d2cc58-dd2b-42c4-b757-3eb3b5c08802 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.706039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.706039Z digest=sha256:fa4625a4454ad38094e27d35aaa5c4d27af7beb07f0aaaf0833d8ce770c29495

Observation e6387e3a-4a47-4d00-a53e-77ceee7a840d · outbound

This paper cites Prompttts 2: Describing and generating voices with text prompt.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Prompttts 2: Describing and generating voices with text prompt

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.768103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.710656Z digest=sha256:07115e42e261737c1d98cc1ae7700ce3ffb53c96ac610254835afa2798fb7e12

Observation d0193d99-198c-42e4-9d4e-1182085b31cb · outbound

This paper cites Hashimoto.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Hashimoto

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.714729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.714729Z digest=sha256:3f91e6fa95deaf67695f706cb40d78ab0c2bcccae24290b35bc0753267cb926b

Observation 7094f46e-75da-491a-b045-e36a3f40d431 · outbound

This paper cites Schuller, and Jianhua Tao.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Schuller, and Jianhua Tao

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.746680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.719166Z digest=sha256:c26df156532cb35ee2d750dc286f515bcc2d10a458307a041f5de1cd91d178db

Observation 889e0ee1-4e31-4af4-af29-f99ef5429cca · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Rouge: A package for automatic evaluation of summaries

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.723507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.723507Z digest=sha256:26a0157baffcb9ea960f09299c576d3534ccc23bd64abb8d42141eafcdbf3f6b

Observation 8a60179a-80ad-419c-855e-fdad8c3e90c3 · outbound

This paper cites Advancing large language models to capture varied speaking styles and respond properly in spoken conversations.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Advancing large language models to capture varied speaking styles and respond properly in spoken conversations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.726850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.727738Z digest=sha256:3dfe9234814174e6c1a1ca9ab5f3b3c9bca6b3e4cbcee7e497e51438318f6761

Observation 2d5c28ff-e9a7-4c01-8b6b-0bbcb4b6157b · outbound

This paper cites Paralinguistics-enhanced large language modeling of spoken dialogue.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Paralinguistics-enhanced large language modeling of spoken dialogue

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.712989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.731863Z digest=sha256:ff9cefedcee0c4c76ca415799db407ea73a8c28adea46ab4fef496970f828876

Observation 5ef7d7dd-6c2f-4802-bf93-294c488d2413 · outbound

This paper cites Recording for eyes, not echoing to ears: Contextualized spoken-to-written conversion of ASR transcripts.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Recording for eyes, not echoing to ears: Contextualized spoken-to-written conversion of ASR transcripts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.700307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.736530Z digest=sha256:4772167300e99dbcd562efa7d7a74548b3bf7addb8f24597cee64b722ee315f6

Observation 7e68416f-9e9a-4439-86ba-dfc0db082199 · outbound

This paper cites Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.687714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.740390Z digest=sha256:b286cb208987cb0ff560bac742eb9b6f0cc2178aaeb1596a4eaeb96a796a1665

Observation cbeae967-33b4-41c2-b95d-50f8d569f215 · outbound

This paper cites Language Model Can Listen While Speaking.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Language Model Can Listen While Speaking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.744635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.744635Z digest=sha256:4d07b6a6d1c7b08b3ebbff0e5ce66af329e8514af52c98f7363b8c215036f777

Observation 911251f6-75d4-4f90-8400-568096a12af4 · outbound

This paper cites The MSP-Conversation Corpus.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction The MSP-Conversation Corpus

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.675342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.749203Z digest=sha256:7265388b0465a859bd88c71ea059576944363b512dd7b366969dd2e767090ce0

Observation 84e2f0cb-fa15-4900-8389-cbfd6b2111dc · outbound

This paper cites PSLM: parallel generation of text and speech with llms for low-latency spoken dialogue systems.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction PSLM: parallel generation of text and speech with llms for low-latency spoken dialogue systems

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.663335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.753158Z digest=sha256:187980c21c5711d1c0ad361b85a82f3e33d6c9f79bf345400e91a41d6b2646ed

Observation bff93507-2102-43b9-8186-a55e0555fe3d · outbound

This paper cites Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.757453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.757453Z digest=sha256:4c61ddc6fa69114340023269a80e7a3ec2d5b1f9d2f90b7a1d114186feb88d66

Observation 36fdad6d-380d-4d6f-8ff8-bb27e436d29e · outbound

This paper cites Generative Spoken Dialogue Language Modeling.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Generative Spoken Dialogue Language Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.762195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.762195Z digest=sha256:82cc4e51ed847c11e7ea43cc50c2076087a295146428e09d6fa9e97d8ba89f97

Observation a36f4b3f-4457-4e60-b27f-03e504e37a83 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Librispeech: an asr corpus based on public domain audio books

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.766632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.766632Z digest=sha256:747a1de8c269271d32a90628414c2567590228f8615c25a2856b46937da60a73

Observation a72cac6a-5ebe-4d8e-8991-ff6c3cba4fa3 · outbound

This paper cites BLEU : a method for automatic evaluation of machine translation.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction BLEU : a method for automatic evaluation of machine translation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.770544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.770544Z digest=sha256:751632c676c86d1f8ed38c151fe9be686d75fb1fd0d8fd9ac959efd4a0f69eca

Observation 2d893e1d-033e-4fe3-a22f-a07f82f98528 · outbound

This paper cites MELD : A multimodal multi-party dataset for emotion recognition in conversations.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction MELD : A multimodal multi-party dataset for emotion recognition in conversations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.640787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.774178Z digest=sha256:8a9e7b33c8cb24ea7562445312b0a4906ca9d00a81f87988b6fb0940323e0552

Observation 7435291b-23f5-46c8-9ac0-793c95dad7fe · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Robust speech recognition via large-scale weak supervision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.777992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.777992Z digest=sha256:96e0d5af7ea20526676bceb6ecd033c3680c77cde08e140da19444c1fd5ee7be

Observation fb7bfb36-54af-4332-9c51-2136b9e466f3 · outbound

This paper cites an unresolved cited work.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.781835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.781835Z digest=sha256:60c65f9bd6dbf5501f751134ee96b95551e3722aae42dfc78873cc2ccc1be72e

Observation 59465bfa-f433-45f8-aa6b-0d272310c90e · outbound

This paper cites Seaco-paraformer: A non-autoregressive asr system with flexible and effective hotword customization ability.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Seaco-paraformer: A non-autoregressive asr system with flexible and effective hotword customization ability

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.620815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.786165Z digest=sha256:e12f0c7a217db98c106039ff73d5fd5e8357202e0b290e03c45d776d64dcfe82

Observation f5c666d7-4760-4326-84cd-bad81b94ae6f · outbound

This paper cites Prompttts++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Prompttts++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.609235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.791034Z digest=sha256:c5c1c2d1a3a2b071fde150c6a911b419ae50d83e2076fccec8b037cbce1bdc06

Observation ca76bd80-41f1-42f3-9797-533d3b1f3607 · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction PandaGPT: One Model To Instruction-Follow Them All

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.795509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.795509Z digest=sha256:4810e2d77d3d003b883e961623865b5948436f07bdb84668243e4646e92fea4a

Observation d99fce26-a44b-4516-8955-5b97bac8d132 · outbound

This paper cites Moss: An open conversational large language model.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Moss: An open conversational large language model

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.597838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.799574Z digest=sha256:f60008ca41eabef5d83d624ff1b08a1e269419b45305a98baaa06e47365c1beb

Observation 25f289c4-ade9-4e54-a546-d8f560828858 · outbound

This paper cites SALMONN : Towards generic hearing abilities for large language models.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction SALMONN : Towards generic hearing abilities for large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.584495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.803552Z digest=sha256:116216554e6b015b5bd85f770e1b3baf98be49ea9b17e893b4e10d6f5b06121e

Observation d7b22c3e-de65-435a-a1c0-23efd78a4267 · outbound

This paper cites Hashimoto.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Hashimoto

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.807716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.807716Z digest=sha256:31ed97b6548377ec1c9686fe04bcf7e27c9bd5d2d3b8c5fc7be25d37084648cd

Observation d156f465-36ba-4816-97d5-944e6a7f4280 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Qwen2.5: A party of foundation models, September 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.811718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.811718Z digest=sha256:3f6877ed21e661ec249179420081b2bb04a827170b92ce8c6caac28b476bfc67

Observation a213b5d4-e9de-4abf-a876-3f9f2f6d79e8 · outbound

This paper cites Peloquin, Bokai Yu, Hongyu Gong, and Shyamnath Gollakota.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Peloquin, Bokai Yu, Hongyu Gong, and Shyamnath Gollakota

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.556413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.816002Z digest=sha256:d8494350037b382e52f72665f99fcc97476a60bbbb1b8718c2389b6f512de718

Observation c93fdb21-b21b-42c5-86c9-2728b3491a27 · outbound

This paper cites Covost 2 and massively multilingual speech translation.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Covost 2 and massively multilingual speech translation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.543472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.820362Z digest=sha256:db10bf08608b0b24a713d54b02c081be8fe708e03263560f2e0b59981d721e6e

Observation e29f0643-f91d-499b-ba34-09985e2a4d06 · outbound

This paper cites A Full-duplex Speech Dialogue Scheme Based On Large Language Models.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction A Full-duplex Speech Dialogue Scheme Based On Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.824128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.824128Z digest=sha256:403f1f78de606167957af55c624ac8c2ab59ad942c4c1f249a41d8aecf785d86

Observation 81d59fda-cf2e-4c2e-88ed-5fc67522dd23 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.828694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.828694Z digest=sha256:d5c0ecb5a108021f97c9599dccc2429cfa52c04dfe53ef72dd30069b6607fc93

Observation 3c4f0230-ea80-4bd2-ac0e-75de7eed028d · outbound

This paper cites Next-gpt: Any-to-any multimodal LLM.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Next-gpt: Any-to-any multimodal LLM

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.530172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.833801Z digest=sha256:9fc47aabc30e3d7ae5ae67b1eca9b335f568531d27031c2e385e9f49f24cf2bc

Observation 56bc4ed6-4168-466e-944c-b4cccbe421fb · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.838954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.838954Z digest=sha256:41797f4268177950b070b09089dbb869a0bf92a14f054d3603ca63e99009f7e7

Observation ac98cf7a-db57-4ed4-a003-6ba6518043f0 · outbound

This paper cites E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.843007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.843007Z digest=sha256:b42927e26f7150bd91986ccccb9b28553bd2f5fd48bcd8955c5bd69703ddc50a

Observation 0a112542-b1cb-40a5-af00-25387a883c67 · outbound

This paper cites Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.516275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.848030Z digest=sha256:fed2a5259bb52d9f2745e273e62b0505427aa07b079a7f093fa0628d30c3482d

Observation 697b6272-1014-4899-b639-1b9f9f60678a · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.851945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.851945Z digest=sha256:c491733560e46a46859faf9e1a27a6db551774a3f322469a877fedb3f15b95b7

Observation 1c5cc91c-664b-4965-8b11-72aae04067f9 · outbound

This paper cites M2met: The icassp 2022 multi-channel multi-party meeting transcription challenge.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction M2met: The icassp 2022 multi-channel multi-party meeting transcription challenge

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.503780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.856089Z digest=sha256:25fa917a397b8dd41345d1553bbc1242ecd99ad00ae14aa6d5455a01fbc2d36f

Observation 82f5b4c4-359b-4b56-a715-6311a995534e · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.860538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.860538Z digest=sha256:6713fedab086deef2727814f9070f7df8e34aa9c19f81847f2a6672a1bd23a91

Observation 2f7c16d6-8c46-40dc-89dd-7c30043b047a · outbound

This paper cites Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.864877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.864877Z digest=sha256:47941e4a9673b60e082775db0f297ca19d7bd677c5564131aa56b5ccebe6436e

Observation 7344f4fc-08be-4d16-9191-f9f9449a7afc · outbound

This paper cites Design of speech corpus for mandarin text to speech.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Design of speech corpus for mandarin text to speech

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.868618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.868618Z digest=sha256:19de5a0d50ce33ebcc98ff41b2a7b6310293423acfea96f61e53e77e777f596a

Observation 4884d303-fbbd-4329-8b8d-13bd2356acca · outbound

This paper cites OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.872982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.872982Z digest=sha256:1f2876d2248d08c1b7e29b6000e0cf2d2c6b7a816f772bfad8fc83dc0db27403

Observation 2ea82a65-33ba-4ab6-9c0e-28d651428f87 · outbound

This paper cites IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.876710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.876710Z digest=sha256:b058ba9d82aa928afc7ee661d956687de4ef1a0f597209e8fb4906dfdff47c45

Observation 412d4691-bfa3-4730-a4cd-b0560506266f · outbound

This paper cites Xing, Hao Zhang, Joseph E.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Xing, Hao Zhang, Joseph E

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.476501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.881469Z digest=sha256:913e138b3df4013c2037491f74f1b63aa8db54a8014f105fbc0d7ed0e8dd9bed

Observation 00d21988-b6ed-4f4e-9c09-0e49937b5209 · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset, 2021.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset, 2021

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:57.464168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T21:10:56.885277Z digest=sha256:f202c8014d0e2c254c974570b893140bda4b9d271d759efd93010c7fad4ec81e

Observation bf0ef60b-364e-4c1c-b2b2-d313a8f6cf1e · outbound

This paper cites @esa (Ref.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction @esa (Ref

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.889278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.889278Z digest=sha256:7e30e9eff7aa6517810dd4723e93ec894fc95bf2f1805de1f734176987fdd892

Observation 3486519a-7d85-4a2e-830d-771eefe1f76a · outbound

This paper cites an unresolved cited work.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.894668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.894668Z digest=sha256:a12d84d1e8269266e4df0d0fb0940dd510b236f8fb59990f2f1cab4a8638a699

Observation 8cae7078-9d07-49f5-b004-9cd8c21c982b · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction WavChat: A Survey of Spoken Dialogue Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.898917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.898917Z digest=sha256:8752c638dce2f590df4387dde9679b0727208a17711f78fa305db7c76c69a239

Pith citing papers

Observation 19db4261-2075-4f44-8b65-a0faa0eeb656 · inbound

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models cites this paper.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.131575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.131575Z digest=sha256:e3057b282c4cc67ce66d97a99bc14d7d3405f982d3682f8104f7fb547d9576e0

Observation 486a5b6a-ea91-4e3c-94d9-7d547a21d731 · inbound

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her cites this paper.

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:25.637672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:25.637672Z digest=sha256:bb59434e768a93a1c9333ad42edfdc3a96c3a14e09a0cd33bbeb483bb9fba2fc

Observation b8366d02-27ad-4e39-8c1c-979a30e28278 · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.414015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:5faca8f5ab705d8a3bd84554f5a94acd3a0a995af2ce4a03b2db33ca1ffffcfc

Observation cf9fb5dd-37db-4f48-bcff-41420203c652 · inbound

Qwen2.5-Omni Technical Report cites this paper.

Qwen2.5-Omni Technical Report MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:54:03.278501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:54:03.225439Z digest=sha256:495aa47c677d152af82d5d54deec8d05ee90a5ac1ee16381ec29d33b290a1cdc

Observation 289e2685-f9ab-426c-96da-39c594ffe397 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:21:27.216305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:5f098a39c61712220618ff9a7379d971791fff30c3fe7d1aa510267dfd15f175

Observation 2f5e0890-292c-4033-8655-acc12cb9abaf · inbound

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis cites this paper.

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:07.794411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:52:07.794411Z digest=sha256:9031cbda4cfed8b806f0d2e68fe1bc6b36d8ef23df10fa694895673a78d05a1b

Observation f8de665c-47f6-46ec-bace-b58b78ad2f12 · inbound

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model cites this paper.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.249905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.249905Z digest=sha256:0012cbb51b90c04331734f24f2a01eecf4d9037a1aadc4e6ce8e6c82e62dd1e5

Observation 34b70932-5755-41b0-90f4-6ac9cbb1fd62 · inbound

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation cites this paper.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.175649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.175649Z digest=sha256:92e0355dc5921b5bec7f0401f549d7acfb0cebc662b66eac2ffba7ad58b32462

Observation 5802628b-d704-4ac7-be71-bddc3a1e7dcb · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:27:25.575517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:ecbbe603ebdbbabafd95470a9d50ff7c258774a7b71fa8e6abb0d52ed28eb1c1

Observation d2aea3f9-24c4-499c-ab07-bcca87572d2a · inbound

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation cites this paper.

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:25.262496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:25.262496Z digest=sha256:911031af76675f566975d25592e259a6c1a7a68aa760187e9a02f3125ad436cb

Observation b35f5ec1-0f86-42ca-a4e5-1425f62ae257 · inbound

RoboEgo System Card: An Omnimodal Model with Native Full Duplexity cites this paper.

RoboEgo System Card: An Omnimodal Model with Native Full Duplexity MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:37:29.400192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:37:29.400192Z digest=sha256:7246d8c8b8847081ff1959feaecf5c753c0752b4040f69f40ab63f6b72189fcd

Observation 0c79f311-4e44-4233-a9bf-7396fc98ea25 · inbound

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training cites this paper.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:14.172217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:14.172217Z digest=sha256:dd7c8455725a1b0d10ab52ab11e0ae20b46742c884763a629db462193ee26faf

Observation a47b8e75-4b61-4b57-b1ed-d6e17c8c917d · inbound

SHNU Multilingual Conversational Speech Recognition System for INTERSPEECH 2025 MLC-SLM Challenge cites this paper.

SHNU Multilingual Conversational Speech Recognition System for INTERSPEECH 2025 MLC-SLM Challenge MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:17:22.225974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:17:22.225974Z digest=sha256:233746234b485b82b6ee8029268ee746287b7ce3144bd82c93f49fc383a7de70

Observation 9690060e-f981-4df2-969d-edbf34cee12c · inbound

Differentiable Reward Optimization for LLM based TTS system cites this paper.

Differentiable Reward Optimization for LLM based TTS system MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:03.998457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:03.998457Z digest=sha256:1b8eb3b568800596a6c1c06f3802757fca580cd9adfcc50bf0d54692e1b50944

Observation 6665ef24-3b40-480a-bf0d-30f4124a50d6 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:59:50.972489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:626c7a80e34a3bd3d2636af3fc490d60f3cf1378e8fd364f4a8532c5259d7180

Observation 4fe26263-f164-4659-a8fe-f347056a6859 · inbound

FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems cites this paper.

FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:06:21.259264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:06:21.259264Z digest=sha256:e3b875c9a63b6cc2477fc9802990908e3dd29316c6d1df1550f73bb738446805

Observation 835e4410-3f65-4d47-bf3a-bba964b3caf9 · inbound

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation cites this paper.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:57.014492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:57.014492Z digest=sha256:e19d7a93b07cac9ec235f742aa39bc5b0ef3c59fa0093b9358f4add339eebdd2

Observation f5052c7e-0f0d-4940-b654-d55c810e0ddf · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.393443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.393443Z digest=sha256:42400daccac54a1aca2f196235bf0e1fca03ac6e08a50b8da0d75e8d4f3cc8f5

Observation 3111c3be-fa8c-4298-9f55-830136d93b89 · inbound

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models cites this paper.

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:52:35.659434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T11:51:43.561210Z digest=sha256:68918cfc0132087eec6fb7c3411249fafa0da7fbe55460185ed9009f75f72e66

Observation d153856f-7e16-4451-80dd-62b539f890d7 · inbound

Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models cites this paper.

Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:49.151182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:49.151182Z digest=sha256:5b8c0cb8b68a3b54a12da40b4b3fb56882d69e1ca7a846e9b17711ea511cae08

Observation 3084213f-457d-4427-a41e-431578c03d96 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:f80b669557710889e4667b8670fdc3dba412bb93c49184f4b70ef604a81f6ec1

Observation f87ebdfc-9267-41f8-b353-9a15b04da1a1 · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:35:18.933536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:02219b199f8a9079454b6d7497748d39235ffd5ece1ab590975138f81aa66ede

Observation 4bd9864a-415f-4561-9c89-e1e4e80f81fd · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:46:15.309003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:1f2711f055e4f9713f0056cce422d1f3b57d8b09a7291b0bc631400a3dc7b96e

Observation 6e833eaf-2b6b-4a9f-b76d-736e5db9cb9e · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:50:49.658682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:933f02698ed4f504a0233691f8dd5e727494b42fdd10bad7da811c3c6a7468d0

Observation f98c64bc-ba00-4f4c-8738-a43e47c1706f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.814043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:a2f857b99476608516172e19d8115767a56e91cb7224b20ab8022ba3095ee4be

Observation bc6f7e98-2b30-46dc-9b13-9ca9073d695e · inbound

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue cites this paper.

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:11:19.082803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T03:08:55.753359Z digest=sha256:5dc9bc3c9faac0dc8d4c363961fac86384c1dbe3d5fd0fac1aa989dbc361107b

Observation de3cdd73-cdc7-4589-bf53-8de655a8d4ad · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:52:27.005157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:e43e4751535e3f7245408ac2b649e7f582722ded88de43d9b0d7b8acbac85c49

Observation 261905a0-a092-4246-8f4c-c7f6c1ffe5ae · inbound

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems cites this paper.

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:37:06.845645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T23:42:38.203116Z digest=sha256:2bcf175ce2785e0f85119967cc0902088652d12520a8b606df47be6321a0a32f

Observation 45174887-b8b0-46f6-8832-7439fbf48ad6 · inbound

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models cites this paper.

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:20:57.401988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T12:53:00.569655Z digest=sha256:c403a6ca0cc793d5fc1e30982cc2d017d5b90b9a2c488f6f7643a3e88373de50

Observation 8c92bdd4-85d7-4681-b4cc-64134ed8f291 · inbound

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents cites this paper.

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:28:39.528756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T05:32:27.371597Z digest=sha256:d8bde097a98bdc2e81b8dbff6ef440b755de6c016e6d8b01de029daad0b72efa

Observation d85abf04-d89f-4a3a-bf8d-9ff602672990 · inbound

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents cites this paper.

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:43:51.440118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T05:01:44.202712Z digest=sha256:f3d32776dbc57fc36e5cfc9fe64555eb68ff6eb4fdaf118be1cca0f79d2f3337

Observation 1c775997-0ae1-4ebb-bcca-133882e01dff · inbound

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine cites this paper.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:49:24.994509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:8f1e7ec4d93beb92f95ecf0460bd37a0183c3e9fd1537d9c390bffc119c58303

Observation 0083f0bf-be5f-402f-a7a8-fddd8bc5ad46 · inbound

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices cites this paper.

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T03:14:14.952294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T03:09:57.069086Z digest=sha256:2d24c1856d3a1208751eb0c69c55cedba277c18c7856682577201d37b6b0d03e

Observation 1cf1750e-18f3-48bc-b49e-15a5843b5a95 · inbound

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices cites this paper.

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:35:39.871015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T06:28:07.670564Z digest=sha256:92d776d4ce30316a2ff59203c14df35d25be211bb7f9e56716531c1ae35d131a

Observation 32763740-303e-4d41-870f-d964311c304e · inbound

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation cites this paper.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:05:45.810988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:05763f416b8b3da1671bcb2b811650c8e08ae766731f4ccf6c5250dd27de156f

Observation be3d8eac-f471-4e53-86b3-0a7ecd8971f1 · inbound

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm cites this paper.

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T23:35:22.292690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:35:22.292690Z digest=sha256:8523d0d89235174209ac4b44ccfc7c228f356619b1d3176a9e32bb027873dc92

Observation 25b295e9-c9a7-4c43-bf82-ef8fd55297a9 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:47.364591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:47.364591Z digest=sha256:3db1f150fd71ff4ed2709617e25f2305dce23e3e37917c8a95d221937a67ed98