Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2501.06282.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:41.249905Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation b8366d02-27ad-4e39-8c1c-979a30e28278 · inbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf9fb5dd-37db-4f48-bcff-41420203c652 · inbound
Qwen2.5-Omni Technical Report MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 289e2685-f9ab-426c-96da-39c594ffe397 · inbound
Kimi-Audio Technical Report MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8de665c-47f6-46ec-bace-b58b78ad2f12 · inbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5802628b-d704-4ac7-be71-bddc3a1e7dcb · inbound
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d2aea3f9-24c4-499c-ab07-bcca87572d2a · inbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b35f5ec1-0f86-42ca-a4e5-1425f62ae257 · inbound
RoboEgo System Card: An Omnimodal Model with Native Full Duplexity MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c79f311-4e44-4233-a9bf-7396fc98ea25 · inbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a47b8e75-4b61-4b57-b1ed-d6e17c8c917d · inbound
SHNU Multilingual Conversational Speech Recognition System for INTERSPEECH 2025 MLC-SLM Challenge MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9690060e-f981-4df2-969d-edbf34cee12c · inbound
Differentiable Reward Optimization for LLM based TTS system MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6665ef24-3b40-480a-bf0d-30f4124a50d6 · inbound
Step-Audio 2 Technical Report MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 835e4410-3f65-4d47-bf3a-bba964b3caf9 · inbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5052c7e-0f0d-4940-b654-d55c810e0ddf · inbound
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3111c3be-fa8c-4298-9f55-830136d93b89 · inbound
Game-Time: Evaluating Temporal Dynamics in Spoken Language Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d153856f-7e16-4451-80dd-62b539f890d7 · inbound
Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3084213f-457d-4427-a41e-431578c03d96 · inbound
A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f87ebdfc-9267-41f8-b353-9a15b04da1a1 · inbound
Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4bd9864a-415f-4561-9c89-e1e4e80f81fd · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6e833eaf-2b6b-4a9f-b76d-736e5db9cb9e · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f98c64bc-ba00-4f4c-8738-a43e47c1706f · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc6f7e98-2b30-46dc-9b13-9ca9073d695e · inbound
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation de3cdd73-cdc7-4589-bf53-8de655a8d4ad · inbound
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 261905a0-a092-4246-8f4c-c7f6c1ffe5ae · inbound
IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 45174887-b8b0-46f6-8832-7439fbf48ad6 · inbound
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c92bdd4-85d7-4681-b4cc-64134ed8f291 · inbound
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d85abf04-d89f-4a3a-bf8d-9ff602672990 · inbound
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1c775997-0ae1-4ebb-bcca-133882e01dff · inbound
A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0083f0bf-be5f-402f-a7a8-fddd8bc5ad46 · inbound
COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1cf1750e-18f3-48bc-b49e-15a5843b5a95 · inbound
COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32763740-303e-4d41-870f-d964311c304e · inbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be3d8eac-f471-4e53-86b3-0a7ecd8971f1 · inbound
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b295e9-c9a7-4c43-bf82-ef8fd55297a9 · inbound
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.