{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:X2MB53XENK52WZGWY6O4W7TYSM","short_pith_number":"pith:X2MB53XE","schema_version":"1.0","canonical_sha256":"be981eeee46abbab64d6c79dcb7e78930880a4974c6b2a377d95fb3bb48be160","source":{"kind":"arxiv","id":"2501.06282","version":1},"attestation_state":"computed","paper":{"title":"MinMo: A Multimodal Large Language Model for Seamless Voice Interaction","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.HC","cs.SD","eess.AS"],"primary_cat":"cs.CL","authors_text":"Baosong Yang, Bin Ma, Changfeng Gao, Chong Deng, Chongjia Ni, Chong Zhang, Fan Yu, Guanrou Yang, Haoneng Luo, Hao Wang, Hui Wang, Jialong Tang, Jiaqing Liu, Jinren Zhou, Mengzhe Chen, Nan Zhao, Pei Zhang, Qian Chen, Qinglin Zhang, Ruize Gao, Shiliang Zhang, Tianyu Zhao, Wen Wang, Xiang Lv, Xian Shi, Xian Yang, Yabin Li, Yafeng Chen, Yanni Chen, Yexin Yang, Yingda Chen, Yunlan Xu, Yuxuan Wang, Zhifu Gao, Zhihao Du, Zhijie Yan","submitted_at":"2025-01-10T15:55:27Z","abstract_excerpt":"Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and human-like conversations. Previous models for voice interactions are categorized as native and aligned. Native models integrate speech and text processing in one framework but struggle with issues like differing sequence lengths and insufficient pre-training. Aligned models maintain text LLM capabilities but are often limited by small datasets and a narrow focus on speech tasks. In this work, we introduce MinMo, a Multi"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2501.06282","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2025-01-10T15:55:27Z","cross_cats_sorted":["cs.AI","cs.HC","cs.SD","eess.AS"],"title_canon_sha256":"6545e2d4679b4c16a438e2ce824c36f7d3d60c5c7d8198babee38877c4372ee1","abstract_canon_sha256":"b81e6c8a5e4cef12060ce4b751e9a1a5234e57b15fed0d772004c2448bbbc76f"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:00:00.678501Z","signature_b64":"uV2P5wXUYmfiU6O2qRQ4KKV6m33St0x6uft3vHKT8q0bdp09UvogJlDmEmOw5kWhPEy/kqAgVAzj76Z3uKuABA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"be981eeee46abbab64d6c79dcb7e78930880a4974c6b2a377d95fb3bb48be160","last_reissued_at":"2026-07-05T10:00:00.677996Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:00:00.677996Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"MinMo: A Multimodal Large Language Model for Seamless Voice Interaction","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.HC","cs.SD","eess.AS"],"primary_cat":"cs.CL","authors_text":"Baosong Yang, Bin Ma, Changfeng Gao, Chong Deng, Chongjia Ni, Chong Zhang, Fan Yu, Guanrou Yang, Haoneng Luo, Hao Wang, Hui Wang, Jialong Tang, Jiaqing Liu, Jinren Zhou, Mengzhe Chen, Nan Zhao, Pei Zhang, Qian Chen, Qinglin Zhang, Ruize Gao, Shiliang Zhang, Tianyu Zhao, Wen Wang, Xiang Lv, Xian Shi, Xian Yang, Yabin Li, Yafeng Chen, Yanni Chen, Yexin Yang, Yingda Chen, Yunlan Xu, Yuxuan Wang, Zhifu Gao, Zhihao Du, Zhijie Yan","submitted_at":"2025-01-10T15:55:27Z","abstract_excerpt":"Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and human-like conversations. Previous models for voice interactions are categorized as native and aligned. Native models integrate speech and text processing in one framework but struggle with issues like differing sequence lengths and insufficient pre-training. Aligned models maintain text LLM capabilities but are often limited by small datasets and a narrow focus on speech tasks. In this work, we introduce MinMo, a Multi"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2501.06282","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2501.06282/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2501.06282","created_at":"2026-07-05T10:00:00.678055+00:00"},{"alias_kind":"arxiv_version","alias_value":"2501.06282v1","created_at":"2026-07-05T10:00:00.678055+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2501.06282","created_at":"2026-07-05T10:00:00.678055+00:00"},{"alias_kind":"pith_short_12","alias_value":"X2MB53XENK52","created_at":"2026-07-05T10:00:00.678055+00:00"},{"alias_kind":"pith_short_16","alias_value":"X2MB53XENK52WZGW","created_at":"2026-07-05T10:00:00.678055+00:00"},{"alias_kind":"pith_short_8","alias_value":"X2MB53XE","created_at":"2026-07-05T10:00:00.678055+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":20,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.19453","citing_title":"A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2606.13544","citing_title":"Adaptive Turn-Taking for Real-time Multi-Party Voice Agents","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2606.06559","citing_title":"IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2606.30944","citing_title":"Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2606.30553","citing_title":"COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2606.30553","citing_title":"COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2606.13544","citing_title":"Adaptive Turn-Taking for Real-time Multi-Party Voice Agents","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01016","citing_title":"PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2606.11167","citing_title":"Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2502.11946","citing_title":"Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2509.26388","citing_title":"Game-Time: Evaluating Temporal Dynamics in Spoken Language Models","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2507.16632","citing_title":"Step-Audio 2 Technical Report","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2505.17589","citing_title":"CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10199","citing_title":"How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05927","citing_title":"Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2504.18425","citing_title":"Kimi-Audio Technical Report","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06765","citing_title":"VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing","ref_index":48,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05927","citing_title":"Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2503.20215","citing_title":"Qwen2.5-Omni Technical Report","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2604.14604","citing_title":"Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection","ref_index":44,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/X2MB53XENK52WZGWY6O4W7TYSM","json":"https://pith.science/pith/X2MB53XENK52WZGWY6O4W7TYSM.json","graph_json":"https://pith.science/api/pith-number/X2MB53XENK52WZGWY6O4W7TYSM/graph.json","events_json":"https://pith.science/api/pith-number/X2MB53XENK52WZGWY6O4W7TYSM/events.json","paper":"https://pith.science/paper/X2MB53XE"},"agent_actions":{"view_html":"https://pith.science/pith/X2MB53XENK52WZGWY6O4W7TYSM","download_json":"https://pith.science/pith/X2MB53XENK52WZGWY6O4W7TYSM.json","view_paper":"https://pith.science/paper/X2MB53XE","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2501.06282&json=true","fetch_graph":"https://pith.science/api/pith-number/X2MB53XENK52WZGWY6O4W7TYSM/graph.json","fetch_events":"https://pith.science/api/pith-number/X2MB53XENK52WZGWY6O4W7TYSM/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/X2MB53XENK52WZGWY6O4W7TYSM/action/timestamp_anchor","attest_storage":"https://pith.science/pith/X2MB53XENK52WZGWY6O4W7TYSM/action/storage_attestation","attest_author":"https://pith.science/pith/X2MB53XENK52WZGWY6O4W7TYSM/action/author_attestation","sign_citation":"https://pith.science/pith/X2MB53XENK52WZGWY6O4W7TYSM/action/citation_signature","submit_replication":"https://pith.science/pith/X2MB53XENK52WZGWY6O4W7TYSM/action/replication_record"}},"created_at":"2026-07-05T10:00:00.678055+00:00","updated_at":"2026-07-05T10:00:00.678055+00:00"}