Pith. sign in

Paper Citation Record · LEDGER

Ming-Omni: A Unified Multimodal Model for Perception and Generation

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 30 inbound Pith citation observations for arXiv:2506.09344.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09344 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:58:10.657355Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:11.838092Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved36
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 97c7ba52-f118-4652-bb40-c2f3d27a5a33 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:06.798089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:06.798089Z digest=sha256:f08910c6be2d4d6ac1535508cb7967914e67a740dd32d58a5dc389fcd4c3379a

Observation c3067baf-60f0-4653-b297-95bb30c869df · outbound

This paper cites Qwen2-Audio Technical Report.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Qwen2-Audio Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.024845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.024845Z digest=sha256:8a0d5e068b5926ec2a856873162d7efbc798e54a6b68b6b6e09c4e0a78a5fcdc

Observation 682af26a-c8b6-492e-ba35-32edade64f8a · outbound

This paper cites Image-to-Markup Generation with Coarse-to-Fine Attention.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Image-to-Markup Generation with Coarse-to-Fine Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.289619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.289619Z digest=sha256:feb7095432c480639770efcfdb3cfbf4712505e6864ce3a4605d9883142957c5

Observation 374d689e-3863-4d1c-8886-b7d8710bd148 · outbound

This paper cites Chaoyou Fu, Y uhan Dai, Y ondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Y unhang Shen, Mengdan Zhang, et al.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Chaoyou Fu, Y uhan Dai, Y ondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Y unhang Shen, Mengdan Zhang, et al

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.341301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.341301Z digest=sha256:15a6716544519a747c5d8f743b8b6cb735953e5def31b09ad4637a727a681f08

Observation 007423dc-82e9-40c5-be88-0607471da412 · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

Ming-Omni: A Unified Multimodal Model for Perception and Generation The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.524656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.524656Z digest=sha256:cdeb757e38bd7824779c20e8e516b6461f327c039a63909d606b3de28e174e1c

Observation 70a0d441-835b-48cb-9407-933b084066dc · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

Ming-Omni: A Unified Multimodal Model for Perception and Generation FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.617554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.617554Z digest=sha256:a29e0dd7b5241f66b5c4873a92691b671f66393269eaef35b5c37a93dfc2e29b

Observation fa731db0-4fbb-40fc-9444-209b2dd5062e · outbound

This paper cites M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance.

Ming-Omni: A Unified Multimodal Model for Perception and Generation M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.736425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.736425Z digest=sha256:0692a62cb7509edb1e62986d1b5d22b349c87fb2392fec2accca10f1d973957e

Observation 608617b7-2ee3-443b-8385-e8a2114b71e8 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Ming-Omni: A Unified Multimodal Model for Perception and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.970950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.970950Z digest=sha256:7890f98dc68182e21a7947e6b3a5921307bd03e8e4f50a5136f1c2cc4ef0dbc5

Observation 852eed35-e376-4708-b0b5-c93c3bc3d385 · outbound

This paper cites Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:08.031807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:08.031807Z digest=sha256:d3c927f029ef380352d0cbeae6237356dec93c8940e4ba9a5d2e14bab33f393d

Observation c10826b0-b811-4ac8-848d-710487f0c234 · outbound

This paper cites EgoTaskQA: Understanding Human Tasks in Egocentric Videos.

Ming-Omni: A Unified Multimodal Model for Perception and Generation EgoTaskQA: Understanding Human Tasks in Egocentric Videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:08.132037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:08.132037Z digest=sha256:010cae60eebb4cd499e0e77e8df440b44e2f0d2cff5b557e67f15b8253bd9836

Observation 1d6e47b9-39d2-47d4-8707-2ace687eb27f · outbound

This paper cites GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning.

Ming-Omni: A Unified Multimodal Model for Perception and Generation GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:08.217484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:08.217484Z digest=sha256:4ef8a3ad419668c5244b3517c0722b4a63c1275c5da442f0cdbac44afca93f5d

Observation 45760fc5-91b9-4438-85c6-7bc341c48d89 · outbound

This paper cites Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-07T04:58:08.521682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:08.521682Z digest=sha256:781ac4671834c12c59fb384ee5411ccbf6605c027f59e0b037b88cbb60c38e64

Observation 241556af-9e91-45bc-b9e4-46a607fba19f · outbound

This paper cites SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering.

Ming-Omni: A Unified Multimodal Model for Perception and Generation SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:08.699216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:08.699216Z digest=sha256:e6f45116170e804835485135fc8a4dab636553ffb5629667e8ab0de9498d4325

Observation 37d6edeb-42af-4c5b-a1e8-f2372a242aa4 · outbound

This paper cites doi: 10.18653/v1/2022.findings-acl.177.https://aclanthology.org/2022.findings-acl.177/.

Ming-Omni: A Unified Multimodal Model for Perception and Generation doi: 10.18653/v1/2022.findings-acl.177.https://aclanthology.org/2022.findings-acl.177/

Reference 22

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T04:58:10.687629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:58:08.892938Z digest=sha256:7dd3800b18b810858f67731989b83dbaabd2ed2ef48e85475b53364020319c89

Observation ac6a8a02-1cf9-4fc5-ba9e-97f7ff1cced6 · outbound

This paper cites an unresolved cited work.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Unresolved cited work

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-08-07T04:58:09.033058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.033058Z digest=sha256:88826712621d01e3e22e5368901345f5f74ac56f887b5db0d638acbdb9727370

Observation 46ecc3aa-3986-4892-9349-3a8f625cc6e7 · outbound

This paper cites OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations.

Ming-Omni: A Unified Multimodal Model for Perception and Generation OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.094956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.094956Z digest=sha256:f77a0cdc2525ac1159358a162ad2d79d6b33360f63f8223933c485d450838f0a

Observation f2333b94-35d2-4b2f-a9d0-565eff4f8b12 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Ming-Omni: A Unified Multimodal Model for Perception and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.209905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.209905Z digest=sha256:8e5049e5e990e844f3b02dbdb36029da445b125c7bbf0870e0cc5adfd7a04723

Observation 7eafccbb-e235-443d-8afd-81a331039c9e · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

Ming-Omni: A Unified Multimodal Model for Perception and Generation MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.373118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.373118Z digest=sha256:662f73bdf1005129d26a4d58337855851dc66a0424883d3dbd17f48173f0a330

Observation f69c249c-b6b3-444c-90c8-e4e86784bdca · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

Ming-Omni: A Unified Multimodal Model for Perception and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.513337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.513337Z digest=sha256:f5346dc2481aea9925d6ed0cae085a12d5c4369d861682d0eaeb00b2313a4167

Observation bf3a7020-ca82-45ec-a2de-6bcdbce815e5 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Ming-Omni: A Unified Multimodal Model for Perception and Generation CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.654968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.654968Z digest=sha256:c25d91ac11bc9120d7430fe0cf21d53fba91529db4d9d2e72807c8cfb3a50f0e

Observation c07e59de-411a-49eb-bf6f-bea0b08dda42 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Ming-Omni: A Unified Multimodal Model for Perception and Generation LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.864477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.864477Z digest=sha256:9103200c7e9239ec9f48d2886b54be9e4248c4db591a852648942b9634cac519

Observation d59c8a55-e419-462d-96a2-21ec867793db · outbound

This paper cites Solving geometry problems: Combining text and diagram interpretation.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Solving geometry problems: Combining text and diagram interpretation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:11.694092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:58:09.991644Z digest=sha256:673922dc3c357ff90bf7583dfdfc606397e92b049182583f41b3dd6deb7a6c3b

Observation 9789e5df-5103-4315-9e4b-142f38f3933b · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Ming-Omni: A Unified Multimodal Model for Perception and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.204457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.204457Z digest=sha256:6210afc6388f8d23768d238eae3fcd683a47343bb7e5a7fe39b3198d46514ac8

Observation b49d0d64-18d9-4b4d-8235-aa46a9201558 · outbound

This paper cites CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus.

Ming-Omni: A Unified Multimodal Model for Perception and Generation CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:58:10.891623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:58:10.339655Z digest=sha256:c55e3dfcd3a544f9e44550797976427de3532b162de9b5b2e37a230563a0c28e

Observation 32ca496a-6bf1-4228-b6be-b1ab8a8463ec · outbound

This paper cites Slidespeech: A large scale slide-enriched audio-visual corpus.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Slidespeech: A large scale slide-enriched audio-visual corpus

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.428391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.428391Z digest=sha256:88359fbcd9473f4cd4bd711b1908ef2f621273c456cd75cd3db7071c30e41d32

Observation 36a390e8-a972-421e-8826-460332d0dfae · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

Ming-Omni: A Unified Multimodal Model for Perception and Generation OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.542544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.542544Z digest=sha256:9fe2438ed14b682ea1b7948922cf26ef56f8facee657ab5bda17c00925fd455e

Observation a82d713d-511c-41f1-8255-6d2bf0604a2a · outbound

This paper cites ISBN 9798400701085.

Ming-Omni: A Unified Multimodal Model for Perception and Generation ISBN 9798400701085

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.633339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.633339Z digest=sha256:41727c9d5f9cd486fb66a9f13930d7554c1d19bcd97af262e3af5341e83ab32f

Observation 4bb92906-f6d6-4b38-a808-d40f05bc41ba · outbound

This paper cites Qwen2.5-Omni Technical Report.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Qwen2.5-Omni Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.636410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.636410Z digest=sha256:8a5045a7ae074b6fb0968717a22367a55156e4e432b3d5374661e4591f789fc4

Observation 049d0006-7368-460b-a253-3cfe11c2066f · outbound

This paper cites Vript: A Video Is Worth Thousands of Words.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Vript: A Video Is Worth Thousands of Words

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.639605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.639605Z digest=sha256:bfc714f7ec561c9382ca58b3edf34178b59db4e3455d24656a28032813835555

Observation 1f598556-2f20-496f-b72b-02e98b4cf9ad · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Ming-Omni: A Unified Multimodal Model for Perception and Generation R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.642605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.642605Z digest=sha256:d7b20e5e53ef8ddaf57a1eb63d16f7d9763967e858d5a4896b9a90dce4f4f3dc

Observation 41c20c2a-d8e2-49e7-88b9-94142f01a28d · outbound

This paper cites Fan Y u, Shiliang Zhang, Yihui Fu, Lei Xie, Siqi Zheng, Zhihao Du, Weilong Huang, Pengcheng Guo, Zhijie Y an, Bin Ma, Xin Xu, and Hui Bu.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Fan Y u, Shiliang Zhang, Yihui Fu, Lei Xie, Siqi Zheng, Zhihao Du, Weilong Huang, Pengcheng Guo, Zhijie Y an, Bin Ma, Xin Xu, and Hui Bu

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:11.532420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:58:10.645410Z digest=sha256:3572e15dfe99a48fe28332b66a80e209a5f477b29ec8342f39b2424c6a403421

Observation 9e141403-3595-47f2-9a8b-2d1e9bd1a02a · outbound

This paper cites Icdar 2023 competition on structured text extraction from visually-rich document images,.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Icdar 2023 competition on structured text extraction from visually-rich document images,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:11.488768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:58:10.648677Z digest=sha256:c4bbb12de34e64168ee6df6421b8632e7fb0653de66f7f41ed672d8a7609ec14

Observation 53a379cc-53de-4149-a087-873789b7006f · outbound

This paper cites ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images.

Ming-Omni: A Unified Multimodal Model for Perception and Generation ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.651374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.651374Z digest=sha256:f2882a2aa5f6326eefa07bfbfc64d027d74ee0dca2c1335f9dc5647a6b1efeb4

Observation 19cd9c47-a240-4442-9e8b-0599fd5ce481 · outbound

This paper cites WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition.

Ming-Omni: A Unified Multimodal Model for Perception and Generation WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.654286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.654286Z digest=sha256:76fcda01d5b61349be31438135b2ae122d365019ffb30a36d4b73851c9365018

Observation 4defd50b-aaff-4b29-9951-93ad0d5ea30d · outbound

This paper cites an unresolved cited work.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:58:11.450100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:58:10.657355Z digest=sha256:90d1460965c413205235615b590f9a3d3a853c910535928c9ccc4ae57a0edfbb

Observation 9d290ed7-dc77-45a3-b4e6-29d20b1418fa · outbound

This paper cites How to Train Data-Efficient LLMs.

Ming-Omni: A Unified Multimodal Model for Perception and Generation How to Train Data-Efficient LLMs

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.750924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.750924Z digest=sha256:581e14dcb0e43b1d5d774625e0253a940ea9cbb966231652be47f943874e2d2d

Observation d2e49b1d-b157-49f3-9bcf-4b9f40b7fff0 · outbound

This paper cites Kimi-VL Technical Report.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Kimi-VL Technical Report

Reference 2014

Resolution
malformed identifier
no resolver link, observed 2026-08-07T04:58:08.350784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:08.350784Z digest=sha256:a4ef88dacbc6829a4265563c68985b64b78317b81d7898fbef30af07cf448f3d

Observation 87b56999-0a42-4485-8fba-9bd04296c0b8 · outbound

This paper cites doi: 10.18653/v1/D15-1171.https://aclanthology.org/D15-1171/.

Ming-Omni: A Unified Multimodal Model for Perception and Generation doi: 10.18653/v1/D15-1171.https://aclanthology.org/D15-1171/

Reference 2015

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T04:58:10.979960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:58:10.103677Z digest=sha256:dfdefe868d022b7f1788d1f97badf6f02c7b19e093d3051959529ba0c3a64e4e

Observation 63b21592-3895-4dd0-b2c3-1a91ca1766ec · outbound

This paper cites Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:08.785940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:08.785940Z digest=sha256:9d0b4d85cf3d38102d80f1111974ce8dae80713281483f8d6a46daa691186a12

Observation df7893b3-3daa-45b7-9005-e5980e6ac6e0 · outbound

This paper cites Chemvlm: Exploring the power of multimodal large language models in chemistry area, 2025.https://arxiv.org/abs/2408.07246.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Chemvlm: Exploring the power of multimodal large language models in chemistry area, 2025.https://arxiv.org/abs/2408.07246

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:08.412448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:08.412448Z digest=sha256:b04a5455361c5954399b281b7919e2bc4db33a69ee260ff1def4a131cb6a5434

Observation 8e9975ef-c8b3-42a2-a663-b1273312b613 · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression.

Ming-Omni: A Unified Multimodal Model for Perception and Generation UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.001154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.001154Z digest=sha256:6efebb94b5b1d1f3850d44ad1fb3624c25b272ad5cbb029c428be81f7045c9e4

Observation 96fbd220-840b-4192-b178-0abc448886f6 · outbound

This paper cites an unresolved cited work.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:58:11.904982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:58:06.958681Z digest=sha256:06ed01760b5b1129ae99d0c6dc3240fa48543a5081224d7700f221fe916ad60c

Observation 414aa10c-dfc5-4e9a-aea4-45c25b413918 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.163409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.163409Z digest=sha256:118c81d22cd8a3380952abcad59227fc3d4b3f12e0e5d17cc5b6c5c6cc43a379

Observation 7bf8485e-b10a-48ca-b658-b84698963118 · outbound

This paper cites Qwen2.5-VL Technical Report.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Qwen2.5-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:06.891957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:06.891957Z digest=sha256:e207deb3f66fbadfff4a815b7eceb061d73f8ae2186b23ab7506c387ae3d5806

Observation 619ad11f-4d92-493b-ac30-edf1ab09139c · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.862140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.862140Z digest=sha256:58cb42d5795707e376335a29b468c894c87b5bd6d4b003121026cdfe94c62b2b

Pith citing papers

Observation f5a44a29-2cb8-4c17-8dc1-78f83a0dfabf · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.838092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.838092Z digest=sha256:1d003ad88163d15fac541cc337bf16427ed606030311844dfc3d99511b827f03

Observation d3abfd34-b80c-4e53-9850-34275c614d48 · inbound

MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications cites this paper.

MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:11.299104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:11.299104Z digest=sha256:479e996e75a5b67fa1433dfdab57ec5b6e3760912bd32ecc5b20223783517b97

Observation 81019c43-5f2f-4939-90ec-c12646c4b2ac · inbound

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models cites this paper.

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:15.104995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:45:37.418493Z digest=sha256:b0cfdddcc87ca1255a875bbb673fda1960d6f804863741d562e60f046a92cd95

Observation 59e6e5ba-9d8b-4c72-a1ef-1317b5dd4ecd · inbound

A Benchmark for Omni-Modal Reasoning in Long Videos cites this paper.

A Benchmark for Omni-Modal Reasoning in Long Videos Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T15:26:12.673638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:26:12.673638Z digest=sha256:c99a822d390e1b78fa9c4ff46a869d31bc364541021fa0fa32b0572e61484296

Observation 74fd0d8d-a4eb-4c6d-80d7-b164acdd6b87 · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:10:43.160927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:9dc113b5a506d9dde7183119741ac4de4e0246db2b87a0076b7a63809825fea5

Observation 2d20c179-b2fc-46c8-9617-2633fc13e696 · inbound

Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training cites this paper.

Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:12:35.432613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T08:12:25.063204Z digest=sha256:e03480d66dfe3bcbaa5a4aff6ee75aaa3cb85efd152feb60158d2f7d19710ff3

Observation 69f0c9b0-bad1-475f-857c-98595b03aa46 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 237

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.929867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:751e5e04b2068b850907971fa4d559462551673b18b6048646a6486908a094ed

Observation 761672d4-f70a-472c-9291-12db54bc66cc · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.289867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T14:10:03.707886Z digest=sha256:c8aafa891f61804797c0e18649315c2360463f1d24cd8bc6f8e037d6d7c98176

Observation c9f1329d-81bd-4abe-9103-0004a9efcaa5 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T21:18:46.566338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:18:46.566338Z digest=sha256:eadb76ff456a9aadfad5f97d34cdc3cd49224cca428cbde2fb8af5e9903cf90c

Observation 9ab75913-f338-49f0-aab7-c3e6c54e60dc · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.069624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:3542a30b956f2e5dff2772f484911a96e5f645e5114d2f3ab1fc365bac70f9fe

Observation 61e884c7-de6c-4fe5-b20f-6fa77e4eefd9 · inbound

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models cites this paper.

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.701157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T07:02:02.752466Z digest=sha256:49f196b9db3c81bd9742b569d8c2e88775c7992dc28235209622d4154a1ae544

Observation acd9c863-5d3b-4c7f-bc21-31b756b26a82 · inbound

Context Unrolling in Omni Models cites this paper.

Context Unrolling in Omni Models Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:04.985281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T22:02:57.841111Z digest=sha256:6597a938c4bb2cbad10babdaae9b7f6f26d34d2d768e1e0e3f401ac837358689

Observation d2be920c-3348-412f-aa57-6314db1164f4 · inbound

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs cites this paper.

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:13.497655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T04:41:52.098355Z digest=sha256:fd7584af890ae05fedc01e6c2ac6d5289c8c8b54d426c7ea78b3962307874137

Observation 6e5184cd-82a6-4083-9a34-8a156e5a6c1b · inbound

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection cites this paper.

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:06:04.247249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T14:04:52.065878Z digest=sha256:f61aeff9b829519dc2d5ad29da0fdcaa62bd279327d11492ea8467d6d56e5700

Observation befb3024-ca75-409d-a5ed-cdabcf8ad0aa · inbound

Accelerating Compound LLM Training Workloads with Maestro cites this paper.

Accelerating Compound LLM Training Workloads with Maestro Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:25.597147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T05:11:46.845356Z digest=sha256:009eb5564f90e4f646cae50f8849ca5bb0ab32d0448224a630e697d98ce3151d

Observation d8a9558d-9a54-4723-b445-58b8bf5b836e · inbound

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models cites this paper.

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:16.317482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T04:52:03.076788Z digest=sha256:1d9e6003f404dd8271398285ea3aea9610e5e9d3b5ff188593d9dfcd06d36d8e

Observation a59024a1-6a22-4c83-b3f2-a717820d62e9 · inbound

When Vision Speaks for Sound cites this paper.

When Vision Speaks for Sound Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.891952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T22:12:52.160596Z digest=sha256:6a4a5b65038723696ef41c4e9bef5fd38166255f9e64350535701e5b317ce7f5

Observation f9ee4736-29c4-4eb7-a044-59007715992e · inbound

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation cites this paper.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.784211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:05249fdf50fb58b8865964e0c3d30f14d92c8984ee25edf92ffb90197d611b1f

Observation c38594da-f5fd-434a-8324-fb58ef1d8085 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:02.008433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:4122f9a62ed27f10a6d8b549c84cec0cfc09101a2ddbbd06b5b7afd640de240a

Observation df39790b-feea-4da2-bf12-88b636dfa4a9 · inbound

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain cites this paper.

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:23:18.014835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T10:20:12.860927Z digest=sha256:9250727a257640588b782d994a56f694a1ce8dd42231f7d70e43f0a8268258dc

Observation d996737e-1647-4160-bc43-d7ba78b8ea99 · inbound

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation cites this paper.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.238741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:24f3cac005abcb970173b4f4eb8c553ab35a8aaa697dae993f7ec04a46fbd4ce

Observation 3b14cc91-248a-4b17-a5c9-f6a3c32c0e3e · inbound

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models cites this paper.

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:00.560543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T22:56:21.783415Z digest=sha256:02d6ae5b44f65d0edada6ee5636de803d2dcb9c34f5b8bb6c0389467ee87bb5d

Observation 08c308ba-da7e-4407-9d0a-15ed987b1931 · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:52:27.077528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:b40a66ed94527896f32ea20735895cdb62068ed159b0cd80309cd7f808bc656c

Observation 02327fc4-30ea-44bc-9b36-2beebe1677fc · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.025507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:7f7a10446aaa68fa3b8034e34842d7f206516e6f9be8712477bfebfd91a48a25

Observation ee1207c0-46e7-4a8e-ab52-f27799133443 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.273485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:446b9b0fd518189ebc80ce95134c821f8f6bc1dd3e55556eda668f429fa572d7

Observation 90f74144-9e31-4f1f-aa25-c288041dabe8 · inbound

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs cites this paper.

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.658325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T07:26:07.356352Z digest=sha256:159c6028724a3ddc075275c88c8bc79ef19c91ba8ef9be574c77dafc629efaca

Observation 6d083d17-7830-4795-8cfb-bb967d85bd16 · inbound

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation cites this paper.

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.906567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T07:24:01.244733Z digest=sha256:81759e9160bef4cbab498a0acab72f54bc525083e2ed97dea917726bad859ecf

Observation 0bf63540-2e1d-4b51-a92a-1afe5db3e18e · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.826743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:47a22edbbae9a54501414ea04b1541141e735a4f2a69c6c349f5746c3e5f1d9d

Observation 2dc81df9-a333-482f-b9d6-c0138f50139a · inbound

HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving cites this paper.

HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:27.262829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:27.262829Z digest=sha256:facecd5cfb3cea98fd793c8c45a05a69980c98b09fd20e9a75c6e713d419ae55

Observation 50206a3c-3d6f-47da-8401-f85120f2b15a · inbound

MMAG: A Multi-Control Mixed Audio Generation Benchmark cites this paper.

MMAG: A Multi-Control Mixed Audio Generation Benchmark Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T18:47:42.569068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:47:42.569068Z digest=sha256:20a17b5e54d127862896a4703c0d6f38d9a2017701fa7e8d064654f5d05d50d7