Pith. sign in

Paper Citation Record · LEDGER

VITA: Towards Open-Source Interactive Omni Multimodal LLM

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2408.05211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.05211 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:02:26.139518Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1f7a9d0f-656d-42cf-bd48-b50f5982ccf4 · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.688477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:4ffe0dccc05723f16f9b1005fbcbb9cd366c70dc89708cf8a0bc273c70e66439

Observation 5c6a522d-e36f-4c50-bf24-b9a962dfe8a9 · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.611554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:12dd88df61f72cecb91b0a6ee8640d2708ef235c178a638ef93e56de01c4eab8

Observation 8a652af5-5952-4562-ac2c-eb56ebb692f0 · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:50:13.976131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:1f26deea3b373414dbf6a7a13e119a293b33e87494ef710a0b6afe1e94884808

Observation e9f845cc-9fb0-439e-bba0-162ba911374a · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.790388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:20651c9f2b0d4e595e22d5f571ee83e3f4a653f491b4580beec58e449f8d8389

Observation 7b3f034a-05a7-460c-97fd-1be3d9d0069f · inbound

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment cites this paper.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.139518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.139518Z digest=sha256:b4865380c0845ab4d812b84f7cd7e01219df289b6b5a9d8e0439f5eda7ed513c

Observation 8cc93ee0-cec6-44eb-9cc5-923a6683089f · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.128649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.128649Z digest=sha256:996a5821b064c9c6dbdfaf322c8c050ee760a2bd742b2753bf3cbee545ede67a

Observation 2eaab4af-c665-4c62-9349-65ff4b72a787 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.836288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.836288Z digest=sha256:7ae05a226b0f31f57f9e793789a0202867b955130cb7900f53623dbfb03fd674

Observation 1660023d-958b-4890-88a7-2a40bb1ad2aa · inbound

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT cites this paper.

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:36:22.224140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:36:22.224140Z digest=sha256:a5f05220ab62a63c22b10453e61df2f935def17546ce7fb2b7135e972ab95d2e

Observation f450c3a2-f6b5-401a-9db7-6da32c58d97d · inbound

FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation cites this paper.

FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.216234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.216234Z digest=sha256:80f127fa4b48cf43ef0347f564b7f7511cbf2db9b36b3c2dc76d8709aba9d50d

Observation 8e9d583b-8ef8-4035-a643-3f27cd917c3a · inbound

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems cites this paper.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.519528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.519528Z digest=sha256:26968f5ceb742a2b52e9516f0c96dc86fa30f03bd150e551a83fdc088ed98a36

Observation 448e7ee4-c789-4f20-b6a3-3023fedf271c · inbound

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities cites this paper.

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:18.654045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:18.654045Z digest=sha256:9393ff7a59e293984424d8fa620733fb73347955ea6131c6d06e3d2822d94574

Observation 6fbb984b-5ef9-49ec-abb0-41193774fd76 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.242431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.242431Z digest=sha256:5689ebd4b498669918aee1a7d712207368f5c6635b88a0420626fe695f9e9775

Observation 8cd04c31-5034-4d57-bb4d-6ba01c0ff23b · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.049331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:ec5b3c252f6ce73dcee4a5767068c7dcf32c083a9569f1fded8ffd4118a9c6cc

Observation c349ad53-5cca-41ad-8693-53519d1865a8 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.791041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:1c60ea8be30db89afe3acdcee022ba3a2799423c5b01a91432e26e0d689ea9aa

Observation f204dd5b-a186-4075-90f4-d271379ccf90 · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.402405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.402405Z digest=sha256:5cc695b57138f39ac4c189354f1655fcbee80cf5cbc2cf42213d2b318c2ca218

Observation e3c490bd-0831-4eb7-b0f5-b81e63db752d · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.705443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.705443Z digest=sha256:696c0e3fbb387a80e9e0b4bb5461608d55c7fb90b99295e8c99521daf8cf6f81

Observation 949449f7-8bea-4dd2-aedf-b8148e916e0a · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.561591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:af669ff320eda7ef5bdcbd6f58631b4db14faabf9f712cf310a2a8ea6f60d5c7

Observation 2fabbe87-6013-4be6-9b76-fe9d13619c93 · inbound

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models cites this paper.

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:15.026920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:45:37.418493Z digest=sha256:614b11e9186f42b68aebc4af41b7b2565b5abdd312951752f8b3dd97b4b49c85

Observation f88d59cb-10d4-4f67-bdfa-7ca045c807f5 · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.093691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:a366e89cf77af74781b9615d2bded128df9e688327fd6b0b0e7490cf8af9a21b

Observation 2b135497-2fdd-4a84-87f4-30298e5069b1 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:37.835015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:37.835015Z digest=sha256:4e880c30f789e7de5fea361f615d960b4f1cef7e605064497c01f5109cd162e7

Observation 113d6a17-9f06-471a-9729-bfa1aa0e3805 · inbound

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation cites this paper.

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T13:21:08.665823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:21:08.665823Z digest=sha256:5d2bf072181aa55c329b2f19357d62dc8567cdc31808df0d3cb4d230c734a7fb

Observation ef426965-d6e5-4257-a0cd-92f627bc1131 · inbound

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion cites this paper.

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-15T13:43:40.241796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:43:40.241796Z digest=sha256:251db9ce15465d821d9cdc0d00acc08aff99d2cc92f0e327e6aee3c802771287

Observation 1695952e-1df4-41f1-b506-c29dd8021035 · inbound

RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models cites this paper.

RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:25:18.402492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:24:18.061622Z digest=sha256:de722ddf473288cd26b3a2932be4e3d0b7e34c4935cc533723ff39a14db33b49

Observation ca753013-6163-4fbc-8f4e-c942c90458a8 · inbound

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models cites this paper.

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.690950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T07:02:02.752466Z digest=sha256:fd6dd980930bf1e05a0f557d727fd79408755ea71e1b0da30340a1901d993ecc

Observation 97036cbc-6701-4384-ab0b-356dbe85f8e6 · inbound

UAF: A Unified Audio Front-end LLM for Full-Duplex Speech Interaction cites this paper.

UAF: A Unified Audio Front-end LLM for Full-Duplex Speech Interaction VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:02.870975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:05:48.624864Z digest=sha256:1c1a6ac9c39263c84b6c36d1fc4c788860838193656e3ba4443e3e8538110283

Observation 4c36f2f1-afd8-4f2c-b1df-76881acfcfc2 · inbound

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness cites this paper.

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:42:01.619900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:24:23.119301Z digest=sha256:4692fa92e806eb1a2a040ce515c34f42a242aad2cdc52cb159f4e1745ea18f6b

Observation 350a4240-24af-42dd-8033-7a91c8674126 · inbound

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model cites this paper.

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:08.314647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T15:34:50.848124Z digest=sha256:d002d32b104312ba4abe47e2255769e868635b08cfbcd05d6fa2de6d767a388b

Observation 5d47c216-597c-43ec-a82e-93a3c66d46a6 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.196125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:5cc5af5a209007a5f4b0431cd8674dd08059a2adfdd3c286ddbce32a837bf2fd

Observation 617f29fc-d179-484e-a997-20ca2f138818 · inbound

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue cites this paper.

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:19.057777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:08:55.753359Z digest=sha256:665d709e5ec0d285ea97ad75061327d07d88bc0c0b3cddfc26d2002a9eca52de

Observation 5efd09f5-0c4c-4cbf-ae02-4af985ca8298 · inbound

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models cites this paper.

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:16.323914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T04:52:03.076788Z digest=sha256:bc22ba2469e826fb028423ca9908bba6a02680309566a0c409c9bf1ea1a47eec

Observation 54d54f17-5860-4175-87ec-a07557b1b8da · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.197089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:f56afa8a4c086139519dd712a81666d35ebba0fde9d980380540d46cddd119e1

Observation 233be047-3760-41e3-9f02-9057080e1e2d · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.323242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:e4a3c20e38624415de0b0baea4faae429f631bfb71b93f3d0cbd1fa9ae453ea4

Observation ea4c177f-4d42-4ebd-9f4b-3cf0d1353fd2 · inbound

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments cites this paper.

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:19:59.971709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T10:16:42.095330Z digest=sha256:b7d087f3e832b7e1358c61c2fb2e61cc92343996122e6907ff4e0d2f9eaff09c

Observation 5f4e2f4a-211e-4af9-84de-35a40c989d24 · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:09:24.302796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:5ca3fb7e735af9dcd1632815f8d27b0b94b2baefd3a33bebdcb586bfbea3d36e

Observation f5effd5b-0ec5-4d52-9c6d-8567efe829f1 · inbound

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding cites this paper.

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:50.869060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T18:36:29.376048Z digest=sha256:94e07bd94e0ba8e595008f1c6bca39e0e9963f8a41530e1430b87453244b8c2b

Observation baedd8d4-0e72-43d1-949d-e5d35620a67c · inbound

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models cites this paper.

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:20:57.394243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T12:53:00.569655Z digest=sha256:4bad9fbc55b5206960e0ef8d4298bd4bcc0e09439f0dc5cf394480be090f3878

Observation b4c4a14f-f9ea-4a76-b00f-8b8505f838fd · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 134

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.859547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:7678e8a5f3e8f6d3a643c0217400ebf578a5243ed5728e2d5595492e89c9ff2d

Observation 79df11e2-4c48-4224-a532-a2c7c461dbda · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.597012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:ec3fd7fb07aae3d2c90e65bf810440f5b856c46683b22907426163e5b78affa4

Observation 2d06a823-eb47-47d0-b9b3-263e1e0e86b9 · inbound

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs cites this paper.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.583375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:4e9534fed5d866e2acbcf630b77b5b7a60adbc7eba48b3d5b68382a958825c36

Observation 1563059e-d21b-4e68-84a5-fffc1a5b873d · inbound

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models cites this paper.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.776575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.776575Z digest=sha256:fe9404f3e9de1530c863b9ae030820149cb6e3b6657bb0ea285362220eb7e9bb