Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2408.05211.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T00:02:26.139518Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
4
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 1f7a9d0f-656d-42cf-bd48-b50f5982ccf4 · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5c6a522d-e36f-4c50-bf24-b9a962dfe8a9 · inbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8a652af5-5952-4562-ac2c-eb56ebb692f0 · inbound
VoiceBench: Benchmarking LLM-Based Voice Assistants VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e9f845cc-9fb0-439e-bba0-162ba911374a · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7b3f034a-05a7-460c-97fd-1be3d9d0069f · inbound
AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cc93ee0-cec6-44eb-9cc5-923a6683089f · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eaab4af-c665-4c62-9349-65ff4b72a787 · inbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1660023d-958b-4890-88a7-2a40bb1ad2aa · inbound
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f450c3a2-f6b5-401a-9db7-6da32c58d97d · inbound
FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e9d583b-8ef8-4035-a643-3f27cd917c3a · inbound
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 448e7ee4-c789-4f20-b6a3-3023fedf271c · inbound
SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fbb984b-5ef9-49ec-abb0-41193774fd76 · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd04c31-5034-4d57-bb4d-6ba01c0ff23b · inbound
Training-Free Multimodal Large Language Model Orchestration VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c349ad53-5cca-41ad-8693-53519d1865a8 · inbound
Training-Free Multimodal Large Language Model Orchestration VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f204dd5b-a186-4075-90f4-d271379ccf90 · inbound
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c490bd-0831-4eb7-b0f5-b81e63db752d · inbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 949449f7-8bea-4dd2-aedf-b8148e916e0a · inbound
DeepEyesV2: Toward Agentic Multimodal Model VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2fabbe87-6013-4be6-9b76-fe9d13619c93 · inbound
OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f88d59cb-10d4-4f67-bdfa-7ca045c807f5 · inbound
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2b135497-2fdd-4a84-87f4-30298e5069b1 · inbound
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113d6a17-9f06-471a-9729-bfa1aa0e3805 · inbound
EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef426965-d6e5-4257-a0cd-92f627bc1131 · inbound
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1695952e-1df4-41f1-b506-c29dd8021035 · inbound
RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca753013-6163-4fbc-8f4e-c942c90458a8 · inbound
Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 97036cbc-6701-4384-ab0b-356dbe85f8e6 · inbound
UAF: A Unified Audio Front-end LLM for Full-Duplex Speech Interaction VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4c36f2f1-afd8-4f2c-b1df-76881acfcfc2 · inbound
EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 350a4240-24af-42dd-8033-7a91c8674126 · inbound
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d47c216-597c-43ec-a82e-93a3c66d46a6 · inbound
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 617f29fc-d179-484e-a997-20ca2f138818 · inbound
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5efd09f5-0c4c-4cbf-ae02-4af985ca8298 · inbound
OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 54d54f17-5860-4175-87ec-a07557b1b8da · inbound
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 233be047-3760-41e3-9f02-9057080e1e2d · inbound
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea4c177f-4d42-4ebd-9f4b-3cf0d1353fd2 · inbound
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5f4e2f4a-211e-4af9-84de-35a40c989d24 · inbound
A Survey of Audio Reasoning in Multimodal Foundation Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f5effd5b-0ec5-4d52-9c6d-8567efe829f1 · inbound
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation baedd8d4-0e72-43d1-949d-e5d35620a67c · inbound
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b4c4a14f-f9ea-4a76-b00f-8b8505f838fd · inbound
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 79df11e2-4c48-4224-a532-a2c7c461dbda · inbound
Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2d06a823-eb47-47d0-b9b3-263e1e0e86b9 · inbound
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1563059e-d21b-4e68-84a5-fffc1a5b873d · inbound
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.