Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2410.05993.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:15.480907Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:49:39.565893Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 63d04862-f1a2-4fe4-9032-23211b9dec67 · inbound
Large Language Model-Brained GUI Agents: A Survey Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 220
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ae4961c-ef96-458c-9340-fd90ed497a19 · inbound
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f7d654c-d3c0-4c4c-a437-1618750f0f72 · inbound
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 760a6262-90f3-4eb9-8b22-62e863cfc59e · inbound
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5f32c4e-ddd6-47aa-8090-8408a836f07f · inbound
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8821343c-b7a2-4f6e-811a-0b539bf251f3 · inbound
TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e7399e0-866e-4842-a160-4c0336e41f05 · inbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21cc2dfd-dfd4-44a1-b1c5-5e5e52635d85 · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 264fb417-cc39-45a4-a218-e19a0cc9fa8d · inbound
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 956df445-02d3-43c9-ba84-4aa92a14787f · inbound
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa974a3-aacc-4d02-a416-d2cbbb1fc859 · inbound
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a95390-c468-4b95-992e-c6f6636ede24 · inbound
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1bce6f4-2bd2-43e1-b54c-0a8033a0e85c · inbound
SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b92b4578-1639-44c9-8c4e-488ab0c0a9d5 · inbound
CyberV: Cybernetics for Test-time Scaling in Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdeedada-c8a1-4732-8f62-8830251ce4c6 · inbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03a9cf9c-51af-453e-893c-cbad769da65f · inbound
SeqPE: Transformer with Sequential Position Encoding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d45607f7-0eeb-4f4a-9271-0dce9d0007ae · inbound
VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d3eefd-96c8-4e8b-ad0a-a00b44db0101 · inbound
MMSearch-R1: Incentivizing LMMs to Search Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84e3669f-09fa-4e89-bf38-aecdd5b59f7a · inbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eead32a8-c627-48ac-b545-2cbbaea125dc · inbound
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81842245-3584-42af-815f-d44d8e3626c9 · inbound
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cf1236-7fa1-4118-8cbe-6711b0775924 · inbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e98d311-ebb1-4b89-86f5-713a3c31b6e6 · inbound
Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e25417-5864-4888-a920-35b88ceee824 · inbound
VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5512c647-879c-48fd-b611-5cd6e6d1ced4 · inbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25639214-899c-434e-956e-1357c905573c · inbound
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27fc9e3e-772a-4fb5-8274-f53805fd91b9 · inbound
"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 967cacbc-4c71-454c-96d2-8db4b59de9d3 · inbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 525d0f11-90c8-43dc-9609-5537117c5b0f · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 748abe4c-ecc3-4ce2-9b2c-bce63cc6da9b · inbound
Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4decabbf-88ca-4b45-9f26-4fc0e6e38fec · inbound
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a133ce72-f350-4e46-b0ab-30c9f0976478 · inbound
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f84a2837-f958-4a6f-b14a-661a3cf0974f · inbound
InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a12fce7-745b-41bb-aa36-4dfe8bc923bf · inbound
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20fe5682-f600-44dc-a371-8c6b83588456 · inbound
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf2041fa-7470-46f5-a5cf-576a911d4d48 · inbound
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 207c8141-ec26-4d31-a20d-2de2f3daa16c · inbound
ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c1a6100-ea69-4415-b286-51f3fb6d7d16 · inbound
OProver: A Unified Framework for Agentic Formal Theorem Proving Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f31884c4-788a-4c9a-a17b-e330fee1ded4 · inbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75332c6c-2cce-4d14-97c2-1754d2c6854d · inbound
MobileMoE: Scaling On-Device Mixture of Experts Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36825118-971b-4ca7-b7a7-0ee70fab0eb3 · inbound
Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78e9bca3-f6e7-40d1-9b2c-92288425112e · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 291
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c308a88b-ff1d-4416-a119-d94f25a1d378 · inbound
CogniRoute: Learning to Route Social Evidence in Omni-Modal Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11cbdc20-f013-4cad-adfc-486ee162293d · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 834e8346-8de0-4ee5-8cd5-66661c36e10b · inbound
WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af36094b-e3bd-4a0e-94ea-65fee3fc4322 · inbound
Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73900df1-0d7c-4843-b1c9-72d742d9c704 · inbound
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.