Pith. sign in

Paper Citation Record · LEDGER

Aria: An Open Multimodal Native Mixture-of-Experts Model

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2410.05993.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.05993 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:15.480907Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:49:39.565893Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 63d04862-f1a2-4fe4-9032-23211b9dec67 · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 220

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.787101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:274d1c847d28d2ebf77636fe6393d1875f5a199ada68c68c4e2200bb8ef17cd2

Observation 6ae4961c-ef96-458c-9340-fd90ed497a19 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:23.665883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:8c4b17f59956d2d7bf67267d2003e1a33f1489b8e82caf686d7059a65dc9f998

Observation 7f7d654c-d3c0-4c4c-a437-1618750f0f72 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.351105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:b9840e02900e45e9dfd9f16471364582e5c0da16c2d2a7eef51b246bb1c624ee

Observation 760a6262-90f3-4eb9-8b22-62e863cfc59e · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.513505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:0dad5c7d61e66ced818d128ba308160faf573e42ade168ab0fd2f2a9db127406

Observation c5f32c4e-ddd6-47aa-8090-8408a836f07f · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.254763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:bf9454c7cc47c952eef2a487321355443e1b1d664fa4ad0b28d065c855e1c92f

Observation 8821343c-b7a2-4f6e-811a-0b539bf251f3 · inbound

TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments cites this paper.

TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:15.480907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:46:15.480907Z digest=sha256:1881f33fb454fa77a8111aecf240eaca6a99c86c382d9c4c8c37fcdb014edf65

Observation 7e7399e0-866e-4842-a160-4c0336e41f05 · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.870434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.870434Z digest=sha256:3b35759301aeec659c406756adf5473278d6b2127ee045396f42191bb5574fbb

Observation 21cc2dfd-dfd4-44a1-b1c5-5e5e52635d85 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.178112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.178112Z digest=sha256:3f26345cb90b3984ed1bc188384bca912e5ee0fd4138b1d9f622ff3198f03131

Observation 264fb417-cc39-45a4-a218-e19a0cc9fa8d · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.321429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.321429Z digest=sha256:a922651c2d191794be011f274d520073ca52b747d208a4ed7e3b045794b2e89e

Observation 956df445-02d3-43c9-ba84-4aa92a14787f · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.699207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.699207Z digest=sha256:20d9ffce2c4591cff24e0b337f5fcc13b22a89a8114f8eff2be2a7edffd273ce

Observation 5aa974a3-aacc-4d02-a416-d2cbbb1fc859 · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.269793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.269793Z digest=sha256:d5c5673c1fe89b3ec855a6e38db8ae349aa01a4695ebf73be0d83843ea951f62

Observation 61a95390-c468-4b95-992e-c6f6636ede24 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.895831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.895831Z digest=sha256:7ed4eac1510519eb1eba98f6cd3813e489c8d7f0db18732473326c603b177ea7

Observation f1bce6f4-2bd2-43e1-b54c-0a8033a0e85c · inbound

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities cites this paper.

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:18.685096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:18.685096Z digest=sha256:a0ec641b7f7e1c4b764b87472bafd139d5f6a8174893d9ce635ddaef87c13b4d

Observation b92b4578-1639-44c9-8c4e-488ab0c0a9d5 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.142398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.142398Z digest=sha256:a42c5d7e2b0c61153cd90385e54282cc796d3bb52fd7b72c29e13c19fccf3213

Observation bdeedada-c8a1-4732-8f62-8830251ce4c6 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.888224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.888224Z digest=sha256:22bd3377b3a47b87d2f1c7ed87fb171563ef3188ba87f3b163770bc65f1b6a41

Observation 03a9cf9c-51af-453e-893c-cbad769da65f · inbound

SeqPE: Transformer with Sequential Position Encoding cites this paper.

SeqPE: Transformer with Sequential Position Encoding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:03.582120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:03.582120Z digest=sha256:e5538b173cca7fbce2c4f866d3e177aedab4d57fd8dd82865f1d7f2ed9ef613d

Observation d45607f7-0eeb-4f4a-9271-0dce9d0007ae · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:03.076564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:03.076564Z digest=sha256:9cc969a6e97d118b525da3d384568efe8b9f0e301eefc55185e3dada10332f78

Observation e9d3eefd-96c8-4e8b-ad0a-a00b44db0101 · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:27:04.381238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:2b9c84d124a1e88b470bea2188981b9e7b5c17d614e208f4da032a0d6a250975

Observation 84e3669f-09fa-4e89-bf38-aecdd5b59f7a · inbound

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering cites this paper.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.651138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.651138Z digest=sha256:51f527d2b94ed9ba66ebb6c3b8d09dc6346f0e2222182f23ece598460104e608

Observation eead32a8-c627-48ac-b545-2cbbaea125dc · inbound

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs cites this paper.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.532425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.532425Z digest=sha256:383d1d54aa23b8259e1a293fef921a3a63d32f1c2867ddc6fb184ae41f8f31cb

Observation 81842245-3584-42af-815f-d44d8e3626c9 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.062315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.062315Z digest=sha256:fff221c0b85fd439e860c48f4bbc0d805c7bf6cc74a3ec32f6298f2feb47fedd

Observation 67cf1236-7fa1-4118-8cbe-6711b0775924 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.704165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.704165Z digest=sha256:c528056fa570be74773ef2217d1d6285017bf647066653ec50a05abecf78363a

Observation 9e98d311-ebb1-4b89-86f5-713a3c31b6e6 · inbound

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis cites this paper.

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:00.379646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:12:00.379646Z digest=sha256:834e8a87fb5bd58774c7ee91e2bf2ee0cd781a1073a822ba8be1cdfbe5e1f37f

Observation a3e25417-5864-4888-a920-35b88ceee824 · inbound

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models cites this paper.

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:28.576320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:28.576320Z digest=sha256:7a6891f6685c12bf83684b871ec604cd9ab6d863075845d154232a7e140b3c24

Observation 5512c647-879c-48fd-b611-5cd6e6d1ced4 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.827141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.827141Z digest=sha256:ee9dc8d8ac8fc77590d22f60cd6d3ba4b2f88c7778498c036ad7aed25a17074b

Observation 25639214-899c-434e-956e-1357c905573c · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.065340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.065340Z digest=sha256:50f13b9fac364d3a8735291850aaa01a86921735f9b70e22831e4f2729227788

Observation 27fc9e3e-772a-4fb5-8274-f53805fd91b9 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:42.093183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:42.093183Z digest=sha256:eddaaf48903e05deef0776ecfcdf022884abd15aa84414b03ffe3328ad8c4f38

Observation 967cacbc-4c71-454c-96d2-8db4b59de9d3 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.830990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.830990Z digest=sha256:8bf7b480324958e3494d2a7e169465c61ec97516d4642688341e5bfd74ac3b58

Observation 525d0f11-90c8-43dc-9609-5537117c5b0f · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:10.590321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:10.590321Z digest=sha256:3843d8702b6e4b86663a58164274ce8353a59ed7d80b365c524529d9c266337a

Observation 748abe4c-ecc3-4ce2-9b2c-bce63cc6da9b · inbound

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation cites this paper.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:46.702453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:46.702453Z digest=sha256:6c50beb353ddbff0f529932d108fa656b2d84f15589b4004deb802cf59c64c13

Observation 4decabbf-88ca-4b45-9f26-4fc0e6e38fec · inbound

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search cites this paper.

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:17:55.677423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:17:55.500268Z digest=sha256:523ded59bde43686deb0e91070b440aa9e81f9e4b4f95141d599d74977d32c63

Observation a133ce72-f350-4e46-b0ab-30c9f0976478 · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.586319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.586319Z digest=sha256:d5352b5ee5b74ee1c21ea25562467e269d16ac6bfd779cc3208a9818daabbc5d

Observation f84a2837-f958-4a6f-b14a-661a3cf0974f · inbound

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning cites this paper.

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:56.236585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:56.236585Z digest=sha256:a0790e1c7448e0ffa7eefd5407c20d25409ab5813fda67807a36090a67331c28

Observation 6a12fce7-745b-41bb-aa36-4dfe8bc923bf · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.431023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:71d61bfe84f14d62e02a2356c5ff59a1104c7ada9638bc2155e66194852bc7c4

Observation 20fe5682-f600-44dc-a371-8c6b83588456 · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.236502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:48a3b6d1616934b9ef457b50aecdb64e28d1efb09d01d10df1e694222d434d07

Observation cf2041fa-7470-46f5-a5cf-576a911d4d48 · inbound

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents cites this paper.

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:08.078362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T03:21:30.732925Z digest=sha256:3dcb513e91c037b4c5a14700ac2fb3cf5f4c211a914e7bb1d7589c541173f902

Observation 207c8141-ec26-4d31-a20d-2de2f3daa16c · inbound

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards cites this paper.

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:10.087841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:12:17.469552Z digest=sha256:5d8554d20b15f9b5924f909db36fa1cf8ab8463c5934335d8e73b2997fff203b

Observation 8c1a6100-ea69-4415-b286-51f3fb6d7d16 · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:48:23.469466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:86b84d4ef6b146de217e95e579f1fc17aa361732313654601a7464df52c44b34

Observation f31884c4-788a-4c9a-a17b-e330fee1ded4 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.644704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:eca9744d2f3c9488fedf0b979341bd8b30aa3b0e51e282f778adbb71c1469572

Observation 75332c6c-2cce-4d14-97c2-1754d2c6854d · inbound

MobileMoE: Scaling On-Device Mixture of Experts cites this paper.

MobileMoE: Scaling On-Device Mixture of Experts Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:53:51.431424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:48:50.656971Z digest=sha256:b7b29da94caa080dfc81b34a3d4a5a397eeef240e098211ce941a2527ea78467

Observation 36825118-971b-4ca7-b7a7-0ee70fab0eb3 · inbound

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models cites this paper.

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.060381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:43:33.929871Z digest=sha256:e4f28ec05fccaf02525700b837c26ca707c4fa8059b0a9f5e978203dbfefd711

Observation 78e9bca3-f6e7-40d1-9b2c-92288425112e · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 291

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:32.081900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:c786cc730d1dabc015b647ed408d7091479b7a5114226c2cccf87200d142e181

Observation c308a88b-ff1d-4416-a119-d94f25a1d378 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.343170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:d0364456ec5296b935f2db51ae16dc84601a821277f2bf3adfa3ca9974b5c8d8

Observation 11cbdc20-f013-4cad-adfc-486ee162293d · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.708094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:3c20d4312da74587874cf6711c95774bdffbdbbed5e887c8694615039a4ac874

Observation 834e8346-8de0-4ee5-8cd5-66661c36e10b · inbound

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware cites this paper.

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:49:39.567377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:40:53.103233Z digest=sha256:805410486d324ac4ec9b28f10b0a7c3b6da583ff059446dc1176be32f1ad9c10

Observation af36094b-e3bd-4a0e-94ea-65fee3fc4322 · inbound

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models cites this paper.

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:36.918993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:51:36.918993Z digest=sha256:c7aa9e5cf6c21108e3e6618b5e1ebe47cfe8dd2f17408bf2592262ac9fa9dc81

Observation 73900df1-0d7c-4843-b1c9-72d742d9c704 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:27.943328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:27.943328Z digest=sha256:4214627a81cf7e1a4d99117d6101f7025dd62bd4886395a6282618f1676d32a3