Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T10:53:26.574350Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 100 inbound Pith citation observations for arXiv:2509.23661.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T10:53:26.574350Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:15:04.579056Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
12 of 12 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
Observation cbc9b1a2-80e2-4ec9-9357-be6e956ad599 · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b62095e3-85b8-401e-96ee-90db625f3260 · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6be7f074-8350-42ff-ae4e-2dfaa8007e4d · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ccb065e-fc53-4243-8d35-1912f4ee181b · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Seed1.5-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0e9fd3db-5621-447e-a107-d9caf1fe1d66 · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Improved baselines with visual instruction tuning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6e7f1377-5f59-4b6e-a128-dc2373558a68 · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4721ba2-2861-46cd-984a-826ed8ccff76 · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cf0b588f-8e48-4056-aa29-6a671db1bf53 · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3b1c8059-5c00-4693-8df0-5593e2087e77 · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9afd1751-b2fb-4b5d-a371-f26283a6438e · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c7c9e0a2-5806-4133-bc6d-1ce5d4549367 · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c92bab24-d994-458e-8c1d-ba831700dccb · outbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Qwen2.5-VL with Same LLM ToenableafaircomparisonwithQwen2.5-VL,wetrainLLaVA-Onevision-1.5-3BbasedonQwen2.5- 3B-Instruct
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b917bfce-3735-45c4-80af-82bfe0998b31 · inbound
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ac081a8-7e28-4f8c-bf34-9c7875622ae9 · inbound
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5356b479-aeb8-46c3-bb09-b613b6c54de4 · inbound
$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56e1fd8-4266-4845-a1f8-4e3861645a2d · inbound
Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5716265c-27fa-4e24-a474-7d6ee936cdbc · inbound
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a194843-e605-457b-8ed1-d8df4eb8ec76 · inbound
Dual Latent Memory for Visual Multi-agent System LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349738ae-b01c-42f3-91c0-9874107e2033 · inbound
R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b4d90e-ff62-4713-ab13-0b5c884139de · inbound
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f3831f-f314-4016-b1fa-e21ad07c8f9b · inbound
How to Take a Memorable Picture? Empowering Users with Actionable Feedback LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1055f925-b328-461c-b4ae-ae9e7c4202ec · inbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed61e2d9-dd4d-4810-bf4c-937f630c789a · inbound
SCP: Spatial Causal Prediction in Video LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 127ea1cc-24a0-4cb0-bdb4-b7814eccc74c · inbound
What if? Emulative Simulation with World Models for Situated Reasoning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aba40587-7c77-41d3-ba35-b127dbd705b6 · inbound
SlowBA: An efficiency backdoor attack towards VLM-based GUI agents LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e00bf86-2b95-41c9-bcc9-462e83cb5da0 · inbound
HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1947a29a-bdd4-4bc0-a9d5-acd628fd37d9 · inbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3142730b-f9d2-408f-9a3f-7ecc346f62ca · inbound
DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5bdf1458-ab84-49fb-bfbe-bae8e742e76b · inbound
Peel neighborhoods LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e90120b-beb8-4475-95d0-04bff2add184 · inbound
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd311869-f0bd-4e8a-b264-8c4536981707 · inbound
Token Warping Helps MLLMs Look from Nearby Viewpoints LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 01e30d6b-ae9a-4cde-9452-8ffd76e762b7 · inbound
BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cfbfdf4f-b70c-4886-bfa7-560635bcffe1 · inbound
Steering the Verifiability of Multimodal AI Hallucinations LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85be7987-200e-4b19-a14a-0db85f1bb29d · inbound
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 787b359f-03c1-4185-a051-bdb401ced8dc · inbound
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5016fc86-0c22-4adc-9087-09269cfb3ac9 · inbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7e4d366d-480f-497a-a5ed-b996f6d4fb22 · inbound
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 108cb2b1-22d1-4f94-adda-44c29233d8dc · inbound
Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 703b41d5-9773-42e5-91b4-932b3c967f8a · inbound
Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5feff337-67bd-4dd0-b6f4-d500cb2c9ad0 · inbound
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8c4c3159-2089-4c1f-8f42-2019eabc6b59 · inbound
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a4065dd2-6766-4ff8-8f39-da886683f0ab · inbound
Boosting Visual Instruction Tuning with Self-Supervised Guidance LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d3f719bc-52d5-461c-a042-fceb0656903c · inbound
PersonaVLM: Long-Term Personalized Multimodal LLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 63e716c8-fc00-47bc-a5af-85a4da609d5f · inbound
SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea5dc257-7758-4616-aec3-430dfd6e0574 · inbound
SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ba838f1f-db2c-41f3-9f83-335e075b4a11 · inbound
Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99182cbd-e134-48a7-92f6-eb31c28a6a36 · inbound
VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d1212e51-d4f0-4391-be2d-9181e458135d · inbound
BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1d8947d4-f66f-4005-ada7-56ffeb0eaded · inbound
Towards Temporal Compositional Reasoning in Long-Form Sports Videos LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14176f09-46b7-4be3-8289-cc7cf2de2db3 · inbound
SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8da5f080-67dc-441f-965e-cb07b768a296 · inbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aa21ed8b-d792-4f37-8981-a01aba58df48 · inbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c4a7cbcb-c5c9-4b51-b8f7-42da4fcf56c7 · inbound
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d57b3385-acb7-4ee0-a523-0827cb857b0a · inbound
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bcc3743c-f4d9-4dd8-b78e-ea3c3e13b86f · inbound
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3f00c458-fd10-4d3a-8806-2a235c9b2cb5 · inbound
GameScope: A Multi-Attribute, Multi-Codec Benchmark Dataset for Gaming Video Quality Assessment LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6d95d961-fc09-4be7-b773-9da99db45b7f · inbound
SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA? LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9aa22368-ea91-4392-ab56-2a7fdc87daa5 · inbound
Causal Probing for Internal Visual Representations in Multimodal Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8597090d-1cd3-4f16-bf27-eea09536c982 · inbound
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 54b888d4-6da8-46bb-85a6-2b2ad46d75a3 · inbound
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d1c355a4-e3f6-43a2-a814-e2cc36b37582 · inbound
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 37d812f7-ce55-421c-85e6-4b13a27bb05e · inbound
ZAYA1-VL-8B Technical Report LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eb7ac702-33ad-4a53-b580-b7dce4cd9dbb · inbound
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73b8cc84-5552-42c3-b5cb-f0bae577c1d0 · inbound
Sens-VisualNews: A Benchmark Dataset for Sensational Image Detection LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e948a9e3-9bbc-4c6f-a842-538d8c5569fe · inbound
LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5383bb70-7c6e-4e24-b80a-001e835af3c2 · inbound
Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 45c912ff-38df-4f70-831e-63ef8f6060cd · inbound
Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d5bfda9-6dea-41ac-bdf6-41c6312e63a4 · inbound
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a476c70a-76b1-4475-aff1-ec1cdfa79336 · inbound
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0259e5b1-bb7b-419c-9120-6bdd82715438 · inbound
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5933b99b-2729-44a8-a312-490afc1eac6a · inbound
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6872e68a-820a-4737-9509-552ffa257b09 · inbound
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 28594ffb-d7c4-4999-b4ed-626fb5da731e · inbound
INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dd4052d0-c5d6-47fe-86fa-3139ad344cc1 · inbound
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1081bbc3-cc39-4d01-bf51-4457b6d742a4 · inbound
RISE: Reliable Improvement in Self-Evolving Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d72db277-76e2-45dd-9276-881537fc991a · inbound
RISE: Reliable Improvement in Self-Evolving Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cfe8a004-722e-446c-8021-a22c89f909e7 · inbound
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3cb2db30-8ff8-449d-a5ed-8a34bfb4b918 · inbound
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 922854f8-c61e-4047-b37f-f0bbb1a97a98 · inbound
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8b00eee8-c5f7-44fb-8807-5c5d731ff68e · inbound
MetaphorVU: Towards Metaphorical Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea9195d4-8be6-4066-ac1e-c19e6997d5b4 · inbound
Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 581aa147-7652-47d4-b5e1-341d2bfcc9db · inbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 22c7d3f0-56c9-4a51-b678-7d0ec76415b3 · inbound
HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation da34d152-e821-4a29-8496-35baa38ce090 · inbound
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e6283773-defb-4958-89f7-07a6107e90ee · inbound
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation df923450-20fe-43a9-a62d-1c3b979d3bdb · inbound
Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9616b357-95b2-49d2-a8fe-551a5dff4033 · inbound
Zamba2-VL Technical Report LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d9117f3e-77ae-447c-b39f-e4aa0fa140cc · inbound
InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9e8a6090-5e48-4629-a5d8-ae8f01b80629 · inbound
Visual Instruction Tuning Aligns Modalities through Abstraction LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b35417bc-f33f-49ef-a6d3-0349718ddb53 · inbound
Benchmark Everything Everywhere All at Once LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 67bf9386-ea11-4b0f-99dd-e994e8c4f025 · inbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a41ade6a-6a63-4175-b91c-519e332121d4 · inbound
TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b6231428-e782-496f-be68-21e46aa0112f · inbound
Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c7231c6-d17f-4431-801e-68e93c8b7e2e · inbound
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e1095336-8700-4016-ad38-021264c5b1fd · inbound
Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9fb5d31f-d23d-43a2-885a-1a023423e044 · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 215
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8dc02e2-c7d2-4d29-b1fa-b908df80c43a · inbound
EventDrive: Event Cameras for Vision-Language Driving Intelligence LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4b6bb2d1-56cb-45d7-8363-9822f964601b · inbound
PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ba74f4e7-5a76-40a1-8d61-6fa971326ead · inbound
PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2ae1fd90-c27c-4c8d-9da8-c5abf5a1c386 · inbound
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 60e186ea-3219-4e3d-9880-16036e32a177 · inbound
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b6c3952d-1f89-4698-ae5e-b28ecc9cb72f · inbound
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 69465d34-1e91-4d0f-8be4-fcbe807e7c05 · inbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0dcb5745-6d38-4870-8e3d-d553d0c16a24 · inbound
Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bb3aa63-3f91-48ce-a715-d44acc1b30da · inbound
The Hidden Evolution of Disguised Visual Context inside the VLM LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 32b7ad4e-1322-4a12-922c-a467a5910ea3 · inbound
MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2380fd62-8a0a-4ffd-96a2-a0a8c2e04b39 · inbound
PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7a49d3f0-710f-4c52-bec9-feceaa80b68e · inbound
SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5d0d744-97d9-4ced-831d-bf813336622c · inbound
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f250ae5c-09ae-45ab-89a6-26bf685e844d · inbound
DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 561bd4ba-7118-41f8-9697-05d8be3b0324 · inbound
Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 67e520c4-024a-4f67-a238-b39f9699f33a · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.