Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T06:01:53.730356Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 100 inbound Pith citation observations for arXiv:2407.07895.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T06:01:53.730356Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:31:15.965901Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
69 of 69 outbound references displayed
External citation measurements
23
pith, observed 2026-08-05T02:28:24.338817Z
Observation c1a6b6b4-7f24-47ae-86c6-5852cf912c73 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b3738015-9edb-4629-99e0-db936c2425cd · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d201b793-59e1-42dc-ab82-3e001f933993 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Scanqa: 3d question answering for spatial scene understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b68a006c-9eb8-4538-948d-4afe93788c11 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Qwen Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation af82c23b-930a-4fa9-91e7-f1d67e9fa388 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61e5f039-a994-4097-8823-244f763dfadd · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Visual question answering on image sets
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 010363d5-1c73-4422-9399-4cab4a8f1a34 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models VideoLLM: Modeling Video Sequence with Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 94fa5daa-c824-44f3-8657-ed84d15d231e · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fd569198-e3d3-4481-bf67-08d9ed3c1e28 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2490bbe3-8e77-4a4e-ae5a-95e5937f0c81 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7be3e54b-3137-42fd-b0d5-754a9281ea62 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 63eae5ec-f086-4854-81d4-29a148cdcf8b · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 946005f6-d8df-44a3-b154-20715311ed17 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Sciverse
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 080577c8-4f50-4d7f-a6ab-4f3b2ac2b91e · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 67017df5-c408-44fb-9adc-c9bdeaafe641 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models ImageBind-LLM: Multi-modality Instruction Tuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6e64181b-a782-4825-9d7e-2cb6cb48973b · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 3d-llm: Injecting the 3d world into large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6fc66ded-ac69-4318-943e-02592dd933e5 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 3d-llm: Injecting the 3d world into large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a9da7cef-728d-47e9-b72c-d75717c57480 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Language Is Not All You Need: Aligning Perception with Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 92774a09-2243-462e-8c7c-4b0a301e10c5 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c324eeda-50c0-4d0a-a6d4-63947787e77b · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Many-Shot In-Context Learning in Multimodal Foundation Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7d6e01e5-b49e-439e-8a6b-de3456d2c396 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models ReMI: A Dataset for Reasoning with Multiple Images
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 37828504-eb27-4e88-b5c2-83a63acadd62 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Obelics: An open web-scale filtered dataset of interleaved image-text documents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 535bc872-3868-499a-95b7-224b248aca5d · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models What matters when building vision-language models?
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 11bbf476-21f5-4f49-b98f-66392cb90281 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 448207a4-f2cb-42c0-9662-5afe17eeae6d · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a52d7edc-059e-43f4-b9f3-2127a1a9817d · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and genera- tion
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a31826b8-5e66-4b5c-8384-49bc1a92bb17 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Fine-tuning multimodal llms to follow zero-shot demonstrative in- structions
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d3fc334e-a87d-4a13-806c-c051054beee2 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Fine-tuning multimodal llms to follow zero-shot demonstrative in- structions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 00f85e25-659b-4493-a6d5-69c64df88929 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models VideoChat: Chat-Centric Video Understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d4ff821-c132-445e-b69a-40f3ee6a4f5d · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 075a7d8a-4684-4e04-b3e0-ef9b26d2067e · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b5d0cc22-b356-4bb9-b792-ea6e52fb3f52 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Video-llava: Learning united visual representation by alignment before projection
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93790b17-d7d6-4ad5-a90f-3c10ec4d9aeb · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Vila: On pre- training for visual language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6b8ac181-2c73-4433-8850-b36323159e1e · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation be1b695a-bf15-4dc1-8f4a-385e1bdf2ab3 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Improved baselines with visual instruction tun- ing
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4b10f6c3-a30e-4f95-803f-6f382f4d8529 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Llava- next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 714843e4-6595-42c8-9ba0-6dfd915c9eb7 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Visual instruction tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a1797f7e-b858-4f8c-aaae-45251b6306fa · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Vista-llama: Reliable video narra- tor via equal distance to visual tokens
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 20a145ad-5579-41ea-9f09-31b4f99eb4c9 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Video-chatgpt: Towards detailed video understanding via large vision and lan- guage models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f2bd38de-4183-4f3c-9720-d52901c0cb0e · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Video-chatgpt: Towards detailed video understanding via large vision and lan- guage models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fda17c4-bc48-4651-9e3e-e1407b6fa0ae · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 631dfacf-4839-4e6e-8a61-3108361a6db9 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Gpt-4 technical report
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d676ccfd-b323-4818-8343-8ba647e835c0 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models GPT-4V(ision) system card
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bcc57d77-533f-44e5-ba13-1e1ea51e6e6d · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models DINOv2: Learning Robust Visual Features without Supervision
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 214c4f86-17d6-4c41-8a3f-ef495a50e597 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Learning transferable visual models from natural language supervision
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c977c08-f0f4-4f82-afc6-a88f86a9a316 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9ae410d6-2a09-4a2e-b1ef-b93d2b6cf2c3 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7ad6d14b-7f6f-40f8-bc6a-811e9ac20a07 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca1fe776-dd21-47ac-9020-1be100527f61 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 24010fe4-7d8f-4485-b2a2-5b10a086e7f7 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models A Corpus for Reasoning About Natural Language Grounded in Photographs
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 18743191-6eb9-4521-833d-2cf5e4cfd3f0 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Generative multi- modal models are in-context learners
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99298421-d2ef-4aaf-bfab-965b5835dc05 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LLaMA: Open and Efficient Foundation Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d09921d-82c0-4f35-977c-bc88e19bc222 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 40bb0bb5-ad2a-4d50-b8b8-2c999dbfe2b8 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c6edefc-924a-4814-96bb-d0ab6ccd99b4 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models STAR: A Benchmark for Situated Reasoning in Real-World Videos
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation baba10dc-c7e9-4a1e-9c9f-1e755cfba758 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 205642d2-b085-413a-b0b3-a6978326b633 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Next-qa: Next phase of question-answering to explaining temporal actions
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 84121b69-b692-4d17-b069-7a39544b3a85 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models PointLLM: Empowering Large Language Models to Understand Point Clouds
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8c17b0ad-8401-41dc-8cbc-5d83bee6f3d7 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 25010481-2a22-4df9-a3e1-949d351b8806 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea90a5a7-a265-4bcd-9651-3cbf7ed84f7d · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T17:38:12.261029+00:00.
Observation 8c7e0da2-81d6-4681-8e2b-2394d3ec103b · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Sigmoid loss for language im- age pre-training
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2be66cc2-d484-4cb3-b4ac-8e26c01ad504 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 43569bcb-7cd3-4831-8cd6-0334e6dd34c7 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LLaMA-adapter: Efficient fine-tuning of large language models with zero-initialized attention
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9a459fd1-0b51-458c-86c6-7c745aa4bb3b · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4a21b82d-8f57-4c5d-8782-6593f43566e9 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a689cc28-8bed-4fd2-a628-1b8e447b91fc · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Llava-next: A strong zero-shot video under- standing model, April 2024
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d8a579f9-681b-4f3e-847f-f2848bf53902 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Multimodal c4: An open, billion-scale corpus of images interleaved with text
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fdf3656d-c70d-46c6-ac01-20c2bc077077 · outbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 211d1d65-7d4e-43bd-85b2-2e4ac21c0a52 · inbound
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 615da092-a5ce-4f55-a172-38ac1fb26e79 · inbound
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16ad22ba-2820-4a99-9532-100dd81de35f · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 220
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0a9b0b24-ddd6-4fe9-8a77-b6c2c21a5797 · inbound
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6ce228db-5c85-4358-ac94-68ac0dde6c16 · inbound
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1b9d8dba-f3b9-4110-a689-ec02bb49593f · inbound
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9512c314-7c3b-47c8-b3c5-fdfe3c0b33bc · inbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca3abdb6-e5d6-46ab-98dc-6f98ad8a3a1b · inbound
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a208d090-c73e-48b8-82a5-cf931a563881 · inbound
VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82223398-adaf-490f-b246-cc2e9b3989d8 · inbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f7f28594-1b24-47e6-a3d8-0760c86a7a36 · inbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cef14f9f-c31c-4f57-81e0-bf0972528647 · inbound
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e5b1fdb-994f-432a-b989-b8dc7c3bf9cc · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 941e3834-430a-4454-93d8-e918e571fb39 · inbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b9fbd42-d719-43f2-948d-1c458390340b · inbound
ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e26760f4-3fcf-4b4e-ac56-e16b76841d14 · inbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4652513e-de24-4f2b-9611-4b62980f1b79 · inbound
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 700fca19-cb51-4cbe-809b-15818d86dd94 · inbound
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fabd24d6-5e9c-487c-ab18-8415e1e5138b · inbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9388e9f1-7733-45f2-a9a0-4616af00ec98 · inbound
Histopathology Multi-modal Embedding for Pathology Composed Retrieval LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4bd1b1-48f3-42f9-a787-d0fba4c35e45 · inbound
Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eef3c2e-3460-4598-bd12-47a07501e139 · inbound
3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c000dd-59bc-471b-992b-76d91a23abb6 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 276
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 00cbfc93-433b-4d54-8a6d-d9dc03f3d662 · inbound
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e66a46c0-3c91-4517-ac37-8e63025eff47 · inbound
Seed1.5-VL Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3341a722-d566-4051-a973-097073331774 · inbound
Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12cec2a7-abdf-4484-a537-a7d49576ec7a · inbound
Visual Agentic Reinforcement Fine-Tuning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf32c404-8e32-4da5-8a41-e18106c4adc5 · inbound
Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113e3bc5-eadf-4c37-92bb-29c545e23c57 · inbound
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e666ddb-be36-425a-86fd-cf359b6216de · inbound
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5cf6810-da43-4cd8-81d0-acd914e552a2 · inbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d07eb18d-c68a-4dc3-984f-9d24bbbf74c5 · inbound
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 981ae193-7e67-442a-9c3b-1bf192d17e9e · inbound
Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f8005a3-abc0-497d-a663-8d9b958afd0b · inbound
TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f855ee83-aae6-49e2-8691-74e7a8b72270 · inbound
Large Language Models for Planning: A Comprehensive and Systematic Survey LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 131
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c500152-2066-4fa1-bda7-fbb376034dcc · inbound
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8485d23b-57a7-4c04-a5a4-44bfb5ed4672 · inbound
Beyond Completion: A Foundation Model for General Knowledge Graph Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7661d267-206a-4624-a7d2-ad635d018b9d · inbound
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 481d3352-d8e1-4c05-97e5-2b773533b625 · inbound
Fostering Video Reasoning via Next-Event Prediction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a96e7a8-45c8-4de3-8a5a-1a800d78b3af · inbound
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58b2d41-789f-4fd5-b5c6-b70b91274d4e · inbound
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ddb09f-9ca1-4b1d-aa5a-6e3e07673ac3 · inbound
DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 344aad25-01f9-4cb8-9749-f1ef562dcc3f · inbound
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396755af-003a-4a95-979d-1e876c601d0a · inbound
SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b379fc0-4ef3-45e3-93b6-51201ac64726 · inbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba1f3937-444b-4ad9-b686-4a5cf4401404 · inbound
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 045796f1-b877-4933-81e6-b55437019dbc · inbound
AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12899f82-3563-44a6-9a78-941a20b1afd3 · inbound
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec1934d-55f9-4823-b709-84fa00b75b98 · inbound
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa1eb206-ca7b-4363-b8fd-77b9c0e6ef69 · inbound
PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9cd6f8d-803e-4d7d-9736-8787ab4efb4a · inbound
Demystifying the Visual Quality Paradox in Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da717749-599d-4916-9ecb-f93dbace5d54 · inbound
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3c7d626-021b-4630-ba82-b41255c2a330 · inbound
Semantic Caching for Improving Web Affordability LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0304ec94-346f-42ef-92bb-ffa955e7c5e6 · inbound
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c4469d-25a7-4db2-ac88-4ab13f38b3a3 · inbound
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491e6990-29ec-4934-a82d-bdedd27f0979 · inbound
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 004e1b95-1576-4438-bb11-86d5d535fae5 · inbound
Room Scene Discovery and Grouping in Unstructured Vacation Rental Image Collections LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b25b168-8e05-4ae4-9e0d-dca0b148fd84 · inbound
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e21e9a31-6b36-4a6f-90ac-6cc483401ae9 · inbound
Kwai Keye-VL Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a86a8290-85a0-4b46-9027-cb17eeb6047d · inbound
Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34958276-5601-46d2-a9fe-fc054773bdba · inbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ad19dd6-f942-4ace-90d9-dab0c91f6bc5 · inbound
From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f989038e-d8a4-4f79-a1a3-12bf90ba02b3 · inbound
Pedestrian Intention Prediction via Vision-Language Foundation Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ff6505-3b8c-489a-93fc-f5953c148163 · inbound
FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4538fcbd-2ec2-4ba8-94e9-b6a0e76935b6 · inbound
LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b730cbda-5bef-43f6-8798-4771a4d6016f · inbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41f92863-7248-4f98-84ce-187f0a7959ac · inbound
ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 025ca9d1-b964-4484-94ed-e1e54b55c98b · inbound
FaceLLM: A Multimodal Large Language Model for Face Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba75bd4-02e0-44dc-8bcf-3d33d8a96cdd · inbound
Describe Anything Model for Visual Question Answering on Text-rich Images LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0188582b-0aa0-412d-bdb3-cdcba7394859 · inbound
InterAct-Video: Reasoning-Rich Video QA for Urban Traffic LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c20f704-e223-4e5b-a652-0cc82321eb7e · inbound
GR-3 Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 50dfb0b0-7468-4715-a1fe-c5ddce93e998 · inbound
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7cad20f-21af-4df1-92e8-4210147ffecf · inbound
Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5af75d5-94d6-40be-ab56-232a5a907760 · inbound
EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede1ea7f-a126-498b-8996-33326c899b63 · inbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a316b7-1e19-4fc6-89ca-7f405d5bf9e2 · inbound
Volume-Distance-Ratio Asymptote and Spacetime Inextendibility for FLRW Spacetimes LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 910ab3a5-948e-4b7d-bcc9-3162ce11c2af · inbound
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7ef91280-3fa7-4007-aa6a-39f50dfe993e · inbound
Multimodal Video Emotion Recognition with Reliable Reasoning Priors LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92fe071e-0cef-4f1d-a299-84da41f4e7b0 · inbound
AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated Content LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 177dc122-8273-4c03-870e-e8cf638b2e03 · inbound
Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3899d92-19a9-4b3f-b1c8-8bab760c7cf2 · inbound
IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d9f69a4-1136-4732-8936-b794417cb6bd · inbound
A Survey on Video Temporal Grounding with Multimodal Large Language Model LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec322ad7-3d90-4a70-ac13-c2e78f070e97 · inbound
Region-Level Context-Aware Multimodal Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eaaec05-f53b-43aa-9ac0-424881140064 · inbound
AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15ecf628-7c09-4a99-a955-bc2e90e71ac8 · inbound
Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8c289cd-b6e4-4188-bf6e-c8e9db1b9195 · inbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f42c18c-c8e8-4144-8a66-04453bac6a48 · inbound
Ego-centric Predictive Model Conditioned on Hand Trajectories LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4097d819-b341-41b9-892a-7aae4935978f · inbound
CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b4468c9a-0187-4eb9-abe4-590c55e5e6a5 · inbound
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab5c75e-59a5-4bb6-8b05-4d7047b3d975 · inbound
Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6e57d26b-0ecc-47ba-9c55-5ba61a521cf0 · inbound
InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8ae2c0-cb60-44fb-9469-3644af3ce3ea · inbound
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a56dd022-f6b6-4b89-b253-71bc8c61c996 · inbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 107633b8-832d-4086-b8c4-cb1733ce28b7 · inbound
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b54e816c-2572-4492-be4c-60f68ca312e9 · inbound
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a20b2645-8f0e-420d-8237-4c5bf9beeb28 · inbound
Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b3603a2e-65cf-45c4-8f50-e6e53fa82e62 · inbound
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c75648d-0ef4-49e2-b309-a9a3f18588ce · inbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c46fbf43-cad5-4e06-b956-1bd23df14fe7 · inbound
Are vision-language models ready to zero-shot replace supervised classification models in agriculture? LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c3be2a57-c51f-4003-8361-00e2d681c6c0 · inbound
BrepLLM: Enabling Large Language Models to Understand Boundary Representations LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.