Pith. sign in

Paper Citation Record · LEDGER

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

As of 11 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 100 inbound Pith citation observations for arXiv:2407.07895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.07895 v2

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T06:01:53.730356Z

measured 169 of 169 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 215 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:31:15.965901Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact31
  • verified fuzzy33
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

23
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c1a6b6b4-7f24-47ae-86c6-5852cf912c73 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.250547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:1769144c7d0451b6bc235fe085fe4a8567bd29d50a40790e0f785e0e89856348

Observation b3738015-9edb-4629-99e0-db936c2425cd · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.553388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:9a4b28aa04df0ece41a5d9a13bc5a58e43119ee2bd539a1415e08bc23e11af54

Observation d201b793-59e1-42dc-ab82-3e001f933993 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.111010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:dcb51fa0c9becee65484a54a02063951d645230f0f4be0587de4eed772a58a57

Observation b68a006c-9eb8-4538-948d-4afe93788c11 · outbound

This paper cites Qwen Technical Report.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Qwen Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:01:54.066484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:d0f0105e53e40e2715c85b6ebad5185dabfd3e64b2820d9e11a93e22e12e73d5

Observation af82c23b-930a-4fa9-91e7-f1d67e9fa388 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:01:53.811519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:5318de6ed14a675cda5ec4b1a63643c9975a2b68a0db3b40b4c165939a9d32d9

Observation 61e5f039-a994-4097-8823-244f763dfadd · outbound

This paper cites Visual question answering on image sets.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Visual question answering on image sets

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.118345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:c9d04957df79c1ba59951e8c9b70534963f85fc9029945717aacee7783acd622

Observation 010363d5-1c73-4422-9399-4cab4a8f1a34 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models VideoLLM: Modeling Video Sequence with Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:53.863823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:3df35c7d00814aa35d8fed6f36c6e90358ef7fb70ba4cb882621600480754222

Observation 94fa5daa-c824-44f3-8657-ed84d15d231e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:01:53.907536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:adba88e0ac821663825a859870811987ecff609499525e050dbf076d8721b16a

Observation fd569198-e3d3-4481-bf67-08d9ed3c1e28 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:01:53.941152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:46b6c8b38aa2ae48d91416ae632d3f542a0c93aea0f85ba05ce366f3cc8cca59

Observation 2490bbe3-8e77-4a4e-ae5a-95e5937f0c81 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:856121cb6d19ad2196dda16b7a9c1c9b6188e763253fcd34b7b1de2431c31ae5

Observation 7be3e54b-3137-42fd-b0d5-754a9281ea62 · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.007023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:7234a8070cf50244ee30b9bfe9d556a994304d9bd0a5010a9c6a840c17bf3856

Observation 63eae5ec-f086-4854-81d4-29a148cdcf8b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:01:54.028270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:72b7ee6e586dd3adce273c09d7fa55f4c31dbcaaf4ea882f32f72bb9b897b77a

Observation 946005f6-d8df-44a3-b154-20715311ed17 · outbound

This paper cites Sciverse.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Sciverse

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.129149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:14e9d8dbf598a2c2df07f8538586dbd4ef4e1f3e534a8c6ad520e5c8e9f38020

Observation 080577c8-4f50-4d7f-a6ab-4f3b2ac2b91e · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.075378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:2873b8687748e5c059bb7a630983bd80fdb9b4d3bc764f3365c7175378752070

Observation 67017df5-c408-44fb-9adc-c9bdeaafe641 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models ImageBind-LLM: Multi-modality Instruction Tuning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.096749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:48f88faeed9bac7f38b27073ec5d3e737b0fb52293260abab9a269eb72906089

Observation 6e64181b-a782-4825-9d7e-2cb6cb48973b · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 3d-llm: Injecting the 3d world into large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.133170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:4471df3a7bfd1cada35f16f271fb5bf718f971f7ee2e15ba2bba5ecf8ee37b41

Observation 6fc66ded-ac69-4318-943e-02592dd933e5 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 3d-llm: Injecting the 3d world into large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.137273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:6357557a2bbcb67ccce0347fd1304b722248bd1b319a5cd01b5eee0cc5896d98

Observation a9da7cef-728d-47e9-b72c-d75717c57480 · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Language Is Not All You Need: Aligning Perception with Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:32:23.026112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:36df6890b0b50584c05c726dfa29b9db330a18af24ce4445d3cc48a0fdb27f35

Observation 92774a09-2243-462e-8c7c-4b0a301e10c5 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:01:53.834439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:6fa8e9f20692ad28328bc4afb9971ed4c72dda7187b7d81244271d448cbc2e33

Observation c324eeda-50c0-4d0a-a6d4-63947787e77b · outbound

This paper cites Many-Shot In-Context Learning in Multimodal Foundation Models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:01:53.841270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:b30cca21867ef693283b690b5038dcdbc8bc6a11f8da5866f300c94f1b4aa0eb

Observation 7d6e01e5-b49e-439e-8a6b-de3456d2c396 · outbound

This paper cites ReMI: A Dataset for Reasoning with Multiple Images.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models ReMI: A Dataset for Reasoning with Multiple Images

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:01:53.858154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:d9c2445c41ef74f070ff3531390a9747abc4160230b76f39e1e4372ed893214f

Observation 37828504-eb27-4e88-b5c2-83a63acadd62 · outbound

This paper cites Obelics: An open web-scale filtered dataset of interleaved image-text documents.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Obelics: An open web-scale filtered dataset of interleaved image-text documents

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.144329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:9adcbdd0469bf41e30f8b0e3b926fa907a0d02b91035910a00622cffc4cc782e

Observation 535bc872-3868-499a-95b7-224b248aca5d · outbound

This paper cites What matters when building vision-language models?.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models What matters when building vision-language models?

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:53.879662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:cb8077b790bba443f6f6bd8233b3b96ce0d7b41773e07bdabef1bf4802751aa5

Observation 11bbf476-21f5-4f49-b98f-66392cb90281 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.156857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:f9741e2ead718e01250a96ce8b5ebb308d86eb63b5aa1b22e19c8d11e8174e19

Observation 448207a4-f2cb-42c0-9662-5afe17eeae6d · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:53.915342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:c7c7b6e1cb1cf92fcb1c6e1102ffb74906145e7efdd3b8899e5df842fe8a8183

Observation a52d7edc-059e-43f4-b9f3-2127a1a9817d · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and genera- tion.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and genera- tion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.164224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:7ae156c00058ee39b80349eb452e93a267e03d8847192c115dd458d8ee3f1c62

Observation a31826b8-5e66-4b5c-8384-49bc1a92bb17 · outbound

This paper cites Fine-tuning multimodal llms to follow zero-shot demonstrative in- structions.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Fine-tuning multimodal llms to follow zero-shot demonstrative in- structions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.169809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:b35bd11358d590cf0a55c027236747ed28642b8c308f06a86989e37c17b29e13

Observation d3fc334e-a87d-4a13-806c-c051054beee2 · outbound

This paper cites Fine-tuning multimodal llms to follow zero-shot demonstrative in- structions.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Fine-tuning multimodal llms to follow zero-shot demonstrative in- structions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.177720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:dcbbd13527c943efce0c0d613791188b68ecec98592f4bafa308cc03f5a7d812

Observation 00f85e25-659b-4493-a6d5-69c64df88929 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models VideoChat: Chat-Centric Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:3c691513d06178abca4e21ed1b124c72b371988e448f2371314ceeaf0565eadb

Observation 5d4ff821-c132-445e-b69a-40f3ee6a4f5d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.182237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:f6d71e17aa8e14b9b6b1ed5ff53151faa3a68307d6663846c5ea82a761115550

Observation 075a7d8a-4684-4e04-b3e0-ef9b26d2067e · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.037741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:81d85198eab3ac3c83da88da1d595ddee8361d6e04486e88ef89a4d399680b34

Observation b5d0cc22-b356-4bb9-b792-ea6e52fb3f52 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Video-llava: Learning united visual representation by alignment before projection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.185673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:4e537edaaa6f938fd4b329110c0c3c2ee0e627754f763776b594842204ae8ee8

Observation 93790b17-d7d6-4ad5-a90f-3c10ec4d9aeb · outbound

This paper cites Vila: On pre- training for visual language models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Vila: On pre- training for visual language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.193923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:a18077febfebfe30357a4e7c4b8bb2cc611763c454c1185bdb4e4e88a37a8549

Observation 6b8ac181-2c73-4433-8850-b36323159e1e · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:03:27.047113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:a091812cf49b82dafdf9a819f3fcf136edec5b0440c4cd5f828f5445b9639442

Observation be1b695a-bf15-4dc1-8f4a-385e1bdf2ab3 · outbound

This paper cites Improved baselines with visual instruction tun- ing.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Improved baselines with visual instruction tun- ing

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.197203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:bc7d667a36b9d6a2acce5e7c55c29935dfcf7ed648f4e1c720a0984492bc3778

Observation 4b10f6c3-a30e-4f95-803f-6f382f4d8529 · outbound

This paper cites Llava- next: Improved reasoning, ocr, and world knowledge, January 2024.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Llava- next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.200045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:434c928a141441cfbd4f532daa9760d571246ee1519d2d3d2af5423c97283492

Observation 714843e4-6595-42c8-9ba0-6dfd915c9eb7 · outbound

This paper cites Visual instruction tuning.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Visual instruction tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.225950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:6e80c3825477f3daf0a3bde0e6b42f5a9a38ddefa0f7af24df3dfe765a291b3e

Observation a1797f7e-b858-4f8c-aaae-45251b6306fa · outbound

This paper cites Vista-llama: Reliable video narra- tor via equal distance to visual tokens.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Vista-llama: Reliable video narra- tor via equal distance to visual tokens

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.230003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:ff76a4d75cf29072e6611d45627971fa949e6510e613e3f5526523727594cb7d

Observation 20a145ad-5579-41ea-9f09-31b4f99eb4c9 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and lan- guage models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Video-chatgpt: Towards detailed video understanding via large vision and lan- guage models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.235050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:d74a9607eaa86215bd9cfedde74a3693bf0240ce6dd2ced81bf08c4726684746

Observation f2bd38de-4183-4f3c-9720-d52901c0cb0e · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and lan- guage models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Video-chatgpt: Towards detailed video understanding via large vision and lan- guage models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.241181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:3b04dd67ae2cd4a8b5a9f445d5cb87f5fe5778673c9d738805b66902c5f68000

Observation 1fda17c4-bc48-4651-9e3e-e1407b6fa0ae · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:928542ad982439c6b30680438a359d2c6ff4568fe90ef1a52817486fb1d7e0bd

Observation 631dfacf-4839-4e6e-8a61-3108361a6db9 · outbound

This paper cites Gpt-4 technical report.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Gpt-4 technical report

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.244545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:f3059680a40fffbac4f602d2eb529ac3377a214cebd12a0a6cedc9cc85233595

Observation d676ccfd-b323-4818-8343-8ba647e835c0 · outbound

This paper cites GPT-4V(ision) system card.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models GPT-4V(ision) system card

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.106869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:12a3aaefa61302aa1483eb8618778f434aeece1edc8ba35f6a37cb7553453351

Observation bcc57d77-533f-44e5-ba13-1e1ea51e6e6d · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models DINOv2: Learning Robust Visual Features without Supervision

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:01:53.869766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:da9aa30544c6dba2b293f667b47aba1896b4ddbc38a7f9c03548e00054d5f141

Observation 214c4f86-17d6-4c41-8a3f-ef495a50e597 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Learning transferable visual models from natural language supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.256195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:19cdd66eaeafe330d8dc3907945886777b9a5b621960355d99c3f7e1a2aa2f63

Observation 9c977c08-f0f4-4f82-afc6-a88f86a9a316 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:01.130308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:bc23143148069211a0b44756dcd12c3273cefaf0e8e043741caf8356bda9bd60

Observation 9ae410d6-2a09-4a2e-b1ef-b93d2b6cf2c3 · outbound

This paper cites Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.295762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:c8be5304837fd3ac8a60a7f9940847158f8f6f5d3414f8bb30a7ee65d0fa7968

Observation 7ad6d14b-7f6f-40f8-bc6a-811e9ac20a07 · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.407700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:19beb32461001163da5d3036d0b27ccfc6a40c48732b558f5e138d4d9c1d903a

Observation ca1fe776-dd21-47ac-9020-1be100527f61 · outbound

This paper cites Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:53.925903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:a3aac1401766d4bc57f10ad0179795f352775f09d1c94367158c749849294a6e

Observation 24010fe4-7d8f-4485-b2a2-5b10a086e7f7 · outbound

This paper cites A Corpus for Reasoning About Natural Language Grounded in Photographs.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models A Corpus for Reasoning About Natural Language Grounded in Photographs

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:53.935043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:cbe3b162e0756786ac23da310d61f5ce110c25af7f528fc6996cf01a723e4209

Observation 18743191-6eb9-4521-833d-2cf5e4cfd3f0 · outbound

This paper cites Generative multi- modal models are in-context learners.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Generative multi- modal models are in-context learners

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.420983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:416257393090f49c219c9c0d6fe501e955a9330df672f305e3930377dd57605f

Observation 99298421-d2ef-4aaf-bfab-965b5835dc05 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:01:53.947344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:4104fcdb01fe52157e853fa2a1ced3a492197dea9bd3db6a2e232ee143885862

Observation 5d09921d-82c0-4f35-977c-bc88e19bc222 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:01:53.954672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:c7563aac1783f76058de07af12a5adfb37f53b5d7a5c613bf14bd8753e6fb549

Observation 40bb0bb5-ad2a-4d50-b8b8-2c999dbfe2b8 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:09:30.603679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:6c57bc321815740a7fea801d194f86b8de9e71608e3164225492a655e2ff52dd

Observation 2c6edefc-924a-4814-96bb-d0ab6ccd99b4 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:53.971264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:d65f9f6043304496a9ccf1231547737b2cc909fbc523cdc59eb4d71a0aea379f

Observation baba10dc-c7e9-4a1e-9c9f-1e755cfba758 · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:53.983850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:2bb809c338894b5935466fa95edc2da2e143e3b5b9ecf138318ba6451cfc3bc5

Observation 205642d2-b085-413a-b0b3-a6978326b633 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Next-qa: Next phase of question-answering to explaining temporal actions

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.453723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:0ebc22e1b897c805eb9ef95a2419d5b1b40ed73a77a8fce9168a4763b0728a31

Observation 84121b69-b692-4d17-b069-7a39544b3a85 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:01:54.002228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:016ad7118cc922596882034cd43908fe567e3c0ea917f40dc21e7c248985c03b

Observation 8c17b0ad-8401-41dc-8cbc-5d83bee6f3d7 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.536774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:6269688a561e5ca56593dfd52f064cee46dde84d4db5b146d74e321908137c0f

Observation 25010481-2a22-4df9-a3e1-949d351b8806 · outbound

This paper cites Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.544773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:2fad533bbc61f489d1381e96fcd1a6416f341acd7108a14367efd6ff27a3ade7

Observation ea90a5a7-a265-4bcd-9651-3cbf7ed84f7d · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:37:42.029646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T17:38:12.261029+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:450117234e98121a435501374768b76acc4cca08d420d034752722d1c1cca102

Observation 8c7e0da2-81d6-4681-8e2b-2394d3ec103b · outbound

This paper cites Sigmoid loss for language im- age pre-training.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Sigmoid loss for language im- age pre-training

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.551143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:76c7ad077de6b95520239d62c7aedc5e1b0d7dcb9df1c13de14dc973b3c49061

Observation 2be66cc2-d484-4cb3-b4ac-8e26c01ad504 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.032979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:689293719a4b669e6c02f04bf35a9b2be0f252b3d4160ff9e2b4ab9df24716e5

Observation 43569bcb-7cd3-4831-8cd6-0334e6dd34c7 · outbound

This paper cites LLaMA-adapter: Efficient fine-tuning of large language models with zero-initialized attention.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LLaMA-adapter: Efficient fine-tuning of large language models with zero-initialized attention

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.560615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:3d08b3e7d0cce09dbae266c69a9f3f0df4d08e8d1081d32bee981ab1fdc0d10c

Observation 9a459fd1-0b51-458c-86c6-7c745aa4bb3b · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:29:30.364997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:9d479132276c5866f35f4fea97ebd3f4e7c21e65ce194777d69426f9fe10eafd

Observation 4a21b82d-8f57-4c5d-8782-6593f43566e9 · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.061147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:1b1d6138b38395dd42dcfd27d09510a0eac74e29cb1fe7a9ab5797f5c0130a07

Observation a689cc28-8bed-4fd2-a628-1b8e447b91fc · outbound

This paper cites Llava-next: A strong zero-shot video under- standing model, April 2024.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Llava-next: A strong zero-shot video under- standing model, April 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.565998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:9240c6b920fbc96c1ef25439cf844cc5ebdd30dd3c3f3f45ddb0a3a65dd2355e

Observation d8a579f9-681b-4f3e-847f-f2848bf53902 · outbound

This paper cites Multimodal c4: An open, billion-scale corpus of images interleaved with text.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Multimodal c4: An open, billion-scale corpus of images interleaved with text

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T06:01:54.573443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:6cb6176556d7bc6e0ae24fd8c3178d29e8ae7ec98bd0657fe28f6e1932152955

Observation fdf3656d-c70d-46c6-ac01-20c2bc077077 · outbound

This paper cites an unresolved cited work.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Unresolved cited work

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-05-11T06:01:54.582702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:9634784f8257c33c9aa142587a282061dd20e1c09af767a7520c6cf4d9f3dc96

Pith citing papers

Observation 211d1d65-7d4e-43bd-85b2-2e4ac21c0a52 · inbound

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? cites this paper.

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:29:30.210789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T01:29:30.032408Z digest=sha256:d6874df05f1118121b3122c20f16e423b1d0fae1276e8e1e024a7697af17ca33

Observation 615da092-a5ce-4f55-a172-38ac1fb26e79 · inbound

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone cites this paper.

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.585341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:19:27.255515Z digest=sha256:141901e59edca02a52c15c9ee076016b8827fe34076745a8757e6e79de3b23d2

Observation 16ad22ba-2820-4a99-9532-100dd81de35f · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 220

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:20:36.416394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:17f02a2baec2d356baf405720432fb986cd885ad289a2a8e708598e9ba08f6e5

Observation 0a9b0b24-ddd6-4fe9-8a77-b6c2c21a5797 · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:51:48.357499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:a0ee6394c83335ee1067dd09244e5e1f93ab5009aff1446d5e1f05ce9d43807f

Observation 6ce228db-5c85-4358-ac94-68ac0dde6c16 · inbound

VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks cites this paper.

VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:19:43.983968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T21:19:43.882232Z digest=sha256:2deb448ded8965dd1a8561f0d9edfe6d00fa23568d48ef627c3bd95bf1c76a8e

Observation 1b9d8dba-f3b9-4110-a689-ec02bb49593f · inbound

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation cites this paper.

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:40:00.141940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T14:39:59.870039Z digest=sha256:ffc7f1b5b579192388e95caf3034d0a2ddf095ed958ef24b9da357c0af04b115

Observation 9512c314-7c3b-47c8-b3c5-fdfe3c0b33bc · inbound

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance cites this paper.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.639650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:f2053aa3c36d09f16087811719834095ccb4f67ae63df2bdaa9c4acb6004074b

Observation ca3abdb6-e5d6-46ab-98dc-6f98ad8a3a1b · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:09:23.805919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:1a34a9a87e009078c3779f0e5b3d65b8bfc8be61e44ce01cca5f6b0f7ebaf8cd

Observation a208d090-c73e-48b8-82a5-cf931a563881 · inbound

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation cites this paper.

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T11:49:14.727821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T11:49:14.698249Z digest=sha256:ea4410b381f7c10ba4761986d9a41ffbcc675c06240e8ef4a1857fef0f3111a8

Observation 82223398-adaf-490f-b246-cc2e9b3989d8 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.430481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:917f9f97734d3d340e2a194e89b55c233499e5a1d870b4ed023dee2911c8f1f7

Observation f7f28594-1b24-47e6-a3d8-0760c86a7a36 · inbound

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs cites this paper.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.965901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.965901Z digest=sha256:6543d89cc342d3cd5e8088b87fb2ea421608053cb2e31e34c67d98031f329b1b

Observation cef14f9f-c31c-4f57-81e0-bf0972528647 · inbound

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency cites this paper.

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:26:00.546599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:26:00.546599Z digest=sha256:6320c965e571782f427a847a24580397afc72b7fcf7a0a58a422e0080a93f402

Observation 8e5b1fdb-994f-432a-b989-b8dc7c3bf9cc · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:02:37.355001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:e559e0c3bcba8c660dec066051a0fb29e49b252ab909867a57f0e68ff3b45fba

Observation 941e3834-430a-4454-93d8-e918e571fb39 · inbound

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models cites this paper.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.578987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.578987Z digest=sha256:fd7b8f4b1a8e1efbc75f06612d7fad0aeefa3229324d5c11909c9f2e2bd79ee1

Observation 2b9fbd42-d719-43f2-948d-1c458390340b · inbound

ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation cites this paper.

ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:03:44.464754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:03:44.464754Z digest=sha256:e194d0f8fd96c498ac60c54f1c0b9184fbefb615c48eac6702be55c3b04de74f

Observation e26760f4-3fcf-4b4e-ac56-e16b76841d14 · inbound

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness cites this paper.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.031175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.031175Z digest=sha256:50a82959850d2908380d0b85ba29e4844ac4e4f1d088523cf3989389327d0767

Observation 4652513e-de24-4f2b-9611-4b62980f1b79 · inbound

When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis cites this paper.

When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:52.804312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:52.804312Z digest=sha256:4d27ffa54e9cc849a7e9a010dc69f8f1b3bcf604b9e1661b90f7d075db9c7a80

Observation 700fca19-cb51-4cbe-809b-15818d86dd94 · inbound

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models cites this paper.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.602049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.602049Z digest=sha256:a82658dc0770ea785b3daaffb99906a27bfc6da140a94f44a3221b3cb96e2db3

Observation fabd24d6-5e9c-487c-ab18-8415e1e5138b · inbound

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding cites this paper.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.236516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.236516Z digest=sha256:a95fcb4794896c0313d589c22d4f2b9ebfadec2dc93b2390602c5d978b74f6fa

Observation 9388e9f1-7733-45f2-a9a0-4616af00ec98 · inbound

Histopathology Multi-modal Embedding for Pathology Composed Retrieval cites this paper.

Histopathology Multi-modal Embedding for Pathology Composed Retrieval LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:30:12.974013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:30:12.974013Z digest=sha256:b600fde0948d7819d3db51cd59e5aae0470441d851bc879e2dc2b13802e1a5a9

Observation 4d4bd1b1-48f3-42f9-a787-d0fba4c35e45 · inbound

Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models cites this paper.

Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T12:15:53.477089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:15:53.477089Z digest=sha256:5abb4a3f95fa4a6718262f0d19943ac78cc5f8698c8e282e922479577afde5ac

Observation 5eef3c2e-3460-4598-bd12-47a07501e139 · inbound

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning cites this paper.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.811203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.811203Z digest=sha256:5a07884eeda00c4b36f685748b8918a8374ac0ce3d3366c64983ec3dba629d20

Observation 34c000dd-59bc-471b-992b-76d91a23abb6 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 276

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:18:53.614766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:2379a407bf7b048055b682ac72e81f0af7a7bd2eb406388c68d67ad7ad0715e8

Observation 00cbfc93-433b-4d54-8a6d-d9dc03f3d662 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:57:13.351076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:dd0b43825ea04cf11ff08ae290035444efc4bbaf5d91ea6883440be17bf53578

Observation e66a46c0-3c91-4517-ac37-8e63025eff47 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.585341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d4e10a43e0eb157704f474910f4b77d4675fe53f8cf2fb234dc0c83be6189ff4

Observation 3341a722-d566-4051-a973-097073331774 · inbound

Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering cites this paper.

Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.855615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:24.855615Z digest=sha256:be1c8d21fa3ca926462e5829d625820eb14dcd213567f2564f3e59d134b3ef4a

Observation 12cec2a7-abdf-4484-a537-a7d49576ec7a · inbound

Visual Agentic Reinforcement Fine-Tuning cites this paper.

Visual Agentic Reinforcement Fine-Tuning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.491865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:27.491865Z digest=sha256:67366e9fc49f91e8b1979f742c74494fa61eb4418d3f3f26f8e058a97e3e4bcb

Observation cf32c404-8e32-4da5-8a41-e18106c4adc5 · inbound

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency cites this paper.

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:04.786215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:04.786215Z digest=sha256:525994758691385826c6e0ffdef3306f2cf8179480f18d985d2b8edbb766cd6a

Observation 113e3bc5-eadf-4c37-92bb-29c545e23c57 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.242364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.242364Z digest=sha256:5f390dadfad2f06a3f47466be901f6771d46d7b2a15675cca3388b3dd85b5060

Observation 5e666ddb-be36-425a-86fd-cf359b6216de · inbound

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs cites this paper.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.075306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.075306Z digest=sha256:70638ce07141d4678e22145ecbe9faec34688a949ee1d7c9a290fe64978974f4

Observation f5cf6810-da43-4cd8-81d0-acd914e552a2 · inbound

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models cites this paper.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.857116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.857116Z digest=sha256:f74378b19781fa7f5a89fdf33703b520864f75f6f9d61fde7c0c3adf132d1f91

Observation d07eb18d-c68a-4dc3-984f-9d24bbbf74c5 · inbound

ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding cites this paper.

ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:55.063902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:55.063902Z digest=sha256:e063a66eb0ef6c5d1c8d3f06b5e62b98f85dc6ddb699d908a5c559544937bd22

Observation 981ae193-7e67-442a-9c3b-1bf192d17e9e · inbound

Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation cites this paper.

Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:12.320067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:12.320067Z digest=sha256:5c1083b39c8d44e8482753cfd4ab9fa752f3be18175e5705812fd2841580a647

Observation 1f8005a3-abc0-497d-a663-8d9b958afd0b · inbound

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs cites this paper.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.190049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.190049Z digest=sha256:a7499575a7e11a26a1a0ce1e6e613321bb9535c6ceae1707576f88d01ab7806e

Observation f855ee83-aae6-49e2-8691-74e7a8b72270 · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:59.735118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:59.735118Z digest=sha256:3892ea18361186c605d89b3adbf5d471598cc0c5c5f5a73ff7e0895b68b2e34b

Observation 2c500152-2066-4fa1-bda7-fbb376034dcc · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.801599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:00.801599Z digest=sha256:4970e80cd48a419925e6146770c262f3d141ca43cfcfac9f910deeb062bbbd13

Observation 8485d23b-57a7-4c04-a5a4-44bfb5ed4672 · inbound

Beyond Completion: A Foundation Model for General Knowledge Graph Reasoning cites this paper.

Beyond Completion: A Foundation Model for General Knowledge Graph Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:17.473519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:25:17.473519Z digest=sha256:dcba6835b225c8efcbd3345e0291e820397a046d34ca346316e3a010a6c9dd86

Observation 7661d267-206a-4624-a7d2-ad635d018b9d · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:10.915411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:10.915411Z digest=sha256:3bbd63d2f6c1cf6fb623f507f8bdb510374cd7e3b4deb74084bcafbb80a721b3

Observation 481d3352-d8e1-4c05-97e5-2b773533b625 · inbound

Fostering Video Reasoning via Next-Event Prediction cites this paper.

Fostering Video Reasoning via Next-Event Prediction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.429393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.429393Z digest=sha256:a883b5cabcda42454faa3ec2b99671b0b74b4c6598bc0326aeb70039c0825774

Observation 9a96e7a8-45c8-4de3-8a5a-1a800d78b3af · inbound

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models cites this paper.

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:00:48.885841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:00:48.885841Z digest=sha256:55e0088c868273731994103bb5f808c5d4507e1c21db30871b41f1d02feadc12

Observation a58b2d41-789f-4fd5-b5c6-b70b91274d4e · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:27.183905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:27.183905Z digest=sha256:1eb04e564280bf729a8302cf4190de9fb5020ec1bb55ec1693e2abc6656dc127

Observation 71ddb09f-9ca1-4b1d-aa5a-6e3e07673ac3 · inbound

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models cites this paper.

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:45.182318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:45.182318Z digest=sha256:079b232672c510e7353801362e162d8515f18b86b3d5869d06441471d9cd1b0f

Observation 344aad25-01f9-4cb8-9749-f1ef562dcc3f · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.706996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.706996Z digest=sha256:5291b45a1e1a021ab5a38eac2ca1ae7cee683e58861502fcfd9aa5af4fc3071b

Observation 396755af-003a-4a95-979d-1e876c601d0a · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.522356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.522356Z digest=sha256:a8dc572dec895e33ad9e1dd83de6e753b59c24ccb5ff98d40f6db8535f544a04

Observation 1b379fc0-4ef3-45e3-93b6-51201ac64726 · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.638441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.638441Z digest=sha256:0666955fc756d7f86e4ce383fd0eda43bde89ae65a4948fc4b3404371cbf2a80

Observation ba1f3937-444b-4ad9-b686-4a5cf4401404 · inbound

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images cites this paper.

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:32.559182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:43:32.559182Z digest=sha256:4d09f07d5bb37c51155833bbf413e5773d559dfabe4efa189fb18c09df188344

Observation 045796f1-b877-4933-81e6-b55437019dbc · inbound

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions cites this paper.

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:27.831298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:27.831298Z digest=sha256:a6a57a1013cb2fb4b0e94908dbb54836337e37dedd29bb93094e237b27cf5294

Observation 12899f82-3563-44a6-9a78-941a20b1afd3 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.206849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.206849Z digest=sha256:bfeba705edc916cf25247c7dae75237118691cbf9b004873c7a5af1d043f10dc

Observation aec1934d-55f9-4823-b709-84fa00b75b98 · inbound

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? cites this paper.

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:36:56.903250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:36:56.903250Z digest=sha256:88e35c0565f9998ba172d01ea066e8e87d39716c1221bd2d948ca251478d8d0b

Observation aa1eb206-ca7b-4363-b8fd-77b9c0e6ef69 · inbound

PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning cites this paper.

PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:34.729605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:34.729605Z digest=sha256:0cd30a6bad2b65101e8c502cb68646f46cb57b159577e8815d2f1d72d1ea7733

Observation a9cd6f8d-803e-4d7d-9736-8787ab4efb4a · inbound

Demystifying the Visual Quality Paradox in Multimodal Large Language Models cites this paper.

Demystifying the Visual Quality Paradox in Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:56:38.092536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:56:38.092536Z digest=sha256:4564393a774cad0a968bdb0434678943c7e7a613daa21cc260c26d28a615840d

Observation da717749-599d-4916-9ecb-f93dbace5d54 · inbound

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions cites this paper.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:42.178165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:42.178165Z digest=sha256:1f5bb29c39231db78aa04c73fade11a28ac8f642d17a36718873ed0ff9d28957

Observation a3c7d626-021b-4630-ba82-b41255c2a330 · inbound

Semantic Caching for Improving Web Affordability cites this paper.

Semantic Caching for Improving Web Affordability LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:11.715965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:52:11.715965Z digest=sha256:6d1be79c38f95da28baaa679522387fa7acd19fc97b8f9ade568feb164bfb520

Observation 0304ec94-346f-42ef-92bb-ffa955e7c5e6 · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:36.453279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:36.453279Z digest=sha256:addd2a9c10754e704e8fa878a5ac83b522b2eb0905f9f014d3549a60ba114c94

Observation 81c4469d-25a7-4db2-ac88-4ab13f38b3a3 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:27.690418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:27.690418Z digest=sha256:93b2a79c95c54c341958efa184658fea0a977cb2445baf2f436ff182c9adaed2

Observation 491e6990-29ec-4934-a82d-bdedd27f0979 · inbound

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning cites this paper.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.772845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.772845Z digest=sha256:29b419647212bb23e5eec5743db974142224979dbd86e8b23b098e8d78c76129

Observation 004e1b95-1576-4438-bb11-86d5d535fae5 · inbound

Room Scene Discovery and Grouping in Unstructured Vacation Rental Image Collections cites this paper.

Room Scene Discovery and Grouping in Unstructured Vacation Rental Image Collections LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:25:04.533347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:25:04.533347Z digest=sha256:bd96aa0898bc569d5cdff4651207839566cb05a599efebccefcb26b847100ca7

Observation 3b25b168-8e05-4ae4-9e0d-dca0b148fd84 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:52:07.952982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:027805de6bb9ee012e193184b705e2cb2086e2ea278466c74daa41802053a79b

Observation e21e9a31-6b36-4a6f-90ac-6cc483401ae9 · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.955283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.955283Z digest=sha256:bc400656417caa230f9ba3bde7b3e0906b6c602c66a0c5d79828af15f0e924f3

Observation a86a8290-85a0-4b46-9027-cb17eeb6047d · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.659753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.659753Z digest=sha256:aeedbec09259f236fd8535cef616bce872c24e567fb8a33e7081c3bca55e7539

Observation 34958276-5601-46d2-a9fe-fc054773bdba · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.433115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.433115Z digest=sha256:df9e4a32414400eb2fcf543b0167da2f31428d60ddd7e05bc25f401e5a5f8b51

Observation 8ad19dd6-f942-4ace-90d9-dab0c91f6bc5 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:36.992645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:36.992645Z digest=sha256:8d9f9e2a4c6be3456349b96b334f7d59e7a8ba4fed3710e7b87cbb1c1aa498fd

Observation f989038e-d8a4-4f79-a1a3-12bf90ba02b3 · inbound

Pedestrian Intention Prediction via Vision-Language Foundation Models cites this paper.

Pedestrian Intention Prediction via Vision-Language Foundation Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.791343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.791343Z digest=sha256:b655b1d637de62a050d21b3a501210335b4a88f972313eb65d7a205c9c92feb9

Observation 66ff6505-3b8c-489a-93fc-f5953c148163 · inbound

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval cites this paper.

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:00.229869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:00.229869Z digest=sha256:87cdd42f4630e0948f71ca56135b2652dc6e49b1da60dfbbcc27fd37cc3dcc9e

Observation 4538fcbd-2ec2-4ba8-94e9-b6a0e76935b6 · inbound

LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning cites this paper.

LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:46.949417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:21:46.949417Z digest=sha256:79655b7872fdabbe5c3212c241691661ef14a2b70cf5e193fd513de76d091eb7

Observation b730cbda-5bef-43f6-8798-4771a4d6016f · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.915011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.915011Z digest=sha256:c00248dea8cc7cf3280e7ef539aeea1e4cb80d25857b9e1418cc244d302f9ec2

Observation 41f92863-7248-4f98-84ce-187f0a7959ac · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.322384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.322384Z digest=sha256:67a53c4d341ae236cb909f6994de8ca375cd63ac0a0f6f2f1a38713cc828ab90

Observation 025ca9d1-b964-4484-94ed-e1e54b55c98b · inbound

FaceLLM: A Multimodal Large Language Model for Face Understanding cites this paper.

FaceLLM: A Multimodal Large Language Model for Face Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:39:41.509601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:39:41.509601Z digest=sha256:7df2c0f8334a1991bcec44a5b8f5b53df1c54a7a35c02aa762f4069c19c1b049

Observation 5ba75bd4-02e0-44dc-8bcf-3d33d8a96cdd · inbound

Describe Anything Model for Visual Question Answering on Text-rich Images cites this paper.

Describe Anything Model for Visual Question Answering on Text-rich Images LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.854998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.854998Z digest=sha256:b6573e82e6f3843943af04c7f0f4a2816510a3e0d9db4658c5c3982d51a5b865

Observation 0188582b-0aa0-412d-bdb3-cdcba7394859 · inbound

InterAct-Video: Reasoning-Rich Video QA for Urban Traffic cites this paper.

InterAct-Video: Reasoning-Rich Video QA for Urban Traffic LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:50.665017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:53:50.665017Z digest=sha256:d4d425424e7ffbad9e21bc97737f6485ee35276f7d7f5d503c817bfba07c9674

Observation 3c20f704-e223-4e5b-a652-0cc82321eb7e · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.673817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:a4311ac7811f10fac8a6591f64d41da572a45494f7b7ede6db10e63506e7d10d

Observation 50dfb0b0-7468-4715-a1fe-c5ddce93e998 · inbound

LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs cites this paper.

LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:57.692643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:57.692643Z digest=sha256:a7cfca452e00c79ec66fc390ccdb9950fc055c0a56097a7e16b591f2c736e3f6

Observation f7cad20f-21af-4df1-92e8-4210147ffecf · inbound

Object-centric Video Question Answering with Visual Grounding and Referring cites this paper.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.124422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.124422Z digest=sha256:1b690aa675409f5995713ab0ee5d30a8ccb7aef3e6390a6bd42469e2e9f0b5e3

Observation c5af75d5-94d6-40be-ab56-232a5a907760 · inbound

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO cites this paper.

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:40:09.004173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:40:09.004173Z digest=sha256:9cd7fb4ae10989538e6251cf1f840d50ffeb8267ea0bb2157834e6f0a1e43061

Observation ede1ea7f-a126-498b-8996-33326c899b63 · inbound

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces cites this paper.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.071331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.071331Z digest=sha256:2caa3b50f7ca10617ee4afd0099aa283474d9d3bd50f9dd0dc0b2f258f75944e

Observation e6a316b7-1e19-4fc6-89ca-7f405d5bf9e2 · inbound

Volume-Distance-Ratio Asymptote and Spacetime Inextendibility for FLRW Spacetimes cites this paper.

Volume-Distance-Ratio Asymptote and Spacetime Inextendibility for FLRW Spacetimes LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T04:36:32.561417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:36:32.561417Z digest=sha256:b3a55dd50b8c6c4d95db592d198483144cb26962a5c1a4b7e776387c824ab4a2

Observation 910ab3a5-948e-4b7d-bcc9-3162ce11c2af · inbound

OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset cites this paper.

OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:26:57.847097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T01:24:15.717972Z digest=sha256:1d4d0aa21448b93b09605844d770db792825419f4a2be99c041646facda735b0

Observation 7ef91280-3fa7-4007-aa6a-39f50dfe993e · inbound

Multimodal Video Emotion Recognition with Reliable Reasoning Priors cites this paper.

Multimodal Video Emotion Recognition with Reliable Reasoning Priors LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.553102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.553102Z digest=sha256:fffb7c3a2dcde70a494e4ebb491975e84092ee30fb1f51f82110914bba1e72d6

Observation 92fe071e-0cef-4f1d-a299-84da41f4e7b0 · inbound

AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated Content cites this paper.

AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated Content LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:08.269070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:43:08.269070Z digest=sha256:03bdda321c3d87bac8942f68ef5b6797bd6849b43285a107ae2734d1dfa9e83a

Observation 177dc122-8273-4c03-870e-e8cf638b2e03 · inbound

Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild cites this paper.

Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T21:53:40.542525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:53:40.542525Z digest=sha256:ede49ff2ff00417e36bba54a820c1f7225c60c5b8715be7d2aaeb77b92a24618

Observation e3899d92-19a9-4b3f-b1c8-8bab760c7cf2 · inbound

IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning cites this paper.

IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:12.525660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:12.525660Z digest=sha256:fe9dab98df6f239f2a51e41c7d6f8ca0b47ac270d26ea1562956866a382170ad

Observation 8d9f69a4-1136-4732-8936-b794417cb6bd · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.799225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.799225Z digest=sha256:3d5a00d0e1990b523a69469467f935b17cd279a24b11beca90172c87a6e358ad

Observation ec322ad7-3d90-4a70-ac13-c2e78f070e97 · inbound

Region-Level Context-Aware Multimodal Understanding cites this paper.

Region-Level Context-Aware Multimodal Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.167701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.167701Z digest=sha256:97ed67689acb1ae62a22c53b4bc9a22c2c5cf14e75badf9e937becb0347b85dd

Observation 5eaaec05-f53b-43aa-9ac0-424881140064 · inbound

AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings cites this paper.

AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T19:02:49.929085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:02:49.929085Z digest=sha256:6c6980cd966467f527ca56613ac6b2de5680eea2f7704f082b758194d6ffd2a8

Observation 15ecf628-7c09-4a99-a955-bc2e90e71ac8 · inbound

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark cites this paper.

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:07.599560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:07.599560Z digest=sha256:3166647460e71f94e2bbe762a42f7a87fa458c1195ec8e3f638ebb5143d13404

Observation c8c289cd-b6e4-4188-bf6e-c8e9db1b9195 · inbound

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models cites this paper.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.040398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.040398Z digest=sha256:cc0b14ab764cbf7f7e059639a1ede4d4b4f08697f3cdfd2db5422e7bbbb426a2

Observation 7f42c18c-c8e8-4144-8a66-04453bac6a48 · inbound

Ego-centric Predictive Model Conditioned on Hand Trajectories cites this paper.

Ego-centric Predictive Model Conditioned on Hand Trajectories LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:29:25.465694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:29:25.465694Z digest=sha256:68542496137990533370f9b104dbe1abd611d6c6844a53b147eae6fc086300f8

Observation 4097d819-b341-41b9-892a-7aae4935978f · inbound

CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving cites this paper.

CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:11:50.824721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T20:10:43.416488Z digest=sha256:17fe4d03c725c2bf0ffb940fea1dc80b0b10d4a9762254225dc0a08733afb280

Observation b4468c9a-0187-4eb9-abe4-590c55e5e6a5 · inbound

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs cites this paper.

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:12:44.181052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:12:44.181052Z digest=sha256:d664029f99e979098f195eceb70fd1fd5cf9f600e3d9948d3d602df8bfa903cb

Observation 7ab5c75e-59a5-4bb6-8b05-4d7047b3d975 · inbound

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images cites this paper.

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:42:47.591189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T17:37:31.837022Z digest=sha256:a36504d62c2c682fe169806afa5265eb2208998b1df5cabf4e0c9a6717583e24

Observation 6e57d26b-0ecc-47ba-9c55-5ba61a521cf0 · inbound

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning cites this paper.

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:56.239297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:56.239297Z digest=sha256:946a87dd4587f7be67f669daaa8d953cb5a170db754d37c61d7ea4851ab35313

Observation 7a8ae2c0-cb60-44fb-9469-3644af3ce3ea · inbound

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction cites this paper.

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:11:27.506328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T14:09:22.942238Z digest=sha256:fae6bd90ac48590a9122491f667fa4ab9368f8fe18e2c0a34ddc56defa1f9f26

Observation a56dd022-f6b6-4b89-b253-71bc8c61c996 · inbound

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency cites this paper.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.874949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.874949Z digest=sha256:60cf8e0d9b0c7411eba897ede23c8a2d612777fc0209f3fd7b80e9914b6a2a78

Observation 107633b8-832d-4086-b8c4-cb1733ce28b7 · inbound

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation cites this paper.

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:44.073688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:44.073688Z digest=sha256:9bdb6aad08a21ed5466ec5e76e11259a5184b3d6b70667a866d0f2c2c0bd05f1

Observation b54e816c-2572-4492-be4c-60f68ca312e9 · inbound

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding cites this paper.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.428946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.428946Z digest=sha256:aeeac1fd33876b1fac8b691f488fb51a2e4d1a392b6054e1c4d5fe9384373114

Observation a20b2645-8f0e-420d-8237-4c5bf9beeb28 · inbound

Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation cites this paper.

Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:19:09.881697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T06:16:53.140133Z digest=sha256:70e1bd00055b9accb4ecf5fd91cce6ad602cb001dcdc598db31e15c965695715

Observation b3603a2e-65cf-45c4-8f50-e6e53fa82e62 · inbound

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension cites this paper.

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T18:15:05.459491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:15:05.459491Z digest=sha256:f5dd9c2558446cce9aae363033729a31bc6c280206d0b929217560617d4ee196

Observation 3c75648d-0ef4-49e2-b309-a9a3f18588ce · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.618430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:fe75739bb96aef177c2f6a0116a91df380c98409bc56e3d82e202f579168c2e9

Observation c46fbf43-cad5-4e06-b956-1bd23df14fe7 · inbound

Are vision-language models ready to zero-shot replace supervised classification models in agriculture? cites this paper.

Are vision-language models ready to zero-shot replace supervised classification models in agriculture? LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:18:32.238282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T21:15:20.717705Z digest=sha256:123fd92632dc1a736e93b4e9787b2eb22b6e2b8794ba93cc124bb6ea20f2e992

Observation c3be2a57-c51f-4003-8361-00e2d681c6c0 · inbound

BrepLLM: Enabling Large Language Models to Understand Boundary Representations cites this paper.

BrepLLM: Enabling Large Language Models to Understand Boundary Representations LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T15:37:20.281619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:37:20.281619Z digest=sha256:465c05c66dd0b6938f5aa222820366076d9fd73490df5eae077a04ba2ccf2c3e